Multi-Agent Coordination Lets Us Pour More Compute Into Post-Training

·LessWrong··

Training models to coordinate across multiple agents gives us a way to productively spend substantially more compute during post-training than current single-agent RL setups.From the recent Dwarkesh Podcast with Noam Brown:Noam BrownBut that’s one data point. We don’t know how long it would take a single agent to solve Navier-Stokes, because we haven’t done that experiment yet. Maybe we will, but that’s also only one data point.If we want to do a thorough ablation, the experiments are just too e...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 9h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
Ask HN: Is it impossible to disable Siri on macOS 27?
semidror · Hacker News · 15h ago
HERMES radio enables voice and data communication over vast distances
SamuraiLion · Hacker News · 12h ago