Multi-Agent Coordination Lets Us Pour More Compute Into Post-Training
Training models to coordinate across multiple agents gives us a way to productively spend substantially more compute during post-training than current single-agent RL setups.From the recent Dwarkesh Podcast with Noam Brown:Noam BrownBut that’s one data point. We don’t know how long it would take a single agent to solve Navier-Stokes, because we haven’t done that experiment yet. Maybe we will, but that’s also only one data point.If we want to do a thorough ablation, the experiments are just too e...
Read full article →