A Four-Axis Bayesian Epoch Capabilities Index with Human Baselines
This is a crosspost from the General-Purpose AI Policy Lab research blog.The Epoch Capabilities Index compresses many benchmark scores into one for each model, following the framework of the Rosetta Stone paper. In a previous post, we added human baselines to the same scale to see how the models compare to humans. But one issue is that most humans score near-perfectly on abstract reasoning benchmarks like ARC-AGI or VPCT and sit near chance on GPQA-type benchmarks, while many models show the opp...
Read full article →