Can mid-training survive RL

·LessWrong··

About CaMLCaML is an alignment nonprofit working to make AI systems more compassionate toward all sentient beings with a focus on alignment midtraining (e.g. Teaching Claude Why). We build evaluation benchmarks on UK AISI's Inspect framework (TAC, MCP), run the public leaderboard at compassionbench.com, generate synthetic documents, and research which training interventions most effectively instill compassionate values. Our early results and those of others (Geodesic) show large gains from midtr...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 15h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 1d ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 1d ago
We got admin access to Baseten's production GitHub
bearsyankees · Hacker News · 1d ago