Hint-based CoT faithfulness evals still mostly work on Claude

·LessWrong··

Thanks to Fabien Roger (Anthropic), who pointed out this system card mistake to me. This mistake will likely be fixed in the relevant system cards after this post comes out.This work was done by an automated research scaffold developed at Redwood Research. For this project, essentially all of the experiment ideas were designed by a human, and the scaffold only executed on the experiment ideas. We think this project is similar to or slightly below the level of rigor of a mid-MATS research update....

Read full article →

Related Articles

GCC steering committee announces AI policy
arto · Hacker News · 11h ago
Why is everyone trying to build a solid-state battery?
crescit_eundo · Hacker News · 10h ago
Stacked PRs are now live on GitHub
tomzorz · Hacker News · 6h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 1d ago
Why DNA damage from smoking and UV rays cause cancer in some but not others
gmays · Hacker News · 8h ago