Where does hint-following and concealment arise? A case study on OLMo-3 checkpoints

·LessWrong··

This work was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I'm grateful to Harshul Basava and Vanessa Ng for their feedback. This is an extension to a prior replication which can be found here.Introduction and MotivationIn an earlier post, I showed that the “necessity effect” of Emmons et al. replicates across eleven models, where LLMs readily follow simple hints, even incorrect ones, but when hints require actual computation, the mode...

Read full article →

Related Articles

My security camera shipped a GitHub admin token in its login page
hhh · Hacker News · 9h ago
JEP 541: Deprecate the macOS/x64 Port for Removal
pmg1991 · Hacker News · 4h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 1d ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 1d ago
Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
1vuio0pswjnm7 · Hacker News · 1d ago