a recurrent llm is quite easy to interpret but complex to steer

·LessWrong··

TLDR;Ouro-1.4b-thinking is broadly interpretable with logit lenses and linear probes. It's also steerable but does 'clean' foreign concepts out of the residual stream if they're injected before the last loop. This could have nasty implications for safety.Code + data: https://github.com/mild-rgb/ouro-experiments + https://huggingface.co/datasets/mild-rgb/ouro-1.4b-thinking-evalsIf you're not familiar with the Ouro family recurrent models, I recommend taking 5 minutes with your favourite AI agent ...

Read full article →

Related Articles

Two-Tier Encryption in the UK – Identical Apple Devices, Different Protection
ReturnoftheHack · Hacker News · 6h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 23h ago
Early rogue AI agent activity and attempts to hack found on urlquery.net
snikolaev · Hacker News · 11h ago
Claude Opus 5.5
km144 · Hacker News · 2d ago
Linux support is coming to Snapdragon X2 Series
aaronday · Hacker News · 18h ago