Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs

·LessWrong··

TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’s influence flows through such a narrow, monitorable channel, we argue that this achieves near-maximal safety in our BashArena setting. We also discuss the general concept of infor...

Read full article →

Related Articles

US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Should you wash your solar panels?
surprisetalk · Hacker News · 12h ago
Kimi-K3 Technical Report [pdf]
vinhnx · Hacker News · 10h ago
The Strongest El Niño Ever
ndsipa_pomu · Hacker News · 1d ago
Exploiting Volvo/Eicher's fleet platform to gain control over all users/vehicles
EatonZ · Hacker News · 10h ago