Should we combine protocols for AI Control Research?

·LessWrong··

In AI control[1] research, we've developed many protocols: trusted monitoring, untrusted monitoring, resampling, etc. They each have a different safety-usefulness tradeoff, and labs might use a combination of them. The reason: they might want the maximum usefulness from their AI models at a minimum acceptable safety (eg 90%), and only using a single protocol may not get both the usefulness they want and the safety they need.We typically model combining protocols as a Defer to Trusted, where we f...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 18h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 14h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 13h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 12h ago
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
HenryNdubuaku · Hacker News · 10h ago