Model ethology for understanding average case alignment

·LessWrong··

TLDR:We want to tell if models are aligned enough to use in important use cases. I call this property trustworthiness. Our current ways of assessing trustworthiness seem mostly based on case studies or vibes. I think we should aim to systematically search for realistic honeypot cases where models misbehave. We should do this by first characterizing the types of misaligned behaviors models engage in, and the conditions and frequency with which they occur. This will take a lot of data.I think we s...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 3h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 14h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 5h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 4h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 4h ago