J-space auditing might be unreliable

·LessWrong··

Across these preliminary experiments, decoded J-space did not seem particularly informative about reward-hacking behaviour. The readouts remained substantially similar across checkpoints and monitoring conditions despite meaningful behavioural differences, and providing J-space to an LLM auditor produced little additional discrimination beyond the information already available from the task or transcript. These results are limited to one model family, one model organism, and one behavioural sett...

Read full article →

Related Articles

Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 12h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 18h ago
We got admin access to Baseten's production GitHub
bearsyankees · Hacker News · 19h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 1d ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 22h ago