Learned Chain-of-Thought Obfuscation Generalises to Unseen Tasks

·LessWrong··

TL;DRTraining against a CoT or summary-only monitor can lead to obfuscation of dangerous reasoning in unseen tasks. This strengthens the “don’t train against a monitor” claims.Figure 1. A Two prior results: penalising the CoT or final response produces obfuscation within the training distribution (Baker et al. 2025; Skaf et al. 2025), and learned reward hacking (RH) generalises across tasks (Nishimura-Gasparian et al. 2024). We combine them and find that obfuscation itself generalises: pressure ...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 9h ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
Pi 1.0
sergiotapia · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 4h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago