When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

·LessWrong··

By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele PoppiThis post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available hereTL;DRThe Setup: We tested three distilled large reasoning models (DeepSeek-R1-7B, Phi-4-Reasoning-Mini, Qwen-4B-Thinking) against a fixed attacker attempting to extract restricted information under constant adversarial pressure in a multi-turn conversation.The Findings: We have particularly identifi...

Read full article →

Related Articles

US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Should you wash your solar panels?
surprisetalk · Hacker News · 1d ago
Kimi-K3 Technical Report [pdf]
vinhnx · Hacker News · 23h ago
New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 1h ago
DMARC Has Been Public Since 2012. 68.4% of Domains Still Don't Enforce It
adulion · Hacker News · 4h ago