When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

·LessWrong··

By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele PoppiThis post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available hereTL;DRThe Setup: We tested three distilled large reasoning models (DeepSeek-R1-7B, Phi-4-Reasoning-Mini, Qwen-4B-Thinking) against a fixed attacker attempting to extract restricted information under constant adversarial pressure in a multi-turn conversation.The Findings: We have particularly identifi...

Read full article →

Related Articles

Measuring the sloppiness of code
doppp · Hacker News · 14h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
HuggingFace: Security.txt
yarapavan · Hacker News · 13h ago
Rune is now open source
ernestrc · Hacker News · 12h ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago