When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele PoppiThis post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available hereTL;DRThe Setup: We tested three distilled large reasoning models (DeepSeek-R1-7B, Phi-4-Reasoning-Mini, Qwen-4B-Thinking) against a fixed attacker attempting to extract restricted information under constant adversarial pressure in a multi-turn conversation.The Findings: We have particularly identifi...
Read full article →