MIT's New Method Flags AI Models Trained on CASM Without Generating It

·Hacker News··

MIT and Thorn built a 100% accurate probe that detects CSAM-tuned AI models by inspecting internal code—no illegal images needed.

Read full article →

Related Articles

Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 20d ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 1mo ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 1mo ago