Prism: Automating Science-of-Evals Research

·LessWrong··

tl;dr – we present [Prism], a scaffold for automating science-of-evals research: work that makes the evaluation the primary object of study. The scaffold provides Claude Code with sub-agents and resources for carrying out scientifically rigorous investigations into eval dynamics and, by extension, model behaviours.We talk through an autonomous Prism run on the Agentic Misalignment setting which demonstrates how minor perturbations to GPT-4.1's prompt cause the model to adopt more indirect method...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 9h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 9h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 17h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 8h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 12h ago