A non-generative model as a trusted monitor for AI Control: Testing TypeSafe's Jev
TL;DRTypeSafe AI has introduced Jev - a new class of frontier model trained to make fast, structured decisions, rather than generating free-form text like a chatbot. It takes unstructured state as input and returns type-safe, structured outputs with confidence scores.I aim to use Jev as the trusted monitor of the ControlArena APPS backdoor setting - to analyze how a non-reasoning model performs as a cheap alternative.One yes/no question gives AUROC 0.976 against LLM-written honest code and catch...
Read full article →