A Multi-Agent Extension for Petri

·LessWrong··

IntroPetri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the Auditor, which runs the evaluation, the Target agent, the model being evaluated and the Judge, which evaluates the specified behaviour. A natural language description of the desired evaluation is given to the Auditor which then designs the evaluation by creating prompts for the Tar...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 8h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 8h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
South African diamond mines are closing due to weak sales and lab-grown stones
bookofjoe · Hacker News · 7h ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago