A Multi-Agent Extension for Petri

·LessWrong··

IntroPetri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the Auditor, which runs the evaluation, the Target agent, the model being evaluated and the Judge, which evaluates the specified behaviour. A natural language description of the desired evaluation is given to the Auditor which then designs the evaluation by creating prompts for the Tar...

Read full article →

Related Articles

LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 22h ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 1d ago
Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 1d ago
Everyone Should Know SIMD
WadeGrimridge · Hacker News · 7h ago
GigaToken: ~1000x faster Language model tokenization
syrusakbary · Hacker News · 7h ago