Evaluating Red Team and Blue Team Capability for AI Control Research

·LessWrong··

This post suggests a methodology to measure red team and blue team capability in AI control research, where each team gets an ELO rating. The methodology can help answer questions like "Are monitors getting better faster than attackers?" We attempt to answer questions like these using runs on LinuxArena.Epistemic Status: High confidence that the method works and is a good standard for measuring monitoring and attacking capability. It is an extension of existing ELO methods and is very general, a...

Read full article →

Related Articles

LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 9h ago
Smartphone makers don't bother to comply with EU repairability requirements
mdp2021 · Hacker News · 4h ago
Asahi Linux on M3
mdp2021 · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 19h ago
Making a Python interpreter in 1024 bytes
azhenley · Hacker News · 16h ago