Constraining the capacity of physical side channels for AI verification and security

·LessWrong··

Based on recent events, including Dario Amodei's essay on pacing the frontier and the subsequent response, a coordinated slowdown in AI capability developments is now in the Overton Window (however it may go down in Washington longer-term). An important aspect of such an effort is the verification of such coordinated measures, which is as yet an unsolved problem in many areas. Borrowing language from arms control and nuclear safeguards, verification involves confirming that claims by a frontier ...

Read full article →

Related Articles

What is it like to be a neural net?
David Balduzzi · LessWrong · 22m ago
Agents let AI safety share experiments hourly, not just papers monthly
Jason Fantl · LessWrong · 26m ago
Measuring alignment drift via trajectory prefixes
Owen Terry · LessWrong · 27m ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 34m ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago