The Case for Model Forensics

·LessWrong··

If we had a misalignment warning shot, would we be able to tell?Suppose an AI company catches their model taking an egregious action, like deleting oversight code that monitors its actions. Should they sound the alarm? A key piece of evidence to determine what to do next – such as what mitigations to take – is to understand why the model took the action. If the model was just confused (e.g. it may have been trying to reduce latency), a simple mitigation like a regex classifier that blocks destru...

Read full article →

Related Articles

Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 10h ago
F-Droid 2.0
daveoc64 · Hacker News · 5h ago
Creatine uptake enhances antitumor immunity
lormayna · Hacker News · 2h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 7h ago