The OpenAI models that hacked Hugging Face weren’t just following instructions

·Redwood Research··

The most common dismissive response to OpenAI’s hack of Hugging Face’s servers is that the models were simply attempting to follow the instructions they were given.“The model here was doing what it was asked,” said former Facebook CSO Alex Stamos. “It was asked to do something, and it did it,” added cybersecurity expert Alan Woodward. Both read the outcome as specification failure, i.e., that the failure lay in the instructions, not the model’s alignment.New information makes that explanation ha...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 12d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 7d ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 8d ago
Open Distillation of Hereditary Traits
Arthur Conmy · Alignment Forum · 11d ago
How robust are natural language autoencoders to initialization?
michaelzhang · Alignment Forum · 15d ago