OpenAI President Brockman says HuggingFace incident model had not been alignment-trained

·LessWrong··

On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet." I assume Brockman is specifically referring to the "Highly Persistent Internal Model" as it's called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not. Discuss

Read full article →

Related Articles

The case against JPEG XL
contact9879 · Hacker News · 16h ago
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
tosh · Hacker News · 5h ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 1d ago
Ubuntu 26.10 completes transition to Rust-based coreutils
theanonymousone · Hacker News · 3h ago
JetKVM Mini
taubek · Hacker News · 1d ago