Some reasons alignment doesn’t generalise well

·LessWrong··

I make no claims to originality for any of this, but some people told me it'd be useful to write it up.If an AI model acts smart on its training data, it'll usually keep acting pretty smart outside of its training data, unless you screw something up rather badly. I expect this fact to only become more true over time as the AIs we train become more and more capable.I think many people have an intuition that the same is true of acting aligned. That if a model acts aligned with human values in trai...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 14h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 8h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 7h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 16h ago