Study 3: Steering welfare-relevant directions moved the representation, but not [detectably] the behavior

·LessWrong··

Epistemic status: an exploratory report. These results are from the calibration process intended to produce a preregistration for the third study in my series on welfare-relevant indicators. Calibration showed the planned procedure wasn't worth running, so I am publishing the calibration data and analysis instead.TLDR: steering moved the frozen directions' projections linearly, but no direction produced a judged-behavioral effect distinguishable from zero or from a random direction of the same n...

Read full article →

Related Articles

The case against JPEG XL
contact9879 · Hacker News · 1d ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 2d ago
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
tosh · Hacker News · 15h ago
Why don't machine learning research agents overfit?
Betelbuddy · Hacker News · 10h ago
Distributed Systems Classics (2017)
grep_it · Hacker News · 11h ago