Steering towards “automated grading” degrades alignment

·LessWrong··

TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret...

Read full article →

Related Articles

Pre-Release of Polars 2.0
komape · Hacker News · 12h ago
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
jakobgreenfeld · Hacker News · 1d ago
Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2014)
DamonHD · Hacker News · 12h ago
Paint.net 5.2 alpha now runs on Linux
judah · Hacker News · 1d ago
Aging brains blend memories together instead of just forgetting them
mdp2021 · Hacker News · 1d ago