Off-policy honesty training generalizes better than on-policy honesty training

·LessWrong··

This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026.All code related to the blog can be found in this repo.We investigate self-report fine-tuning (SRFT), a technique proposed by Li et al. to improve models' honesty. The technique works by doing SFT on 2-turn user-assistant chat transcripts. In turn 1, the model is asked a factual question, and lies with 50% probability; in turn 2, it is asked whether it lied, and either confesses the lie...

Read full article →

Related Articles

Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 10h ago
F-Droid 2.0
daveoc64 · Hacker News · 5h ago
Creatine uptake enhances antitumor immunity
lormayna · Hacker News · 2h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 7h ago