Off-policy honesty training generalizes better than on-policy honesty training

·LessWrong··

This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026.All code related to the blog can be found in this repo.We investigate self-report fine-tuning (SRFT), a technique proposed by Li et al. to improve models' honesty. The technique works by doing SFT on 2-turn user-assistant chat transcripts. In turn 1, the model is asked a factual question, and lies with 50% probability; in turn 2, it is asked whether it lied, and either confesses the lie...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 8h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 2h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 5h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 3h ago
Tail-call optimization in C is relatively recent (2025)
prakashqwerty · Hacker News · 7h ago