Training Models to Predict and Explain Their In-the-Wild Behavior

·LessWrong··

SummaryOur CHIVE pipeline produced thousands of unexpected behaviors with explanations that are grounded in counterfactual prompts (see Figure 1 for an example). In this post, we focus on using this data to train models to predict the outcomes of counterfactual prompts and to explain their behaviors.We build two training targets from our CHIVE-generated data (see Figure 2 for examples): counterfactual prediction, where the model answers a binary question ("would this specific edit to the prompt ...

Read full article →

Related Articles

Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 11h ago
Solving the Jane Street reverse engineering challenge
anitil · Hacker News · 7h ago
US Military disables ad trackers on troops' phones
tencentshill · Hacker News · 4h ago
Google AI Mode shows same products 21.6% more expensive than traditional search
DeepLogin · Hacker News · 6h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 20h ago