Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits

·LessWrong··

Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model’s default behaviour. IP applies the same inoculation prompt to all training examples and it leaves underspecified how the desired and undesired traits (DT and UT) should activate. Two failures follow:The model develops backdoors: conditional vulnerabilities through which prompts that do not directly request the UT can still elicit it. We call this UT leakage.The desired trait weakens under ord...

Read full article →

Related Articles

US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 13h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 6h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago
Adults over 65 will outnumber children by 2029
brandonb · Hacker News · 7h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 1d ago