Reducing synthetic markers makes some SDF false facts linearly indistinguishable from pretraining-acquired knowledge
SummarySynthetic document finetuning (SDF) — the state of the art for implanting false beliefs into LLMs — produces training documents with features that distinguish them from pretraining text.These features (i.e. synthetic markers) are partially responsible for making SDF false facts distinguishable from pretraining-acquired beliefs by linear probes on middle-layer activations.Reducing synthetic markers enables SDF to implant subtly-implausible false facts that are indistinguishable from pretra...
Read full article →