Self-Modeling Interventions Modulate Emergent Misalignment

·Hacker News··

Code: https://github.com/atagade/sgtr-em • Model checkpoints: https://github.com/atagade/sgtr-em/blob/main/MODEL_README.md …

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 19d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 13d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 14d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 21d ago