How robust are natural language autoencoders to initialization?

·Alignment Forum··

Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model might be thinking. How robust are NLAs to these guesses? We change Claude's guesses in various ways and measure the impact on the NLA's statements as well as on reconstruction accuracy. We show that Qwen2.5-7B NLAs have some robustness to irrelevant statements and prevailing sentimen...

Read full article →

Related Articles

Self-Modeling Interventions Modulate Emergent Misalignment
gmays · Hacker News · 2h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 19d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 13d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 14d ago