Linear probes tell you where quantization will hurt

·LessWrong··

Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I am not inventing anything; I am just connecting quantization with a very general idea from mech interp: the model does the easy syntax work first, the semantic work in later layers, and the prompt-relevant task work in the latest layers. That's the main idea I lean on in this proje...

Read full article →

Related Articles

My security camera shipped a GitHub admin token in its login page
hhh · Hacker News · 16h ago
JEP 541: Deprecate the macOS/x64 Port for Removal
pmg1991 · Hacker News · 11h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 1d ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 1d ago
Postgres LISTEN/NOTIFY actually scales
KraftyOne · Hacker News · 8h ago