Using Base-LCM to Monitor LLMs

·LessWrong··

Epistemic status: experimental results. This is an exploratory work examining an alternative approach to the interpretation of language models.SummaryWe aim to determine whether the LCM model — which predicts sentence embeddings rather than token embeddings — can predict the outputs of LLM. We compare four architectures, the most efficient model predicts the following paragraphs with a cosine similarity of 0.53.The code is available hereMotivationThis work is driven by the need to monitor and un...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 14d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 16d ago
The Hobbesian Bootstrap Paradox in Frontier AI
Claudio Di Meglio · EA Forum · 19d ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 22d ago