Modern LLMs have tiny GPTs hidden inside them

·LessWrong··

Experiments into predicting GPT2 completions via Qwen modelsThis is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects.----Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only i...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 20h ago
AMD's random number generator can't generate a 0?
BruceEel · Hacker News · 7h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 12h ago