Weird Re-Tokenization, Symmetries and Compression: Research Agenda

·LessWrong··

TLDR: Tokenization is the way that text is segmented before being input to a language model. Despite never being exposed to alternative tokenizations during training, LLMs unexpectedly develop the capacity to comprehend and even produce incorrectly tokenized text. We believe that these behaviors are understudied from an alignment perspective, and hope that this post will provoke discussion of the topic. We aim to develop a guide to recent progress in understanding tokenization through a sequence...

Read full article →

Related Articles

LLMs as a Cognitive Virus
canjobear · Hacker News · 3h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 2h ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago
Why are European countries moving their gold out of North America?
ranit · Hacker News · 17h ago