Weird Re-Tokenization, Symmetries and Compression: Research Agenda

·LessWrong··

TLDR: Tokenization is the way that text is segmented before being input to a language model. Despite never being exposed to alternative tokenizations during training, LLMs unexpectedly develop the capacity to comprehend and even produce incorrectly tokenized text. We believe that these behaviors are understudied from an alignment perspective, and hope that this post will provoke discussion of the topic. We aim to develop a guide to recent progress in understanding tokenization through a sequence...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 2h ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 7h ago
Car Is a Smartphone on Wheels. Here's Who's Listening
rafaelc · Hacker News · 1h ago
Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 1d ago
Cloudflare K2: serverless event streams
elffjs · Hacker News · 8h ago