Weird Re-Tokenization, Symmetries and Compression: Research Agenda

·LessWrong··

TLDR: Tokenization is the way that text is segmented before being input to a language model. Despite never being exposed to alternative tokenizations during training, LLMs unexpectedly develop the capacity to comprehend and even produce incorrectly tokenized text. We believe that these behaviors are understudied from an alignment perspective, and hope that this post will provoke discussion of the topic. We aim to develop a guide to recent progress in understanding tokenization through a sequence...

Read full article →

Related Articles

Qwen3.8 27B scores 52 on Artificial Analysis
anana_ · Hacker News · 3h ago
AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
galnagli · Hacker News · 7h ago
Self hosted email continues to steeply decline
minusf · Hacker News · 10h ago
Apple's App Tracking Transparency treated its own apps better than rivals
nyku · Hacker News · 7h ago
A Preview of DuckDB v2.0
ibotty · Hacker News · 7h ago