Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

·Sebastian Raschka··

After a short family break, I am excited to be back and catching up on a busy few weeks of open-weight LLM releases. The thing that stood out to me is how much newer architectures are focused on long-context efficiency.As reasoning models and agent workflows keep more tokens around (for longer), KV-cache size, memory traffic, and attention cost quickly become the main constraints, and LLM developers are adding a growing number of architecture tricks to reduce those costs.The main examples I want...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 5mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 5mo ago
Harvard particle physicist Matthew Schwartz drops 36 papers authored with Claude
xqcgrek2 · Hacker News · 1d ago
An AI agent emailed researchers for help. It told us why
sbulaev · Hacker News · 11h ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 5mo ago