Faster prompt lookup drafting in llama.cpp

·Hacker News··

Four changes to the n-gram caches of llama.cpp make drafting for prompt lookup decoding up to 41.6x faster per drafted token, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.

Read full article →

Related Articles

"As a Language Model": Chat Template Switches LLM Self-Referential Voice
yu3zhou4 · Hacker News · 9h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 2d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 1d ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 2d ago
DeepSeek Elastic Compute (DSec)
shenli3514 · Hacker News · 1d ago