common : read ngram cache parts by reference instead of copying by jadidbourbaki · Pull Request #2 · jadidbourbaki/llama.cpp

· GitHub

1 min read Original article ↗

Conversation

JohannesGaessler

alainnothere added a commit to alainnothere/llama.cpp that referenced this pull request

Sep 28, 2026
…ast customer's homework

try_draft copied whole unordered_map parts by value on every drafted token; they are read by reference now, 8.54 to 0.64 us per drafted token in lookup-stats with a static cache (jadidbourbaki#2, plus lemire's threshold pre-check from ggml-org#12). begin() was a no-op, so a reused slot drafted from the previous request's n-grams at 13% acceptance; it now clears the context cache (ggml-org#27866), T2 15.2 to 39.3 t/s. finished context caches go to ngram_cache_done so the -lcd write-back still sees every request.

Repository owner locked and limited conversation to collaborators

Sep 28, 2026