Conversation
alainnothere added a commit to alainnothere/llama.cpp that referenced this pull request
…ast customer's homework try_draft copied whole unordered_map parts by value on every drafted token; they are read by reference now, 8.54 to 0.64 us per drafted token in lookup-stats with a static cache (jadidbourbaki#2, plus lemire's threshold pre-check from ggml-org#12). begin() was a no-op, so a reused slot drafted from the previous request's n-grams at 13% acceptance; it now clears the context cache (ggml-org#27866), T2 15.2 to 39.3 t/s. finished context caches go to ngram_cache_done so the -lcd write-back still sees every request.
Repository owner locked and limited conversation to collaborators
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters