---
format: "aidr-story-markdown/v1"
id: "cf8a692268d7b087b06ede07e23b5370ae1745497258b0bb4f1dc916987a1eeb"
canonical_url: "https://aidr.today/cf8a6922?lang=en"
title: "Faster prompt lookup drafting in llama.cpp"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-26T19:57:24.000Z"
category: "Infra"
topics: ["llamacpp","inference","open-source","prompt-caching","memory-optimization"]
source_urls: ["https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/","https://news.ycombinator.com/item?id=49859982"]
summary: "Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x."
---

# Faster prompt lookup drafting in llama\.cpp

> [Open the canonical story](<https://aidr.today/cf8a6922?lang=en>)

**Published:** 2026-09-26T19:57:24.000Z
**Category:** Infra
**Topics:** llamacpp, inference, open\-source, prompt\-caching, memory\-optimization

## Summary

Four changes to the n\-gram caches of llama\.cpp make drafting up to 41\.6x faster, load the static cache up to 23\.5x faster, and lower peak memory up to 2\.65x\.

## Sources

- [Story source](<https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/>)
- [Discussion](<https://news.ycombinator.com/item?id=49859982>)

