Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.

Sign in to suggest edits

Key sources

  1. DISCUSSIONpptadversarynews.ycombinator.com
Markdown