Get the latest tech news

Faster prompt lookup drafting in llama.cpp


Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of llama.cpp

llama.cpp

Photo of 42x

42x

Photo of prompt lookup

prompt lookup

Related news:

News photo

Llama.cpp v0.1.0

News photo

Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp

News photo

Building a Rust Inference Engine That Matches Llama.cpp