The article describes performance optimizations for prompt lookup decoding in llama.cpp, achieving up to 42x faster drafting and 2.6x less memory usage through techniques based on work by Daniel Lemire and Martin Ankerl. Additional contributions by Lemire increased overall speedup to 140x. Prompt lookup decoding uses n-gram models as a draft mechanism for speculative decoding to accelerate token generation.