Hands-on writing about local AI models, the software built on them, and the machines you run it all on.

Detailed close-up of a red circuit board highlighting electronic components and connectors.
AI & Local Models

KV Cache Quantization in llama.cpp for Consumer VRAM

llama.cpp KV cache quantization lets 8B models run 32k–64k context on 8–16 GB cards. The math, the -fa prerequisite, and the flags.