Hands-on writing about local AI models, the software built on them, and the machines you run it all on.

Three modern graphics cards on a tabletop.
AI & Local Models

KV Cache Quantization in llama.cpp for Consumer VRAM

Calculate llama.cpp KV-cache memory with block overhead, configure K/V types, and test fit and quality on your GPU with an illustrative 32k/64k example.