AI & Local Models

Llama 3.1 8B Quant Shootout: 4‑bit to 6‑bit on RTX 3060

On an RTX 3060, Q5_K_M runs Llama 3.1 8B nearly as fast as 4-bit with better output. See the numbers for Q4, Q5, Q6.