This article reports no RTX 3060 measurements. It gives a repeatable test for comparing Llama 3.1 8B GGUF quantizations on a 12 GB RTX 3060: record storage, VRAM, prompt processing, generation speed, and task quality for the exact artifacts and context settings you run.
What the Quantization Names Mean
GGUF is a model container format used by llama.cpp and compatible runners. Names such as Q4_K_M, Q5_K_M, and Q6_K describe tensor encoding schemes, but they are not simply “every weight uses exactly four, five, or six bits.” llama.cpp documents effective bit rates for the underlying schemes and allows different encoding choices for individual tensors.
The suffix is part of the llama.cpp quantization preset name. The project lists Q4_K_M and Q5_K_M as distinct presets, and its documentation notes that a file can contain tensors with different encodings. Treat the converter, source revision, and exact GGUF artifact as part of the result; a preset name alone does not describe every tensor.
The official Ollama Llama 3.1 tag listing illustrates the file-size progression for one set of artifacts: its 8B instruct tags list Q4_K_M at 4.9 GB, Q5_K_M at 5.7 GB, Q6_K at 6.6 GB, Q8_0 at 8.5 GB, and FP16 at 16 GB. Those are model artifact sizes, not total GPU memory use.
Why Disk Size Is Not VRAM Use
An inference process also needs:
- a key-value cache that grows with context;
- compute and graph buffers;
- runtime/backend allocations;
- memory for the display or other GPU applications; and
- sometimes CPU memory if part of the model is not offloaded.
Do not estimate a 32K-context fit from an idle 2K run. Measure each context configuration, verify the offloaded layers, and retain the runner log.
Pin Every Input
Create a manifest before benchmarking:
GPU model and VRAM:
GPU board vendor:
Power limit:
NVIDIA driver:
CUDA version:
Operating system:
llama.cpp commit:
Model repository and revision:
GGUF filename:
SHA-256:
Quantization:
Chat template:
Context setting:
KV-cache types:
GPU layers:
Flash attention: on/off
Use quantizations made from the same source model and, ideally, the same conversion process. A community fine-tune or different importance matrix makes the comparison about more than bit width.
Benchmark Prompt Processing and Generation Separately
Build llama.cpp according to its current instructions, then use the included llama-bench tool. A starting command for one file is:
./build/bin/llama-bench \
-m llama-3.1-8b-instruct-q4_k_m.gguf \
-p 512 \
-n 128 \
-r 5 \
-o json
Run the same command for Q5_K_M and Q6_K. The official tool documents -p as prompt processing, -n as text generation, and -r as repetitions. Its timings exclude tokenization and sampling, so do not present them as complete application latency.
Test more than one context depth if context is part of the recommendation. llama-bench supports pre-filling the key-value cache with -d; record the exact command and confirm that the intended GPU backend is actually doing the work.
Store raw JSON and the command line in version control or an attached archive. A rounded table without raw output cannot be independently checked.
Measure Memory Directly
Sample nvidia-smi during load, prompt evaluation, and generation rather than quoting model file size as VRAM consumption:
nvidia-smi \
--query-compute-apps=pid,process_name,used_memory \
--format=csv
Also record total card use because the desktop may consume memory outside the compute-process list. State whether the RTX 3060 drives a display and whether other GPU applications were closed. On Windows in WDDM mode, NVIDIA says per-process GPU memory use may be unavailable; do not report an unavailable value as zero.
Report the highest sampled memory with the context setting and sampling interval; a short-lived peak may fall between samples. “Fits on a 12 GB card” should mean the intended workload completed without hidden CPU offload, not merely that the process started.
Evaluate Quality with Saved Tests
Throughput alone cannot choose a quantization. Build a task suite that can be scored:
- deterministic code tasks with unit tests;
- structured output checked against a JSON schema;
- document questions with known supporting passages;
- long-context facts located near different context positions;
- instruction-following constraints; and
- refusal questions whose answer is absent.
Use the same prompt template and generation settings for every artifact. Repeat stochastic tests or use a fixed seed when supported. Save every response, including failures.
Perplexity on a disclosed corpus can measure distributional change, while task checks measure usefulness for the application. Neither justifies broad statements such as “Q5 is indistinguishable from FP16” unless the scope and uncertainty are shown.
Publish the Complete Results
Use a table that keeps the performance numbers connected to their conditions:
| GGUF SHA-256 | Quant | Context | Prompt tok/s | Generation tok/s | Highest sampled VRAM | Task pass rate | Runs |
|---|---|---|---|---|---|---|---|
| exact hash | Q4_K_M | tested value | mean ± spread | mean ± spread | sampled max | disclosed suite | 5+ |
| exact hash | Q5_K_M | tested value | mean ± spread | mean ± spread | sampled max | disclosed suite | 5+ |
| exact hash | Q6_K | tested value | mean ± spread | mean ± spread | sampled max | disclosed suite | 5+ |
Include the date, thermal conditions, failures, and whether the difference exceeds run-to-run variation. A three-token-per-second gap is not meaningful if the measurement spread is larger.
Choosing a Starting Point
Lower-bit artifacts generally use less storage and memory; higher-bit artifacts generally retain more numerical precision. That broad trade-off can guide the first download, but the best operating point depends on the task and context budget.
Start with an artifact that leaves enough headroom for the intended context, then move upward only if a retained quality test shows a benefit. Move downward if memory pressure, partial offload, or concurrency is the constraint. The outcome is a configuration decision, not a permanent ranking of Q4, Q5, and Q6.