Use this repeatable benchmark to compare local models on an RTX 4070 12 GB. It captures model hashes, context length, software versions, throughput, VRAM, and task quality so the result reflects the exact setup rather than a generic GPU ranking.
Plan Around 12 GB of VRAM
NVIDIA lists the desktop RTX 4070 with 12 GB of GDDR6 or GDDR6X memory and a 192-bit memory interface. Record board cooling, driver version, background GPU use, and power limits for the system under test.
For local inference, the 12 GB capacity must hold more than the model file:
- model weights;
- the key-value cache used for context;
- temporary compute buffers;
- any desktop or display workload using the same GPU; and
- sometimes an embedding or reranking model running beside the chat model.
A model file fitting on disk therefore does not prove that the same model will run entirely on the GPU at the desired context length.
Pick Comparable Models Before Testing
“Llama versus Mistral versus Phi” is not a controlled comparison until the exact artifacts are recorded. Use instruction-tuned models intended for the same task, then capture:
- the complete model tag or GGUF filename;
- the model repository and revision or checksum;
- parameter count and architecture;
- quantization type;
- declared context limit;
- prompt template or chat template; and
- license and model-card restrictions.
Do not compare one base model with another model’s instruction-tuned release and call the result a quality ranking. Do not describe Phi-3 Medium as “4-bit native training”: a 4-bit GGUF file is a quantized distribution artifact, not evidence that the original model was trained as a 4-bit model.
Model families and serving tools also change. If reproducing an older comparison, keep the old versions and label the result historical. If choosing a model for current use, begin again with currently available releases rather than silently substituting a newer tag.
Freeze the Software and Hardware State
Record enough detail to distinguish a model difference from an environment change:
GPU and VRAM:
NVIDIA driver:
CUDA runtime:
Inference runner and version:
Operating system:
CPU and system RAM:
Model tag or file SHA-256:
Context length:
Batch settings:
GPU offload setting:
Flash-attention setting:
KV-cache type:
Power limit:
Display attached to test GPU: yes/no
Rebooting is rarely necessary, but close other GPU-heavy applications and note remaining VRAM before each run. Verify where the model was loaded. With Ollama, ollama ps reports the processor split for loaded models; with llama.cpp, retain the startup log that shows the selected backend and offloaded layers.
Separate Throughput from Quality
There are at least three different performance questions:
- Load time: how long the runner takes to make a model ready.
- Prompt processing: how quickly an existing prompt or document is ingested.
- Token generation: how quickly new output is generated after ingestion.
Combining them into one “tokens per second” number hides the user experience. A retrieval workload with long source documents cares heavily about prompt processing; short interactive chat may care more about time to first token and generation rate.
Ollama’s non-streaming API response includes load_duration, prompt_eval_count, prompt_eval_duration, eval_count, and eval_duration. Save the complete JSON response rather than copying a number into a table by hand. If using llama.cpp directly, its official llama-bench utility can test prompt processing and generation separately and can emit JSON or JSONL.
For example, after building llama.cpp and placing one exact GGUF file at model.gguf:
./build/bin/llama-bench \
-m model.gguf \
-p 512 \
-n 128 \
-r 5 \
-o json
Run the same command for each model. llama-bench does not include tokenization or sampling time, so describe its output as a runner benchmark, not total application latency.
Test Context Length Instead of Assuming It
A model’s advertised maximum context is not a promise that every quantization and runner will use that context efficiently on a 12 GB card. Test the context sizes relevant to the workload—perhaps 2K, 8K, and 16K—while watching for partial CPU offload, out-of-memory failures, and sudden latency changes.
For each context size, record:
- GPU memory immediately after load;
- GPU memory after prompt ingestion;
- whether all intended layers remain on the GPU;
- time to first output;
- prompt-processing rate; and
- generation rate.
Do not infer a 32K result from a 2K run. The key-value cache grows with context, and runner settings such as cache quantization can materially change memory use.
Use a Task Suite You Can Publish
Quality claims need retained evidence too. A useful small suite might include:
- one function-completion task with tests;
- one structured JSON task checked against a schema;
- one fact question whose answer exists in supplied context;
- one long-context retrieval question;
- one summarization task scored against required points; and
- one refusal task where the provided material lacks an answer.
Run identical prompts with deterministic settings where the runner supports them. Save raw outputs, expected answers, and the scoring script. Human preferences can be recorded separately, preferably with model names hidden during review.
Avoid claims such as “indistinguishable from unquantized,” “best at reasoning,” or “the only model that can solve” unless a published protocol and results directly support that scope. A handful of favorite prompts cannot establish a model-wide ranking.
Record the Complete Results
Keep every measurement connected to the conditions that produced it:
| Model artifact | Quant | Context | Prompt tok/s | Generation tok/s | Peak VRAM | Task score | Runs |
|---|---|---|---|---|---|---|---|
| Record exact tag or hash | Record exact type | Measured value | Mean ± spread | Mean ± spread | Measured value | Defined score | At least 5 |
Include failures and CPU offload rather than discarding inconvenient runs. A fast model that answers the test incorrectly should not win because only throughput was reported.
How to Make the Choice
Use the smallest quantization and context combination that passes the actual quality threshold while leaving operational headroom. A general pattern—not a guaranteed result—is that smaller quantizations use less storage and memory, while larger quantizations preserve more numerical precision. The size, speed, and quality differences must still be measured for the exact model conversion and backend.
For a desktop that also drives displays, leave more VRAM unused than a headless server would. For retrieval-augmented generation, budget for the context and any companion embedding model. For code completion, prioritize first-token latency and fill-in-the-middle behavior rather than chat benchmark scores.
The RTX 4070 is a capable 12 GB inference card. A trustworthy comparison starts where a spec-sheet comparison ends: with pinned artifacts, saved raw output, repeated measurements, and task-specific scoring.