Llama 3.1 8B Quant Shootout: 4‑bit to 6‑bit on RTX 3060
On an RTX 3060, Q5_K_M runs Llama 3.1 8B nearly as fast as 4-bit with better output. See the numbers for Q4, Q5, Q6.
Hands-on writing about local AI models, the software built on them, and the machines you run it all on.
On an RTX 3060, Q5_K_M runs Llama 3.1 8B nearly as fast as 4-bit with better output. See the numbers for Q4, Q5, Q6.
Build a local document Q&A system that doesn't lie. Exact chunking strategies, retrieval setup, and prompt constraints to stop LLM fabrications cold.
Compare Faster-Whisper vs whisper.cpp for local speech-to-text. Dictation vs subtitles: setup, speed, hardware trade-offs — all offline, private.
A sleeping laptop kills your model. Run Ollama on a cheap DigitalOcean droplet for always-on inference—$200 free credit with referral link.
Intel N100 vs Ryzen 5 5500U for local LLMs on a $300 budget: token speeds, memory limits, and which mini PC keeps the conversation flowing with Ollama.
Head-to-head local LLM benchmarks: Llama 3.1 8B, Mistral 7B, and Phi-3 Medium on RTX 4070 12GB. Tokens/sec, VRAM, and quantization trade-offs—see which mod
Shingled magnetic recording drives can silently cripple ZFS pools with timeouts, faults, and multi-day resilvers. Here's how to spot and avoid them.
Set up Linkwarden to self-host bookmarks with automatic page archiving. A step-by-step guide to a permanent, private archive you control on your own hardwa
Get your Proxmox host to shut down cleanly before the UPS battery dies. NUT config, USB passthrough, and a shutdown script that actually fires.
Ditch tracking and ads by self-hosting SearXNG, a metasearch engine that aggregates results privately. Step-by-step guide for your homelab.
A $20 ConnectX-2 promises 10GbE—along with kernel panics, counterfeit EEPROMs, and DACs that won't link. Here's how to make cheap 10GbE reliable.
Ditch risky port forwarding. Use Tailscale's WireGuard mesh VPN on Proxmox to access your homelab securely from anywhere—no open ports required.