From the engine room
DE — Deutsche VersionInsights
Technical field reports from our work with local AI — hardware, models, benchmarks and what actually holds up in day-to-day use. Real measurements instead of marketing.
AI Glossary
AI terms explained
From A3B to zero point: a reference for local language models, quantization, hardware and speculative decoding.
Open glossary →#8 · Laguna S 2.1 INT4 on 8×3090: A Fast Specialist with Sharp Edges
Most models we have run on our 8x3090 rig have fallen into one of two camps: HY3 was smart and painfully slow, while the early MiniMax quants were fast but not always technically up to the job. We explored how those differences play out under sustained agent workloads in our comparison of...
Read →#7 · 35 Billion Parameters, Three Billion Active: Why Qwen3.6 Runs So Fast Locally
A 35-billion-parameter model, around 150 tokens per second in single-stream decode, and only two RTX 3090s to run it: that combination was not what we expected. On paper, Qwen3.6-35B-A3B looked like one more large model with a serious appetite for VRAM. In practice, it became the model we kept...
Read →#6 · Qwen3.5 122B on Eight RTX 3090s: The First Large Model That Felt Like Infrastructure
By the time we had worked through reasoning loops, unexpected language switches, erratic rankings, and hardware faults masquerading as model problems, we had settled on a fairly sober standard: a model that barely squeezes into 192 GB of VRAM is not yet a usable server. It needs enough headroom...
Read →#5 · HY3 on 8x RTX 3090: A Sharp Mind in a Slow Body
After our extended experiments with MiniMax M2.7, Tencent HY3 was an obvious next candidate. The 295B-A21B MoE had a strong reputation for coding and reasoning, a promising mixed GGUF quant was available locally, and our eight GPUs had enough combined VRAM to make the idea seem at least somewhat...
Read →#4 · Beyond 60,000 Tokens: MiniMax M2.7 Under Sustained Agent Load
For production use, we needed to know how quickly, coherently, and consistently the model could still perform with 60,000, 80,000, or 100,000 genuinely active tokens. MiniMax M2.7 showed us that high short-context throughput and impressive individual answers can coexist with unstable behavior over long sessions.
Read →#3 · Why Five Bits Aren’t Five Bits: A MiniMax M2.7 Quantization Odyssey
A 4-bit quant was fast and useless, and two 5-bit quants were worlds apart. Three attempts, one breakthrough, and a plot twist in the proxy: why nominal bit width tells you nothing about how smart an MoE quant is — and which hybrid finally delivered speed and quality at once.
Read →#2 · Thrown Under the Bus: Power, Heat, and Other Lessons from Our 8×3090 Rig
450 W at idle, spikes up to 3.1 kW, GDDR6X in summer heat, and a GPU that mechanical pressure shoved clean off the PCIe bus: the operational lessons from our 8×3090 rig — including the diagnoses nvidia-smi won't show you.
Read →#1 · Eight RTX 3090s, Seven Slots, and an Awful Lot of Cables: How We Built Our Local LLM Rig
192 GB of VRAM from eight used consumer GPUs: bifurcation, proper PCIe 4.0 risers, four power supplies, and a wooden frame — everything that actually sits between owning eight GPUs and running them stably.
Read →