From the engine room
DE — Deutsche VersionInsights
Technical field reports from our work with local AI — hardware, models, benchmarks and what actually holds up in day-to-day use. Real measurements instead of marketing.
Why Five Bits Aren’t Five Bits: A MiniMax M2.7 Quantization Odyssey
A 4-bit quant was fast and useless, and two 5-bit quants were worlds apart. Three attempts, one breakthrough, and a plot twist in the proxy: why nominal bit width tells you nothing about how smart an MoE quant is — and which hybrid finally delivered speed and quality at once.
Read →Thrown Under the Bus: Power, Heat, and Other Lessons from Our 8×3090 Rig
450 W at idle, spikes up to 3.1 kW, GDDR6X in summer heat, and a GPU that mechanical pressure shoved clean off the PCIe bus: the operational lessons from our 8×3090 rig — including the diagnoses nvidia-smi won't show you.
Read →Eight RTX 3090s, Seven Slots, and an Awful Lot of Cables: How We Built Our Local LLM Rig
192 GB of VRAM from eight used consumer GPUs: bifurcation, proper PCIe 4.0 risers, four power supplies, and a wooden frame — everything that actually sits between owning eight GPUs and running them stably.
Read →