From the engine room

DE — Deutsche Version

Insights

Technical field reports from our work with local AI — hardware, models, benchmarks and what actually holds up in day-to-day use. Real measurements instead of marketing.

AI Glossary

AI terms explained

From A3B to zero point: a reference for local language models, quantization, hardware and speculative decoding.

Open glossary →

#18 · Ten Cards, Over 100 tok/s: The Team Now Works Entirely on the Large Qwen

Five or six developers here now spend the whole day working against Qwen 3.8 Flash Next, spread across all ten RTX 3090s — the small 27B is switched off for the time being. The service delivers 105 tok/s in the short run and 112 tok/s at 100,000 tokens of context, in a context window of 524,288 tokens. Along the way we had to take back one of our own findings from the previous article: that more cards are slower only held as long as one card sat in the wrong slot. The article also tells why the reported pool size of 3.3 million tokens is really just under 2 million and why the even layer split is not the best one.

Read →

#17 · Qwen 3.8 Flash Next 176B: Our Workhorse’s Well-Read Brother

There is a saying about doctors: surgeons know nothing and do everything, internists know everything and do nothing. Our workhorse, the small Qwen 3.8 27B, plays the surgeon — strong at coding, thin on knowledge. Qwen3.8-Flash-Next is considerably larger and should close part of that gap while matching it on tool calling; that, though, we have not measured yet. What we did measure: 83.9 tok/s at 100,000 tokens of context, a window of 524,288 tokens, on six of ten cards. The road there from 28.1 tok/s ran through a different engine, speculative decoding and the finding that more cards are slower.

Read →

#16 · Scaling Up: The Monster Rig With Room for Up to 19 GPUs

Our inference rig has grown from eight RTX 3090 to ten, and all ten are running. What makes that possible is a dual-socket board with 19 connectors for graphics cards, up to ten of which can hang off a single CPU. The road there ran through three cards we thought were dead, two series of crashes, and a suspicion we held ourselves: that the new platform was slower. It is not — after the memory expansion and socket-local cabling, both clusters are back over 120 tok/s, the level of the old rig. The 11.7 percent deficit we measured came from a broken interim state, and not one single card was defective.

Read →

#15 · One Model for the Team: Qwen3.8-27B on Two Four-GPU Clusters

Since late August, several of our developers have been pointing their coding agents at the same Qwen3.8-27B endpoint: two clusters of four RTX 3090 each, with a proxy of our own handing out the requests. The hard part was not throughput but what happens when several agents turn up at the same moment with a 120,000-token prompt apiece. This is the story of the collapse that forced the split, the routing rules that came out of it, and the single flag that cut idle power by 775 W.

Read →

#14 · A Full-Fledged Office Assistant on a Single GPU: Qwen 3.8 27B

Artificial Analysis ranks Qwen 3.8 27B first in its size class on the Intelligence Index, with coding and tool calling among its strong suits — the two disciplines office automation is built on. The model runs here on a single RTX 3090 with 24 GB and still decodes at roughly 58 tok/s with 81,000 tokens of context. How we got it there, and where our front office already works with a model like this every day.

Read →

#13 · DeepSeek V4 Flash on TensorSharp: The Server Path Is Ready, the OpenCode Configuration Isn’t

At the end of Post 10, we made a deliberately unglamorous promise:

Read →

#12 · The Six Strings Laguna Cannot Write

Some bugs swallow an entire workday while every system involved reports success and apparently nothing happens.

Read →

#11 · DeepSeek-V4-Flash on 8x RTX 3090: One Million Tokens, but No Tensor Parallelism

A model can run on eight GPUs. That statement, however, can mean two very different things.

Read →

#10 · The LLM Rig Built from eBay Parts: 2 x RTX 3060, 120 Euro for Board and CPU, 77 tok/s

Our large local LLM rig has eight RTX 3090s, an EPYC, and a power supply that could probably handle some light metalworking on the side. For that build, throwing hardware at the problem was part of the brief. This time, we wanted to find the opposite limit: how cheap can a local LLM machine get...

Read →

#9 · Three Models, Eight GPUs: How We Split 192 GB of VRAM

Local inference often revolves around one question: how do we get this one large model running? One model, all eight GPUs, TP=8, done. It makes a good story for the first photo of the server in action.

Read →