← Insights

DE — Deutsche Version

#18 · Ten Cards, Over 100 tok/s: The Team Now Works Entirely on the Large Qwen

· LYTH-REST-Publisher

For a few weeks now, five or six of our developers have spent the whole day working against a single model: Qwen 3.8 Flash Next, spread across all ten RTX 3090s in our rig. The small 27B, our workhorse for day-to-day work until then, is switched off entirely for the time being.

The impression after those weeks is genuinely good. The large model is noticeably smarter than the 27B, and that shows up most clearly in code review. The honest assessment belongs with that, though: it does not reach the big frontier models. For our coding pipelines it is enough.

Running it turns out to be surprisingly undramatic. We operate a context window of 524,288 tokens — YaRN makes that possible, a method that stretches the positional encoding and hands the model a larger window than it was trained for. The inference engine vLLM copes very well with several developers working in parallel, and even at large context the model barely slows down: over 100 tok/s at 100,000 tokens of context is noticeably pleasant in daily use. The shared buffer that all running sessions draw from holds more than one and a half million tokens.

Two things have changed since the state described in #17. First, we now run a mix of two kinds of parallelism. Two cards at a time share the model’s individual weight matrices; that is called tensor parallelism. Five such pairs of cards in sequence then share the 48 layers, and that is pipeline parallelism. Second, we have taken back one of the findings from the previous article: the claim that more cards make the model slower. It held as long as one card sat in the wrong slot.

What runs now

The final state: all ten cards, two cards per tensor pair and five pipeline stages, the layers split 11/10/10/10/7, plus a self-patched container image and speculative decoding with four draft tokens. Speculative decoding already came up in #17: the model ships with a small draft head that proposes several tokens at once, and the large model checks that draft in a single pass.

Metric State in #17 (6 cards) Today (10 cards)
Split 2 cards per matrix, 3 stages 2 cards per matrix, 5 stages
Short run about 71 tok/s 105.17 tok/s
At 100,000 tokens of context 83.9 tok/s 112.59 tok/s
Reading in a prompt of 100,000 tokens 4,992 tok/s 8,570 tok/s
Context pool, as reported by the server 793,921 tokens 3,334,471 tokens
Context pool, measured not measured about 1,950,000 tokens

By context pool we mean the KV cache here: the buffer in which the service holds the context already read in for all sessions running at the same time. It is the real capacity limit in multi-user operation, not the context window of a single session. Reading the prompt in — the prefill from here on — is a work step of its own with its own speed; it runs before the first output token.

Against the state of the previous article that is 4.2× the reported pool at 34 percent more throughput. By the capacity we actually measured, it is enough for about 15 sessions of 130,000 tokens each, or just under four sessions of 500,000. Why the reported and the measured figure sit so far apart is a story of its own, further down.

“More cards make it slower” — that no longer holds

#17 records the finding that extra cards push throughput down: eight cards slower than four, ten slower than eight. All ten run now, and they deliver more than any earlier setup. The explanation comes in two parts, and the second one is the more uncomfortable.

The first part is the kind of parallelism. The old statement applied to tensor parallelism, that is, to the case where all participating cards share every single weight matrix and have to reconcile their intermediate results after every part-computation. With five pipeline stages that problem spreads out: a slow pair of cards now carries only a fifth of the layers instead of a third, and the reconciliations inside a pair are short.

The second part: One card sat in the wrong slot. Our board has two CPU sockets. Nine of the ten cards hung off the second CPU, a single one off the first; moving it had been on our list since the rebuild, but the price had never been measured. Of all the cards, that was the one that formed the first tensor pair with its neighbor in the six-card setup — and tensor parallelism synchronizes after every part-matrix. Every one of those reconciliations therefore ran across the socket boundary, over the link between the two CPUs. The system itself rates that route with a distance of 32, against 10 within a single socket.

The counter-test was cheap: same setup, same configuration, just a different set of cards whose six members all hang off the same CPU. The short run went from 70.93 to 99.21 tok/s, so up 40 percent. At 100,000 tokens of context it went from 83.86 to 98.97 tok/s, plus 18 percent. Every measurement in the previous article was therefore taken against a handicapped setup.

Setup Cards Short run 100k Prefill at 100k
2 per matrix, 3 stages, crossing the socket 6 70.93 83.86 4,992
2 per matrix, 3 stages, all on one CPU 6 99.21 98.97 6,597
2 per matrix, 5 stages, crossing the socket 10 98.99 103.63 8,629
8 cards per matrix, 1 stage, all on one CPU 8 97.38 89.90 1,558

The last row is why we don’t run pure tensor parallelism across eight cards. Generation looks usable; the prefill collapses by a factor of 5.5. Without NVLink every reconciliation goes over PCIe and main memory, and reading in a long prompt produces a great many of them. On top of that comes a quirk of the model that already decided the number of cards in #17: for the context buffer it has only two heads, so two parallel part-computations across which that memory could be spread. Two heads cannot be divided across eight cards, so the memory gets multiplied instead of divided, and the pool falls accordingly.

Then we actually moved the card, and honesty calls for the addendum: with ten cards and the same layer split, socket locality was worth only another six percent at 100,000 tokens, from 103.63 to 109.92 tok/s. In the short run the median even dropped from 98.99 to 91.02 — with a mean of 105.77 in the reference run, the short-run values scatter too much to prove differences of that size. The 40 percent from the six-card test certainly do not carry over. With five stages the handicapped pair carried only a good nine of 48 layers instead of sixteen, so it barely mattered anymore. The find was right all the same; by the time we could act on it, it had already been half defused.

The layer split is the bigger lever

With five pipeline stages you are free to choose how many of the 48 layers each stage carries. We measured four splits, all of them after the card was moved, all with ten cards and an otherwise identical configuration.

Layers per stage Context pool, reported Short run 100k Prefill
9/10/10/10/9 (the engine’s default) 2,149,580 91.02 109.92 8,226
10/10/10/10/8 3,130,748 98.37 108.38 8,540
11/11/11/11/4 3,334,471 89.94 102.24 8,117
11/10/10/10/7 3,334,471 105.17 112.59 8,570

The even split, the one the engine proposes itself, yields the smallest pool and the lowest speed in the short run. The largest pool and the highest throughput at the same time come from a slightly lopsided split with eleven layers on the first stage and seven on the last.

The mechanics behind that are unintuitive but simple. The service manages context in blocks of fixed size, and all five stages have to hold the same number of those blocks; each one stores only the context of its own layers in them, though. The card with the least free memory therefore sets the block count for all ten, and free memory above that bottleneck is structurally out of reach. The last stage additionally carries the model’s output layer and is almost always the tightest for that reason — it needs three to four layers fewer than the others so that it doesn’t hold all the rest back. In the best of the cases we measured, 28.3 of 240 GiB of graphics memory go unused, about 16 GiB of that as working reserve that must not be touched.

That reserve is exactly where the optimization ends. The 11/11/11/11/4 split squeezed the buffer the cards need for their coordination so far down that startup aborted out of memory — or, worse, only did so after hours under load. Below roughly 1.6 GiB of free graphics memory on the tightest card it gets unsafe. That is not a theoretical limit: one of the earlier splits ran through a whole night first and then aborted out of memory on one card.

The pool is smaller than the server reports

At startup the service reports a context pool of 3,334,471 tokens. That number is no good for planning, and it is the finding that surprised us most. A control measurement with a single long prompt shows it:

prompt of 362,948 tokens
share of the reported pool used:  18.61%
arithmetically expected:          10.88%

362,948 / 0.1861  =  about 1,950,000 tokens of actual capacity
what gets reported   3,334,471  —  a markup of about 71 percent

The cause lies in the model’s architecture. 36 of its 48 layers are recurrent: they carry forward a state that the next token needs. That state is checkpointed every 1,616 tokens, so the service can resume work on a session without recomputing everything. Those checkpoints occupy the same memory as the context buffer — for a prompt of 363,000 tokens that comes to about 225 of them per session. They do not show up in the reported pool size, because that counts only the buffer of the remaining twelve layers, the ones computing classical attention.

reported measured
Context pool 3,334,471 tokens about 1,950,000 tokens
Sessions of 130,000 tokens each 25 about 15
Sessions of 500,000 tokens each 6 just under 4

For capacity planning, the right-hand column is the one that applies. In practice that means: take the reported number at face value, and you run into eviction at around 60 percent displayed occupancy, because memory is actually full by then and the service starts throwing context out. The caveat belongs with it: this is one measurement at one context length. That the markup per token is constant is plausible and unproven.

The collapse into a single word

Alongside these measurements ran a bug that hit operations harder than any shortage of capacity. Sessions tipped reproducibly into the endless repetition of a single token. It was always the same token, the English word “duct”. After that the model filled its entire output budget with that word and ended at the token limit.

We have it on record in three cases from two exported sessions: collapse after 0 tokens at 14,623 tokens of context, after 303 tokens at 75,593, and after 24 tokens at 108,256. The case with zero generated tokens is the most revealing — there was nothing there that could be read as the model running off into a long chain of thought. This is not overthinking; it is a destroyed state.

We measured the condition for reproducing it, and it matches production exactly:

  • On its own, at roughly 100,000 tokens of context: 24 runs, not a single collapse.
  • Four at the same time, same context: eight runs, one collapse.

The cause was a bug in the engine, fixed in a contribution dated September 5; our container image was four days older. According to the bug report it takes at least two prompts being read in at the same time plus active speculative decoding — both are the normal case for us as soon as more than one developer is working. For operations there is a usable diagnostic: on affected requests the acceptance rate of the speculation, the share of draft tokens that get taken, falls to exactly zero.

That clears up an open item from our operating statistics. 36 of 382 requests had run into the token limit at the time instead of ending normally, and we had put that down to an output budget set too tightly on the client side. In all likelihood it was these collapses: on an instance with purely serial traffic, 139 of 139 requests subsequently ended normally, not one of them at the limit.

The patched image

The fix existed; it just wasn’t usable. The new image wouldn’t start for us at all — a known bug with hybrid models under pipeline parallelism: a stage can end up owning no layer at all for one group of cache layers, and that empty group crashes memory allocation at startup. The solution was two lines from a contribution that is still open, which we copied into the container and committed as an image layer of our own.

There is a trap in that which takes a while to think of: when you commit, the tool takes over the start command of the helper container you patched in. Unless you set it explicitly, what stays there is the sleep command you used to keep the helper container open. The finished image then doesn’t start the server, exits immediately and writes not a single log line. Back in #17 the most instructive startup failure was already one without any error message at all; here, again, the most informative log line was the one that was missing.

The patched image brings one more change that you have to budget for in main memory. The model’s large lookup table, which on our machine sits in the computer’s main memory rather than on the cards, no longer needs a helper process of its own there. That makes one error path unnecessary, but costs about 131 GiB of shared main memory instead of the 93 GiB of private memory before.

Speed against space

Speculative decoding is the single largest item in the context pool, and in a way you cannot tell from the switch. Measured at identically allocated graphics memory:

Context pool Short run 100k
four draft tokens 793,921 70.93 83.86
no speculation 1,560,671 51.12 50.28

The speculation halves the pool, because with four draft tokens the table of context blocks needs five columns per session instead of one — the service reserves the slots for the draft whether they get used or not. That is a genuine trade-off and not a bug: speed against space, in this case a factor of two against 67 percent more throughput at long context. We went with speed.

The pipeline itself scales only weakly with the number of concurrent requests. In one measurement series without speculation on six cards, three concurrent requests delivered 104.4 tok/s together, six requests 128.8 and twelve requests 223.5. The reason lies in how the pipeline works: with three stages the engine’s scheduler distributes three requests across three part-steps with one sequence each. The pipeline fills up, but nothing gets computed as a batch.

In daily use we notice little of that, and that is down to the speculation — whose value we underestimated for a long time. We had measured an acceptance rate of 22.9 percent earlier, though on synthetic test tasks. Under real workload it sits at 64.6 percent, so the model takes 3.58 tokens per verification cycle instead of one. That batches on the other axis what the pipeline does not batch: roughly 71 verification cycles per second times 3.58 accepted tokens works out to about 254 tok/s; in live operation we have observed 132 to 234 tok/s across all sessions.

Context does not slow generation down

One measurement habit led us astray twice along the way: tokens divided by total duration falls as context grows, because the prefill flows into that rate. Measured separately, pure generation is nearly constant across all context lengths — 60.1 tok/s at 71 tokens of context, 63.0 at 55,051 and 58.0 at 140,851. The absolute values come from a different setup than the final state above; all that matters here is that they do not fall off from 71 tokens of context to over 140,000. What grows with context is the prefill. Mix the two into one number and you measure an effect that isn’t there.

The cap on thinking, this time per request

The model likes to think for a long time before it answers, and without a limit it spends its whole output budget doing so. In our series of programming tasks that is literal: no code, only thinking text, abort at the limit. With a cap on the thinking part, code comes out — at 512 tokens after 13.3 seconds, at 2,000 tokens after 43.3 seconds. Without a cap the same task ran 65.8 seconds and delivered nothing executable.

We had the same problem with the 27B, and there it could only be solved for the whole service — the addendum to #15 describes that. With the large model the situation is better. For one thing there is a thinking budget per request that takes effect without restarting the service, so the clients can set it themselves. For another, the setting for thinking effort actually reaches the chat template here, the template from which the service assembles the final prompt. We showed that with a counter-test: invalid levels are rejected with an error, and the three valid levels produce measurably different prompt lengths, 53, 41 and 11 tokens. So the setting arrives instead of being silently ignored.

What the optimal cap for this model is, we don’t know. With the 27B the optimum sat at 4,000 to 5,000 tokens; for the big brother we have only the two single measurements at 512 and 2,000 so far.

An outside claim, measured

On our side the model’s weights sit in a compressed form, four bits instead of sixteen per value. Builds like that are called quantizations, and different providers produce them differently. There was a claim about them: that the model’s wordiness is an artifact of the quantization procedure — so a different quantized build would think more briefly. That was testable: two quantized builds, same engine, same startup arguments.

our build the other one
Short run 70.93 69.24
100k 83.86 66.83
Context pool 793,921 567,729
Without a cap on thinking no code no code

The switch costs 20 percent of speed at long context and 28 percent of context pool. It changes nothing about the thinking behavior: with the other build we ran the series without a thinking cap all the way through, zero of 38 tasks passed, every run consumed the full output budget. The claim is refuted, and we are staying with our own build.

What’s open

Two of the ten cards throttle under load, and not thermally — temperatures sit between 55 and 76 °C, and every card sits on the board with its full eight PCIe lanes of the fourth generation. The two run into their power limit of 250 watts, while other cards in the rig have higher limits set. And of course those two make up the last pipeline stage, and a pipeline is only as fast as its slowest stage. Whether raising those two limits achieves anything we have not measured.

Output quality at very long context stays open as well. The 524,288 tokens come about by stretching beyond the trained length; we measured throughput, and quality at 300,000 or 500,000 tokens we did not measure at all. Anyone working seriously in that range should check it against their own tasks.

Operations are covered so far by a small monitoring script that polls the interface every minute and restarts the service after five minutes of outage, reporting the cause as it does so. That safeguard does not survive a restart of the rig; we have deliberately not set an automatic restart policy for the container yet, because after the experience from #17 it would hide every startup failure inside a loop.

And one tally we had not expected in this form: of the outside recommendations we checked over these weeks, three contradicted our own measurements. A block size for reusing the beginnings of prompts already read in is described in a fork’s documentation as practically mandatory — our hit rate of 88.8 percent on exactly that reuse says something else. One memory management option is supposed to destroy the coordination between the cards; we run it and don’t crash. And a combination of extra flags that according to the source should bring 19 to 39 percent more throughput cost us a measured 23 percent — the same combination that already stood out negatively in #17.

With this model on this hardware, our own measurement is therefore more reliable than the outside figure. That holds even when it contradicts what we wrote ourselves in the last article.

Write a comment

Your e-mail address will not be published. A first comment is approved manually.