Four people have run this test and no two of them held the same things still
The comparison exists. That is worth saying plainly, because the shape of the answer is not "nobody knows", it is "four answers, four setups, and the cleanest one was published by the people who build DFlash2".
| Who ran it | Machine | What was held fixed | Result |
|---|---|---|---|
| DFlash2's authors, on the model card | 1x H200, SGLang | Target model, precision, engine, sampler, and the draft budget matched at seven tokens across every arm | GSM8K at concurrency 1: plain decode 68.9 tok/s, MTP 178.5, DFlash 2 236.1 |
| Stepan Mazurov, 2026-09-07 | 2x GB10 at TP2, GLM 5.3 Flash NVFP4 | Context at 262,144, KV at FP8, chunked prefill, concurrency | DFlash2 46.06 tok/s against native MTP 38.73, aggregate at 128k cached |
| Ama5u, NVIDIA forums, 2026-08-22 | 1x DGX Spark, Qwen3.8 27B NVFP4 | Quantization and context, across 69 benchmark scenarios | DFlash2 32.6 on code against MTP 19.8, about 2.5x everywhere |
| @analogalok, published here in August | 1x RTX 4090, Qwen3.8 27B Q4_K_XL | Model file, engine, card, prompt depths from 26k to 212k | DFlash2 leads decode by 10 to 15% at every depth. Native MTP wins prefill at every depth, by up to 40% |
Three of the four put DFlash2 ahead on decode. The one run by DFlash2's own team puts it furthest ahead, at 3.43x against MTP's 2.59x, on a datacentre H200 nobody reading this owns.
The honest summary is the fourth column. Mazurov left each method at its own default draft budget, five for DFlash2 against three for MTP, so part of his gap is budget. Ama5u ran MTP on vLLM and DFlash2 on SGLang, so his engine moves with his drafter. Ours held the machine and the engine still and swapped only the drafter, and it found the two methods winning different halves of the job.
Nobody has reproduced the H200 numbers on a card you can buy, and nobody has run the clean swap twice.
What the two things actually are
MTP, multi token prediction, is a head trained into the model and shipped inside the checkpoint. Qwen3.8-Flash-Next carries one in its config. So does Qwen3.8 27B. GLM 5.3 Flash ships a bundled NEXTN head that SGLang drives through its EAGLE path. If the lab trained one, you switch it on. If it did not, no download adds one to your checkpoint.
DFlash2 is the other arrangement: a separate 2B draft model that runs inside the server and proposes tokens for the big model to verify. It works on any target its authors have trained a drafter for, and it costs memory the head does not. On a 24 GB card we priced that draft at 1.14 GB as the Q4_K_M GGUF, 3.85 GB at full precision, plus draft state that grows with the draft depth.
Not every model uses either. DeepSeek V4.1 Flash ships DSpark, a third approach, which is why it appears in the measurements below without being in the title.
Every number you find sits on a different model
Go looking for DFlash2 speeds and you will land on Qwen3.8 27B. Go looking for MTP and you will land on Qwen3.8-Flash-Next. Of the measured runs we track, eighteen of the twenty one DFlash2 results are on the 27B, and sixteen of the nineteen MTP results are on Flash-Next.
Nothing about the models forces that, because both of them ship an MTP head. DFlash2 launched with a drafter trained for the 27B, so that is the model its users arrived on. Flash-Next came with its own head ready to switch on, so its users never needed a drafter.
The cost lands on you when you try to settle the question by reading. The two piles of numbers are not comparable, and stitching one to the other produces a result that looks decisive and means nothing. On a single DGX Spark, GLM 5.3 Flash decodes at 17.29 tok/s with MTP and at 41 tok/s with DFlash2. Both are real. Together they tell you nothing, because the quantizations differ, and the MTP one is not even a flat cut: that tester kept attention, the shared expert, embeddings and lm_head at source precision while trellis quantizing the 288 routed experts. The engines differ too, ExLlamaV3 against llama.cpp. So does the context, 64,000 against 11,000. So does the person. A 2 bit model is smaller, and on a memory bound machine smaller is faster whatever else you change.
The before and after pairs, and the one that answers the question
Four testers have published a speed for the same machine with DFlash2 off and on:
| Tester | Machine | Baseline | With DFlash2 | Ratio |
|---|---|---|---|---|
| @analogalok | RTX 4090 | 60, on native MTP | 87 | 1.45x |
| @ViC305 | DGX Spark | 12.37, plain | 24.51 | 1.98x |
| @fahdmirza | A100 | 28.9, plain | 59.1 | 2.05x |
| @danpacary | DGX Spark | 17.2, plain | 41 | 2.38x |
Read the baseline column before the ratio, because they are not measuring the same thing. @analogalok's 1.45x is DFlash2 against a native MTP head that was already running. The other three are DFlash2 against plain decode. His is the smallest number in the table and it is the only one that answers the question this article is about.
Most published speedups move more than the drafter
Testers optimise machines. They are not running experiments with controls, and the numbers that circulate are bundles.
@Tech2Wild publishes enough detail to see it. On one DGX Spark with Qwen3.8-Flash-Next in NVFP4:
| Configuration | Decode |
|---|---|
| Eager, no MTP, no CUDA graphs | 15.4 tok/s |
| MTP4, piecewise CUDA graphs, table in memory | 32.5 tok/s |
| MTP3, staged gather, decode graphs, table on disk | 43.9 tok/s |
The jump from 15.4 to 32.5 gets quoted as what MTP is worth. It is what MTP plus CUDA graphs is worth, because the floor has both switched off.
His two box figures make the point from the other side. At tensor parallel 2, MTP3 with decode graphs and the table in memory gives 53.7 tok/s, while MTP4 with piecewise graphs and the table read off disk gives 35.8. Same model, same pair of machines, same method, and a 50% spread that belongs to everything except the method.
We made this mistake ourselves, on this exact tester. We had one of his runs written down as "20 tok/s with MTP off, 54 with built-in MTP on. About +64%." Twenty to fifty four is not 64%, it is 170%. The 64 was his acceptance rate, sitting a few lines further down the same repository page. An acceptance rate had been copied across as a speed gain, which is what the next section is about.
Acceptance length is the number that travels
Speculative decoding proposes several tokens at once and then verifies them. Two figures describe it, they get confused constantly, and only one is a count.
Acceptance rate is the share of proposed tokens the model keeps, as a percentage. Acceptance length is how many tokens get committed per verification step. It starts at 1, because a step that accepts nothing still commits the token the model would have produced anyway.
@Tech2Wild measured it twice, on two models running two methods, three weeks apart: MTP4 on Qwen3.8-Flash-Next at 3.56 tokens per step, and DSpark on DeepSeek V4.1 Flash at 3.57, ranging 1.88 to 5.92 across 35 windows. DFlash2's own card reports 3.74 to 5.46 depending on the workload, and 5.46 against MTP's 5.02 on the benchmark where it wins hardest.
Then a measurement that shows why acceptance still does not decide it. Running DeepSeek V4.1 Flash on a single RTX 5090, streaming the weights off NVMe, JigSawPT measured the DSpark head at a median minus 4% on cold content and level once cached, across prompts whose acceptance ran 51 to 79%. Individual prompts ran from minus 8% to plus 15%. On verbatim repetition, where acceptance reached 97%, it gave plus 12 to plus 15%.
A drafter accepted three times in four, paying nothing. The reason is worth carrying: a verification step has to fetch the experts every token in its block might need, and on a machine reading weights off a disk that fetch is most of the cost. Drafting saves forward passes. It does not save memory traffic, and on the machines people actually buy, memory traffic is the budget.
Measure it on your own machine
Ten minutes settles your setup better than any table here, including ours.
- Fix everything else first. Same file, same quantization, same context, same batch, same CUDA graph setting. Change graphs and the drafter together and you learn what @Tech2Wild's 15.4 row teaches, which is nothing about drafters.
- Measure the floor. Drafter off, three runs, take the median.
- Switch only the drafter on. Same three runs, same median.
- Read the acceptance length, not the rate. llama.cpp, vLLM and SGLang all report it.
- Test the content you actually use. Prose, code and structured output land in different places, and every tester who has split them reports the same order: structured and code accept best, prose worst.
Our speed checker carries the measured runs behind this page, and the model catalogue will tell you whether the 1.14 GB a quantized DFlash2 draft wants is memory you have spare after the weights and the cache.
FAQ
Is DFlash2 faster than MTP?
On decode, in three of the four published comparisons, yes. The margin depends entirely on who measured. DFlash2's own authors report 3.43x against MTP's 2.59x on an H200 with the draft budget matched across arms. An independent test on two GB10 nodes puts DFlash2 at 46.06 against 38.73, with the budgets left at each method's defaults. Our own RTX 4090 test gives DFlash2 a 10 to 15% decode lead and gives native MTP the prefill lead at every depth, by up to 40% on a 26k prompt. Nobody has reproduced the datacentre result on consumer hardware.
Can I use both at once?
No, and @analogalok's own two commands are the clearest way to see why. His MTP run is --spec-type draft-mtp --spec-draft-n-max 4. His DFlash2 run is --spec-type draft-dflash --spec-draft-n-max 4. One flag, one value, one drafter in the loop. vLLM takes a single method in its speculative config and SGLang a single --speculative-algorithm, the same shape. So where you see a result described as one stacked on the other, check what the baseline was: his 60 tok/s is that 4090 running native MTP, and his 87 is the same card running DFlash2 in that slot instead.
Which one do I get for my model?
Check whether your checkpoint ships a head. Qwen3.8-Flash-Next and Qwen3.8 27B both do, GLM 5.3 Flash ships a bundled NEXTN head, and DeepSeek V4.1 Flash ships DSpark rather than an MTP head. Where a head exists it costs no extra weights, so it is the obvious first thing to try. An external drafter is what you add when you want to beat the head, not when you have no head at all.
What acceptance length should I expect?
Around three and a half tokens per step on a well configured setup. @Tech2Wild landed there twice, on MTP4 with Qwen3.8-Flash-Next and on DSpark with DeepSeek V4.1 Flash, and DFlash2's card reports 3.74 to 5.46 across workloads. Below two, your draft is not helping much. Above five, you are probably measuring repetitive content rather than the work you do.
Does the draft budget matter?
Yes, and it is the one parameter with evidence behind it. In a controlled fit against our own corpus, 42 scored runs from 15 testers, the stated draft budget carried real signal about how far a run beat its predicted speed, and it beat the choice of inference engine outright. Refitted with a whole tester held out it fell apart and started running hot, which on a tool that tells people what to expect is the worse direction to be wrong in. So it is measured and it is not shipping. Budget matters; how much is still unsettled.
Why does my drafter make things slower?
Most often because your weights are not all in memory. When a model streams from system RAM or an SSD, each verification step reaches for whatever its whole block of drafted tokens might need, and that fetch costs more than the forward passes you saved. Measured, not theoretical: a draft head accepted 51 to 79% of the time came out a median 4% slower on a single RTX 5090 running a 502 GB checkpoint off NVMe.