Does DFlash 2 Actually Double Qwen 3.8 27B?
Yes, against a plain baseline. One A100 test went from 28.9 to 59.1 tok/s, and one RTX 4090 test went from 60 tok/s on MTP to about 90 on DFlash 2. The word “again” matters. DFlash 2 is not a faster MTP. It is a second, independent speed lever, a separate 3.85 GB draft model that proposes whole blocks of tokens in one pass, and it costs VRAM rather than a flag. Two testers, two cards, every number attributed below, and all three runs are in our tokens per second tool next to measurements on other hardware. This is the honest version of a claim that started travelling before its caveats did.
What DFlash 2 Is

DFlash is “Block Diffusion for Flash Speculative Decoding”, from the z-lab/dflash project on GitHub, 5,715 stars. Speculative decoding works like a rough first draft: a small draft model proposes tokens, and the full target model checks them, keeping only what it would have written itself. That is why the output is identical by construction. A bad draft just wastes time; it never changes a single token.
Standard speculative decoding drafts one token at a time. DFlash 2 drafts a whole block in one pass, keeps the top candidates at every position, and a selector traces one coherent path through them. Two-tap dynamic convolutions keep the draft from decaying toward the end of the block. The target is still Qwen 3.8 27B. The draft is incoai/Qwen3.8-27B-DFlash2, mirrored at z-lab/Qwen3.8-27B-DFlash2, and it is a 3.85 GB model, not a config flag.
MTP vs DFlash 2: Two Different Kinds of Fast
Qwen 3.8 27B already has a built-in speed multiplier, multi-token prediction. Our speed page covers it. DFlash 2 is the second one, and they are not the same mechanism.
| MTP | DFlash 2 | |
|---|---|---|
| Where it lives | Trained into the model’s own weights | A separate draft model |
| What it adds | One extra prediction head | A block-diffusion drafter, 1.14 GB as Q4_K_M |
| How it speeds up | Predicts the next tokens from the model itself | Drafts blocks in parallel for the target to verify |
| Cost | A runtime flag | VRAM for the draft, plus draft-state memory |
| Status | Shipped in the model | Landing via open pull requests |
They do not stack. They swap. Both run through one llama.cpp flag, --spec-type, and it takes a single value: draft-mtp or draft-dflash. @analogalok’s 60 tok/s is his 4090 running native MTP at 130k context, and his 90 is the same card with DFlash 2 in that slot instead. Both of his reproduction commands are further down this page, and they differ by that one value.
Corrected 12 September 2026. This paragraph used to say the two stack, and read the 60 to 90 jump as MTP with DFlash 2 added on top of it. That was wrong. The deep context table below was always a swap, and so is this. For which one actually wins, and the four published comparisons that disagree about it, see DFlash2 vs MTP.
The Measured Numbers
Three independent tests exist so far, all single testers, all attributed. Treat them as three rows, not one average. All three are in the tokens per second tool, which also estimates any card and model pair without a measured run, and says plainly that it is an estimate.
| Tester | Hardware | Baseline | With DFlash 2 |
|---|---|---|---|
| @fahdmirza | NVIDIA A100 | 28.9 tok/s | 59.1 tok/s |
| @analogalok | RTX 4090, 24 GB | 60 tok/s on MTP | about 90 tok/s |
| @ViC305 | DGX Spark | 12.37 tok/s | 24.51 tok/s |
Fahd Mirza’s benchmark is in his video, and he states the output is provably identical. @analogalok’s number carries a detail worth reading twice: his headline says 90 tok/s, and his own table, shown below, says 83 to 87. The headline travels. The range is what a card owner should plan around. @ViC305’s row is the cleanest: a plain baseline, 1.98×, and he measured the cost the others skipped. Prefill dropped from 73.38 to 37.62 tok/s because the draft has to prefill the same prompt, at a 54.2 percent acceptance rate. Decode is the win, and for agent work decode is what you wait for.
What It Costs in VRAM
DFlash 2 is not free, but it is cheaper than the headline number. The draft model is 3.85 GB at full precision, and 1.14 GB as the Q4_K_M GGUF the 4090 test actually used. On top of that, draft depth is a dial between VRAM and speed.
@analogalok’s full matrix, on a single RTX 4090 with 24 GB, using Unsloth’s Qwen 3.8 27B UD-Q4_K_XL target and the DFlash 2 draft at Q4_K_M:
| Context | Prefill | Decode | VRAM |
|---|---|---|---|
| 30k | 1,725 tok/s | 87.05 tok/s | 22.2 GB |
| 80k | 1,789 tok/s | 84.20 tok/s | 23.3 GB |
| 110k | 1,767 tok/s | 83.35 tok/s | 23.96 GB |
The number that matters is the flag --spec-draft-n-max. At 7, the draft states eat too much VRAM on a 24 GB card. At 4, the overhead drops and throughput actually rises. He reports a 5.39 token acceptance rate at that setting. The physical rule from section 5c applies: speculative decoding only helps while the draft accepts enough tokens to pay for its own memory and compute, and on a card where memory is the constraint, a shallower draft wins.
That 110k row, 83.35 tok/s decode sitting at 23.96 GB, is the ceiling on a 24 GB card. Nothing more fits, and the fact that it holds at all is the point.
How to Run It Today
DFlash 2 is landing through open pull requests, not shipped releases. The README links vLLM PR #52816, SGLang PR #35371, and llama.cpp PR #27342, plus an oMLX fork at 0.6.2-dflash2.
On llama.cpp, the PR is patchable now:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git fetch origin pull/27342/head:pr-27342
git switch pr-27342
cmake -B build -DGGML_CUDA=ON && cmake --build build -j
Then serve it with @analogalok’s flags:
./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
-md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 4 \
-c 110000 -ngl 99 -ctv q4_0 -ctk q4_0
Two caveats, stated plainly. First, “already in vLLM, SGLang, and llama.cpp” means open pull requests, and llama.cpp’s DFlash path has open bug reports: corrupted predicted_ms on some Q4 and Metal requests, and an AMD APU regression. It works on the cards people have tested, and it is not a finished, merged feature. Second, we measured none of this. Every number above is a named tester’s, with their card and their flags.
The target model’s VRAM and quant tiers, for anyone sizing the base model first, are on the VRAM requirements page. The MTP baseline that DFlash 2 builds on top of is on the speed page. The flags and how they change what fits are on the llama.cpp flags guide, and the GPU checker tells you whether your card holds the 20.6 GB base model at all. If it does not, DFlash 2 does not rescue it: the draft needs the target resident first.
It Does Not Fall Apart at Deep Context, and Somebody Measured Where It Lands
Updated 22 August. The standing objection to speculative decoding is that the advantage evaporates once the prompt is long. @analogalok tested it on one RTX 4090, feeding Qwen 3.8 27B at Q4_K_XL prompts from 26,000 up to 212,000 tokens, comparing his own 2-bit Q2_K DFlash 2 drafter against native MTP.
| Prompt | DFlash 2 prefill | DFlash 2 decode | MTP prefill | MTP decode |
|---|---|---|---|---|
| 26k | 1,662 t/s | 75.66 t/s | 2,324 t/s | 68.09 t/s |
| 80k | 1,597 | 59.22 | 1,986 | 55.56 |
| 160k | 1,350 | 48.93 | 1,499 | 42.49 |
| 212k | 1,213 | 41.54 | 1,333 | 37.58 |
DFlash 2 holds a 10 to 15% decode lead at every depth, and at 212,000 tokens it still decodes at 41.5 t/s. The decay is real, but it is gradual and it applies to both.
Read the prefill column the other way, because it reverses. Native MTP is faster at reading the prompt at every single depth, by up to 40% at 26k. DFlash 2 wins generation and loses ingestion. Which one you want is the same question as the MTP section above: if you paste a repository and wait, MTP reads it sooner; if you sit watching tokens appear, DFlash 2 gets there faster. This is one of four published head to head runs, and they do not agree, because no two of them hold the same things still: we went through all four.
The VRAM Nobody Told You llama-server Was Holding
The same tests turned up something that changes what fits, not just how fast it runs. By default llama-server reserves memory for serving several requests at once. If you are one person at a terminal, that reservation is bought and never used.
--parallel 1 hands it back. Combined with a 2-bit drafter instead of a larger one, Alok reports these ceilings on a single 24 GB card, all measured with a 28k prompt on Ubuntu 22:
| KV cache | Context | Decode | Prefill | Peak VRAM |
|---|---|---|---|---|
| Q4 | 250,000 tokens | 73.66 t/s | 1,608 t/s | 23.8 GB |
| Q8 | 150,000 | 75.01 t/s | 1,667 t/s | 23.9 GB |
| FP16 | 90,000 | 80.58 t/s | 1,699 t/s | 23.92 GB |
Our calculator does not model that reservation, and no calculator we know of does. It sizes weights, cache and runtime overhead, which is what determines whether a model can fit. The batching buffer is a property of how you launch the server, not of the model, and on a card this full it is the difference between 170,000 tokens of context and 250,000.
So treat our context figures as the arithmetic and that flag as the thing that decides whether you reach them. If you are serving several users, leave the reservation alone. It is doing its job.
The Same Trick Is Now in Three Places, and Nobody Prices the Drafter
Updated 21 August. Speculative decoding stopped being a Qwen story this week. Liquid AI shipped DSpark draft models for three LFM2.5 models on 20 August, and the single-node DeepSeek V4 Flash build that runs on one DGX Spark is tuned for DSpark too. Same idea as DFlash 2 above: a small model proposes a block of tokens, the big one checks them in a single pass.
The quality guarantee is the part worth understanding. Liquid states that under greedy decoding a draft token is accepted only if it matches the target’s own distribution, so the emitted sequence is identical to baseline by construction and benchmark accuracy is unchanged. That is the difference between speculative decoding and quantization. One is free accuracy-wise and costs memory. The other costs accuracy and saves memory.
So what does it cost? Liquid’s blog says a “minimal memory increase” and never puts a number on it. Neither does anyone else. The drafters are 295.7M and 327.7M parameters, with the embeddings and LM head tied to the target, so that count is the incremental part. At bf16, against each target at a 4-bit class quantization:
| Target | Drafter | Added memory | As a share of the target | H100 mean | M4 Max mean |
|---|---|---|---|---|---|
| LFM2.5-1.2B-Instruct | 295.7M | 0.55 GiB | 88% | 2.10x, 656 to 1384 | 2.54x, 138 to 350 |
| LFM2.5-2.6B | 327.7M | 0.61 GiB | 45% | 2.67x, 323 to 864 | 2.27x, 61 to 139 |
| LFM2.5-8B-A1B | 327.7M | 0.61 GiB | 14% | 2.54x, 418 to 1074 | 1.18x, 90 to 106 |
Read the last two columns against the memory column, because they run in opposite directions. On a Mac, the model where the drafter is cheapest in relative terms is the one where it barely helps: the 8B-A1B adds 14% memory for 18% more speed. The 1.2B nearly doubles its own footprint and returns 2.54x. Whether that trade is worth taking depends entirely on which end of the family you are on, and no announcement puts those two facts in the same table.
Liquid attributes the weak Mac result to how mixture-of-experts models currently run on Metal in llama.cpp, not to the method. It is the one configuration where the drafter looks like a bad deal today, and the reason is software rather than arithmetic.
One number in circulation is the best case, not the average. The 3.18x being quoted is the top of the range for the 8B-A1B on an H100, across datasets. The mean for that same configuration is 2.54x, and the low end is 1.29x. Liquid publishes all three. The post that reaches you usually publishes one.
Our arithmetic, stated so you can check it: the added memory is the published parameter count at bf16, and the share column compares it to the target at roughly 4.5 bits per weight. A quantized drafter would cut that column and Liquid has not published quantized drafter sizes, so treat these as the ceiling rather than the only option.
FAQ
Does speculative decoding still help at 200k context?
Yes. On a single RTX 4090, @analogalok measured DFlash 2 decoding at 41.5 tokens per second on a 212,000 token prompt, keeping a 10 to 15% lead over native MTP at every depth from 26k upward. Native MTP reads the prompt faster throughout, so the choice depends on whether you wait on ingestion or on generation.
How much memory does a speculative decoding draft model add?
For Liquid’s DSpark drafters, 0.55 to 0.61 GiB at bf16, since the embeddings and LM head are shared with the target. What matters is the share rather than the figure: that is 14% on top of an 8B target and 88% on top of a 1.2B one. The smaller your model, the worse the trade looks.
What is DFlash speculative decoding?
A small draft model proposes blocks of tokens in parallel, and the full target model verifies them, keeping only what it would have written. Output is identical to running the target alone, because the target is still the one deciding every token.
Is DFlash 2 the same as MTP?
No. MTP is a prediction head trained into Qwen 3.8 27B’s own weights and switched on with a flag. DFlash 2 is a separate block-diffusion draft model, 3.85 GB at full precision and 1.14 GB as Q4_K_M, that works alongside the target. They stack.
Does DFlash 2 really double Qwen 3.8 27B?
Against a plain baseline, one A100 test measured 28.9 to 59.1 tok/s. On top of MTP, one RTX 4090 test measured 60 to about 90 tok/s. Both are single testers, so treat them as what they are, two data points, not a spec.
How much VRAM does DFlash 2 add?
The draft model is 1.14 GB as Q4_K_M (3.85 GB at full precision), plus memory for the draft states, which grows with the draft depth. On a 24 GB card, --spec-draft-n-max 4 keeps the base model plus draft under the ceiling while holding around 83 to 87 tok/s decode.
Do I need extra VRAM to run DFlash 2?
Yes. You need the full target model resident first, then the draft model on top. A card that cannot hold the base Qwen 3.8 27B at your chosen quant cannot use DFlash 2 to fix that.
Is DFlash 2 in llama.cpp?
It is available as an open pull request, #27342, patchable today but not yet merged, and its llama.cpp path has open bug reports. It also ships as open pull requests for vLLM and SGLang and an oMLX fork.
Does DFlash 2 change the model’s output?
No. Speculative decoding is lossless by construction. The draft proposes; the target verifies. If the draft is wrong, the target corrects it, and the result is exactly what the target would have produced on its own.