Abliterated model speed: faster at one stream on four DGX Sparks
@Tech2Wild served GLM-5.3-Flash twice on the same four NVIDIA DGX Spark (GB10) machines at tensor parallel 4, and the abliterated weights came out ahead on seven of the nine single-stream prompt categories he recorded, level on one and behind on maths. The two lanes differ in one respect: which NVFP4 pack he loaded. Lane A is nvidia/GLM-5.3-Flash-NVFP4, MIT licensed and ungated. Lane B is Blackfrost-AI/GLM-5.3-Flash-DERISKED-NVFP4, MIT licensed, which Blackfrost derived from zai-org/GLM-5.3-Flash-BF16 through its own de-risked BF16 master and @Tech2Wild treats as an abliteration. Weights are by @Blackfrost_AI, whose own evaluation reports 1.3% refusal on harmful prompts (4 of 300). Everything else held still, and what the model needs in memory is on our GLM-5.3-Flash VRAM page. Same vLLM build, same incoai/GLM-5.3-Flash-DFlash2 block diffusion drafter at k=7, same 500,000 token window on a 3,532,196 token fp8 KV pool at 43.76 GiB per rank, same 14 entry ignore list, same W4A16_NVFP4 recipe, same knobs. The weights are the only variable, which is what makes this a single-variable test of abliterated model speed rather than a story about two different setups.
This is one operator’s fleet. Four machines, one recipe, one model, medians of three passes at temperature 0 after a discarded warm-up. Read from his repository on 21 September 2026 and unchanged when re-read on 25 September, with announcement posts on 20 and 21 September 2026.
The numbers
Single-stream decode in tok/s, except the C3 row, which is aggregate throughput at three concurrent streams.
| Category | Abliterated | Censored |
|---|---|---|
| Code | 95.9 | 78.7 |
| Counting | 138.1 | 107.7 |
| Prose | 50.8 | 40.8 |
| Maths | 83.3 | 88.8 |
| C3 aggregate | 125.0 | 117.0 |
| Draft tokens committed per step | 3.90 | 3.76 |
| Draft acceptance rate | 0.415 | 0.394 |
Maths is the one category below, about 6% lower, and @Tech2Wild places it at the edge of its own spread. The code row needs a caveat he writes himself. The censored lane’s code cell is bimodal: 78.7, 92.0 and 78.7 across its three passes, a spread of 1.17x against the abliterated lane’s 1.01x, and he warns that the median gap overstates the typical difference because it leans on the two low passes. The reading that survives is the paired one: the abliterated lane’s slowest code pass, 94.62 tok/s, beats the censored lane’s fastest, 92.04.
Lane A also measured a cold prefill of 1,997 tok/s on a 40,659 token prompt, with a time to first token of 20.4 seconds. Boot to a healthy endpoint is about 12 minutes. His quality gate passes.
Why the earlier abliterated lane broke
The retired lane is the dealignai o_proj transplant. It substituted o_proj tensors from dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4 on layers 12 to 44 before quantization. It did stop refusing. It also garbled in live agentic use in two unrelated harnesses: emoji runs, injected foreign script fragments, duplicated paragraphs and a degenerate tail. Its draft acceptance collapsed to 0.224, or 2.57 tokens committed per step, with position 7 at exactly zero.
@Tech2Wild’s attribution measurement puts refusal in the routed expert down_proj tensors, which move refusal by 0.81, against 0.03 for attention plus the shared and dense MLP. That pack edits 46 o_proj tensors; by his count the real intervention is 12,384 down_proj tensors. His reading is that the transplant bought non-refusal by perturbing attention, and that the same edit made it incoherent over long contexts: one cause, two symptoms. He calls that link strong but inferential, since seven probes failed to reproduce the garbling directly.
Blackfrost does not publish which tensors it changed; its model card calls the production details proprietary. @Tech2Wild located the edit himself by comparing six sampled tensors per class against NVIDIA’s pack. Every routed-expert down_proj sample differed by about 10%, 0.092 to 0.108. Expert gate_proj and up_proj samples ranged from byte-identical to 0.082, which he leaves unresolved at that sample size. Token corruption was not seen on either current lane in operator use. It tracked the retired transplant. He states the mechanism is still unmeasured, and a tool-call loop seen on the retired lane has not yet been re-checked on the Blackfrost one.
That distinction is the part a reader can act on before downloading anything. The question is not whether a pack is abliterated. The question is which tensor class the abliteration touched. A pack that edits down_proj is doing the intervention the attribution measurement points at. A pack that edits o_proj is doing something else and paying for it in attention.
Two traps in the pack itself
Blackfrost’s config.json declares quant_algo NVFP4, which is the W4A4 path, while by @Tech2Wild’s count the pack ships zero input_scale tensors. W4A4 kernels would then multiply real weight scales by uninitialized memory. Its own config_groups says input_activations is null, so weight only is the intent and the string is mislabelled. A script in @Tech2Wild’s repository rewrites it to W4A16_NVFP4.
NVFP4_PATCH is not a tuning flag. It bind mounts two patched vLLM files that stop quant_config being forced to None on the attention projections. Without them vLLM builds those layers BF16, a packed 4 bit weight has nowhere to load, and the build does not come up at all. A reader who copies the launch command without the patch gets a failure that looks like a model problem and is a loader problem.
What this does and does not say
The GLM-5.3-Flash result is one model, one operator, one recipe, four machines. It says that on this pair, with these weights, the abliterated lane measured faster at one stream on seven of the nine prompt categories he ran, and on draft acceptance and committed tokens per step, and slower on maths by about 6%, at the edge of its own spread. Under load the lead does not hold. Aggregate throughput favours the abliterated lane from one to three streams, 125.0 against 117.0 at three, and then the lanes trade places: the censored lane measured 172.4 tok/s against 165.5 at six streams and 352.8 against 351.9 at 32, while the abliterated lane led at eight and 24. It does not say that abliteration raises speed as a general property. The mechanism behind the code gap is not established by the numbers above. The draft acceptance difference is small, and what causes it is not settled by anything in his repository.
One second operator has served the same Blackfrost pack on four Sparks since. GitHub user joesinvestments built it on vLLM 0.30.0 and on 22 and 23 September measured 0.391 draft acceptance and 3.74 tokens per step with his own prompt set, level with @Tech2Wild’s censored lane rather than his abliterated one. The harness differs, so it is not a contradiction, but the acceptance edge has not been reproduced. His speed comparison points the same way as @Tech2Wild’s, with a weaker control: against a stock RedHatAI NVFP4 pack, which differs in more than the weights, he found stock 10 to 25% slower for one user and about even at 32 users.
The weaker prior evidence comes from a different model. Qwen3.8 27B is a dense 27B model, not GLM-5.3-Flash, and it says abliteration cost nothing, which is not the same claim. @filicroval, on 30 August 2026, served three Qwen3.8 27B NVFP4 builds on one DGX Spark with SGLang and DFlash2: the stock RadixArk build and two abliterated ones. Without speculative decoding all three ran at 12.3 tok/s (12.34, 12.31 and 12.33). With DFlash2 the Huihui abliterated build reached 57.95 tok/s, 4.71x its own baseline, against 57.11 and 4.63x for stock. Those are whole-request rates, prefill and first token included. The one build that fell well behind, a second abliterated pack at 3.36x, keeps its lm_head in bf16, and he puts the gap down to that single tensor rather than to the abliteration.
He pinned revisions, published sha256 sums, ran a frozen quality gate that scored 24 of 27 both with and without speculative decoding, and included a reproduction script. That is the standard the GLM-5.3-Flash pair is being held to as well, and it is why the pair is worth reading rather than skimming.
File size will not tell you which pack you have
At full precision an abliterated build is the same size as the original. Blackfrost-AI’s BF16 release of Qwen3.8 27B and Qwen’s own BF16 both hold 27,781,427,952 parameters in 18 shards, 55.56 GB, and the shard sizes match file for file. Abliteration edits weights in place. Nothing is added and nothing is removed.
Quantized builds are worse as a signal. On 25 September 2026, three publishers’ Q4_K_M files of the untouched base model ran from 16.81 GB (lmstudio-community) through 17.44 GB (bartowski) to 18.97 GB (ggml-org), and unsloth’s current Q4_K_M, its UD dynamic build, is smaller still at 16.46 GB. Two abliterated Q4_K_M files sit inside that range: 0bserverx’s at 16.55 GB and Blackfrost-AI’s at 16.81 GB, the same size as lmstudio’s base file to two decimal places. The spread is the quantizer. A reader comparing file sizes to decide which pack is which is reading the wrong column.
Popularity will not tell you either. In the month to 25 September, Hugging Face counted about 2.2 million downloads of huihui-ai’s abliterated Qwen3.8 27B GGUF and 1.6 million of 0bserverx’s, against 6.9 million for unsloth’s standard repository. Download count is a popularity signal, not a quality signal, and it does not tell you which tensor class was edited.
What to check before you pull one
Read the pack’s config.json and confirm the quant_algo string matches the tensors actually shipped. A declared W4A4 path with zero input_scale tensors is a mislabel, and the fix is a rewrite to W4A16_NVFP4, which @Tech2Wild’s repository does with a script.
Confirm the launch recipe includes whatever patch the pack needs. For this pair that is NVFP4_PATCH, which bind mounts two patched vLLM files. Without it the attention projections build BF16 and the packed 4 bit weights have nowhere to load.
Ask which tensor class the abliteration touched. The attribution measurement from @Tech2Wild puts refusal in the routed expert down_proj tensors at 0.81, against 0.03 for attention plus the shared and dense MLP. A pack that edits down_proj is doing that intervention. A pack that edits o_proj is perturbing attention, and the retired transplant is what that looks like when it goes wrong over long contexts.
Check whether the pack ships a quality gate and a reproduction script. @filicroval’s Qwen3.8 27B work pinned revisions, published sha256 sums and scored 24 of 27 with and without speculative decoding. That is the shape of evidence that lets a reader compare two packs without running both.
Then measure on your own hardware. The numbers above are @Tech2Wild’s, on four DGX Sparks at tensor parallel 4, with a 500,000 token window on a 3,532,196 token fp8 KV pool at 43.76 GiB per rank. A single machine, a different quantizer or a different drafter will move the draft acceptance rate and the committed tokens per step. Whether those two numbers explain the rest of the gap is not something his repository settles. For background on how those figures are usually reported, see tokens per second. For the broader category, see abliterated models. The Qwen3.8 27B comparison is covered at Qwen3.8 27B speed.
FAQ
Does an abliterated model run slower than the original?
Not in the two single-variable comparisons covered here. @Tech2Wild’s abliterated GLM-5.3-Flash lane was faster at one stream on seven of nine prompt categories on four DGX Sparks, and @filicroval’s abliterated Qwen3.8 27B ran level with the stock build on one Spark, 57.95 against 57.11 tok/s. Under load the two GLM-5.3-Flash lanes trade places.
Why was the abliterated GLM-5.3-Flash lane faster?
His repository does not settle it. The abliterated lane committed 3.90 drafted tokens per step against 3.76 with the same DFlash2 drafter, but a second operator serving the same pack measured 0.391 acceptance and 3.74 tokens per step, level with the censored lane.
Does an abliterated model need more VRAM?
No. At full precision Blackfrost-AI’s abliterated Qwen3.8 27B and Qwen’s own BF16 hold the same 27,781,427,952 parameters in 55.56 GB. Quantized files vary with the publisher: on 25 September 2026, base Q4_K_M builds ran from 16.46 to 18.97 GB, and two abliterated ones sat inside that range.
How do I choose between abliterated packs?
Ask which tensor class was edited, check that config.json’s quant_algo matches the tensors actually shipped, confirm the launch recipe includes any patch the pack needs, and prefer a pack that ships a quality gate and a reproduction script. File size and download count will not tell you.
What hardware did @Tech2Wild use?
Four NVIDIA DGX Spark (GB10) machines at tensor parallel 4, vLLM with the incoai GLM-5.3-Flash DFlash2 drafter at k=7, and a 500,000 token window on a 3,532,196 token fp8 KV pool at 43.76 GiB per rank.