TensorFold: Speculative Decoding That Never Changes a Token

What TensorFold is

TensorFold is an open-source inference server for Apple Silicon and NVIDIA GPUs that speeds up decoding with draft tokens, and promises that the drafts never change a single token of the reply. It comes from @ashxhart, the repository is MIT licensed, and it serves an OpenAI-compatible API on your own machine.

The idea is speculative decoding: a small, fast predictor proposes several tokens, and the model checks them together instead of writing one at a time. Most engines do this. What TensorFold adds is a rule its release notes check almost every time. A drafted token is accepted only when it equals the token the same engine would have produced serially, and each release checks that drafted replies are byte-identical to the same request sent with drafting off. The README puts it plainly: "A draft is accepted only when it equals the token the same engine would produce serially." The promise has a scope: it holds for the same engine, weights and settings, not between the Mac and NVIDIA engines or between two quantizations.

It runs five model families: Qwen 3.8 27B, Qwen 3.8 Flash Next, Nemotron 3.5 Lightning, GLM-5.3-Flash and Gemma 4 26B-A4B, each through a named checkpoint and drafter. The checkpoints it names come mostly from Vontra, the Hugging Face organisation its author links from his GitHub profile. Macs run MLX checkpoints; NVIDIA GPUs, including the DGX Spark, run a CUDA engine inside NVIDIA's PyTorch container.

Four days, eleven releases

The repository dates from June, but the version people are running now began with a rewrite on 25 September. What followed, from the project's changelog:

Date Release What changed
25 Sep 0.2.0 Rewrite: Nemotron, Qwen 3.8 27B and Flash Next on Apple Silicon
26 Sep 0.3.0 NVIDIA GPUs, one or two DGX Sparks; GLM-5.3-Flash on two Sparks
26 Sep 0.3.1 tensorfold info shows a checkpoint's format and refuses unsupported ones before downloading
26 Sep 0.3.2 tensorfold update; GLM-5.3-Flash reads EXL3 weights on two Sparks, experimental
26 Sep 0.3.3 Qwen 3.8 27B drafting exact on every M1 to M5 GPU
26 Sep 0.3.4 (pre-release) Every model on the lane engine; the 27B on M1 to M4 at 1.9 to 4 times serial
27 Sep 0.3.4.1 Prompt processing back to MLX's own speed
27 Sep 0.3.5 Concurrent requests share each verification round; follow-up turns resume
28 Sep 0.3.5.1 Qwen 3.8 27B loads on M1 and M2 again
28 Sep 0.3.6 GLM-5.3-Flash and Gemma 4 on Macs, EXL3 checkpoints on NVIDIA GPUs
28 Sep 0.3.6.1 CUDA builds inside NVIDIA's containers again

From 0.3.5 on, the notes thank contributors by handle, a dozen names in the last four releases. On 28 September the repository had 538 stars and 54 forks.

The numbers on a Mac

The benchmark behind TensorFold's speed tables uses two public prompts, a short code task and a chat question, 64-token replies, five seeds and the median decode rate. That is a narrow test, and the notes say which cells are code and which are chat because the two behave differently.

Qwen 3.8 27B on an M3 Ultra, 0.3.4, MLX 0.32.0, 64-token replies, thinking off:

Code, sampled Chat, sampled Code, greedy Chat, greedy
MLX serial 38.2 38.2 39.3 39.3
TensorFold 0.3.4 141.3 73.9 158.4 74.2

On code that is 3.7 to 4 times serial on the same Mac; on chat, about 1.9 times. 0.3.5 then measured another 17 to 20% on the same M3 Ultra against 0.3.4.1. For the same model on other hardware and engines, see Qwen 3.8 27B speed.

Qwen 3.8 Flash Next on an M3 Ultra, 0.3.4.1, from the notes: 135.5, 111.1, 142.3 and 122.2 tok/s across the same four cells. GLM-5.3-Flash on an M3 Ultra, 0.3.6: 1.86 to 2.08 times mlx-vlm, and the notes say plainly that this model misses the project's own floor of three times.

Qwen 3.8 27B on an M5 Max, 0.3.4, with the DFlash 2 draft model: 153.2 and 154.2 tok/s on the code cells, 69.8 and 72.8 on chat. The first release, 0.2.0, had given 120 to 124 on short answers.

The numbers on a DGX Spark

On NVIDIA hardware the comparison is against vLLM with multi-token prediction at three drafts, the baseline TensorFold's recipe uses. From TensorFold's Qwen 3.8 27B recipe, one DGX Spark, medians of 15 runs a cell:

Cell TensorFold, EXL3 3.00 bpw TensorFold, MLX 4-bit vLLM, MTP=3
Code, sampled 83.4 57.5 23.4
Chat, sampled 44.2 49.7 25.4
Code, greedy 64.5 53.6 25.8
Chat, greedy 39.9 50.0 24.7

The EXL3 pack is turboderp's, read through TensorFold's experimental EXL3 path. It wins on code and loses on chat, and the recipe gives the reason: the DFlash 2 draft model was trained on the unquantized model, so fewer of its drafts survive on the 3-bit pack in chat, 3.0 to 3.3 tokens a round against 4.2. To check what a card of your own should deliver, the tokens per second checker estimates it.

One independent run. On 26 September @WescheNex1q ran TensorFold on one DGX Spark against vLLM with MTP=3 on NVFP4, on three prompts of his own: 102.9 tok/s against 34.9 averaged over the three (75.7 to 126.9 against 32.7 to 37.0 per prompt), TensorFold on its 4-bit MLX weights, with byte-identical output to serial and a first token in 0.12 s. His own caveat, in the post: the two runs were not weight-matched.

What it needs in memory

TensorFold's README has a table of Mac memory classes and what each can run. On 28 September every model cell in it read "TBD", pending a public prompt fixture and measured peak memory for each result. What the README does give is the rule: the process stays inside 70% of RAM, or the GPU's recommended working set if that is lower, and 3 GiB of that is reserved for everything outside the model.

So the question can be sized from the files. The Qwen 3.8 27B checkpoint TensorFold names is 16.05 GB, and its DFlash 2 draft model 3.85 GB. Our engine puts the 27B's cache at 64 KiB per token of context at 16 bits, on its 16 full-attention layers, plus a fixed 0.15 GiB for the other 48. Against TensorFold's budget:

Mac memory TensorFold's MLX limit (70% less 3 GiB) Left after model and drafter Context that fits, at most
24 GB 13.8 GiB none does not fit
32 GB 19.4 GiB 0.9 GiB about 8K
36 GB 22.2 GiB 3.7 GiB about 56K
48 GB 30.6 GiB 12.1 GiB about 188K
64 GB and up 41.8 GiB and up 23.3 GiB and up the full 256K

These are ceilings. On Macs where macOS gives the GPU less than 70%, the ceiling is lower still. TensorFold also reserves the reply length and a prompt workspace before it admits a request, so the context it reports at startup will be lower, and it prints that figure. Our Mac checker sizes the same model under macOS's own GPU limit.

The larger families follow the same arithmetic. Flash Next's MLX checkpoint with its MTP head is 113.21 GB, so it sits fully in memory only from a 192 GB Mac up. TensorFold's 0.3.6 adds --ple-on-ssd, which reads Flash Next's n-gram tables from disk: on a 128 GB budget it peaked at 85.6 GiB and decoded at 0.91 to 1.03 times the 256 GB run. --ssd-experts goes further and streams routed experts from the SSD, fitting Flash Next in a 64 GB Mac and GLM-5.3-Flash in a 128 GB one, at 0.31 to 0.39 and 0.13 to 0.17 times their resident speed. Those are TensorFold's measurements on an M3 Ultra under emulated budgets.

How it compares with Inco Splash

Both engines arrived in the same fortnight and both bet on drafting for Macs. They make different bets. Inco Splash compiles kernels for each model's exact shapes and, a week after launch, added Unsloth GGUF files. TensorFold verifies drafts in batched lanes, guarantees output identical to serial decoding, also runs on NVIDIA GPUs, where the 27B, Flash Next and GLM-5.3-Flash have published DGX Spark numbers, and reads MLX and EXL3 checkpoints. Their published numbers come from different Macs and different prompts, so this page does not rank them.

How to run it

On a Mac, with Python 3.11 or newer:

python -m pip install git+https://github.com/ashhart/TensorFold.git
tensorfold serve Vontra/Qwen3.8-27B-MLX-4bit

The server listens on 127.0.0.1:8080/v1, and tensorfold pull downloads a checkpoint and its draft model ahead of time. On a DGX Spark, the README runs it inside NVIDIA's pytorch:26.07-py3 container and pulls the DFlash 2 draft model before serving the 27B. tensorfold update installs the newest stable release, which this week has meant most days.

FAQ

What is TensorFold?

An open-source, MIT-licensed inference server for Apple Silicon and NVIDIA GPUs that uses draft tokens to decode faster, and accepts a draft only when it equals the token serial decoding would produce, so replies stay byte-identical.

How fast is TensorFold on a Mac?

On an M3 Ultra, its notes put Qwen 3.8 27B at 141.3 to 158.4 tok/s on a short code prompt against 38.2 to 39.3 for MLX serial, and about 74 on chat. In every Mac result TensorFold publishes, code runs faster than chat.

Is TensorFold faster than vLLM on a DGX Spark?

On TensorFold's own benchmark, with turboderp's EXL3 3.00 bpw pack, Qwen 3.8 27B decodes code at 83.4 tok/s against 23.4 for vLLM with MTP=3, and chat at 44.2 against 25.4. @WescheNex1q measured 102.9 against 34.9 on his own prompts, and noted the runs were not weight-matched.

Which Mac can run Qwen 3.8 27B with TensorFold?

By file sizes and TensorFold's 70% memory limit, a 36 GB Mac holds the 27B and its draft model with about 56K of context at most, and a 48 GB Mac about 188K; both are ceilings, lower where macOS gives the GPU less than 70%. A 24 GB Mac does not fit it.

Which models does TensorFold support?

Qwen 3.8 27B, Qwen 3.8 Flash Next, Nemotron 3.5 Lightning, GLM-5.3-Flash and Gemma 4 26B-A4B, each through named checkpoints, with experimental EXL3 checkpoints for the two Qwen models on NVIDIA GPUs, and an experimental EXL3 conversion of GLM-5.3-Flash on two DGX Sparks.


Leave a Comment