What the Athena Engine does on one DGX Spark
Athena’s Engine is an inference server for a single NVIDIA GB10 box, the DGX Spark or the ASUS Ascent GX10, and it runs two models: DeepSeek V4 Flash and Qwen3.8 Flash-Next. It is written by Marco Palaferri, and its install script pulls the weights for you. It answers on the same HTTP endpoints the big providers use, which is how a local coding agent talks to it: point OpenCode, Hermes or anything else that expects those endpoints at your own box instead of a paid account. Free for personal use, research and study under PolyForm Strict 1.0.0, with a separate licence for commercial use.
The numbers it publishes are unusual in one specific way, and the reposts have all picked the same one. It is worth looking at what is underneath.
The published figures
From the project’s own README, measured on a GB10 with a 256K context configured:
| Model | Prefill at 8K | Prefill at 256K | Decode at 8K | Decode at 256K |
|---|---|---|---|---|
| DeepSeek V4 Flash | 1,126 tok/s | 948 tok/s | 21.4 tok/s | 19.4 tok/s |
| Qwen3.8 Flash-Next | 1,071 tok/s | 961 tok/s | 29.9 tok/s | 32.1 tok/s |
Decode is measured over 256 generated tokens. The files are a DeepSeek IQ2_XXS mix with Q8 projections and shared experts, about 87 GB, and Unsloth’s UD-IQ4_XS for Qwen, 93.68 GB in Hugging Face’s listing. The installer’s 97 GB adds a 2.79 GB Q8 draft head and a 0.90 GB vision projector on top of the weights.
Two things here are worth a reader’s attention, and they are not the same thing. Prefill barely moves between 8K and 256K. And on Qwen, decode at 256K is higher than at 8K.
The flat curve is real, and it is the architecture
A conventional transformer slows down as the conversation grows, because every new token attends over everything before it. These two models are built differently. Qwen3.8 Flash-Next puts 36 of its 48 layers on a fixed-size recurrent state and gives only 12 layers a cache that grows, which is why our Flash-Next memory page shows the jump from 8K to 256K costing 5.9 GiB rather than the 15.8 GiB the same trip costs a dense 27B.
Speed follows memory here, and our own corpus already carried the shape from other people’s runs. @redp314’s ladder on one Spark reads 43.04 tok/s at 4,292 tokens and 44.22 at 258,790. On the dense Qwen3.8 27B, which holds 48 of its 64 layers on the same kind of fixed state, @analogalok’s 4090 ladder sits between 40.68 and 40.96 tok/s from 80,000 tokens to 260,000, dropping the KV cache from f16 to q4_0 as it climbs.
So Athena’s flat curve is a property of the models it serves, not a discovery about the engine. The engine earns credit for not spoiling it.
What the reposts left out: two recipes decode faster at the same depth
This is the part we can answer and a repost cannot, because it needs other people’s runs on the same hardware. For Qwen3.8 Flash-Next on one DGX Spark, our tokens per second tool holds these:
| Recipe | What it holds | Context | Decode |
|---|---|---|---|
| Athena’s Engine, MTP draft head | UD-IQ4_XS GGUF, 87.25 GiB of weights | 256K | 32.1 tok/s |
| @redp314, vLLM with MTP 3 | NVIDIA’s NVFP4 checkpoint, about 76 GiB resident | 258,790 | 44.22 tok/s |
| @ViC305, ExLlamaV3 fork, MTP depth 5, 8-bit KV | EXL3, 4.05 bpw on his earlier pack | 240,000 | 72 tok/s |
The two that win are not running heavier files, and that is the part the reposts cannot reach. NVIDIA’s NVFP4 checkpoint is 132.68 GB on Hugging Face, and what a Spark actually holds is a good deal less. @Tech2Wild’s recipe leaves the 47.68 GiB FP8 n-gram table on the NVMe and reads the sixteen rows each token needs, so about 76 GiB of weights stay resident. Athena holds 87.25 GiB. The recipe that decodes 38% faster is reading fewer bytes per token, not more.
The bit width is not the lever here. Only the routed expert layers of the NVFP4 build are four-bit at all: NVIDIA’s card keeps attention, the shared experts and the rest of the main model in BF16, and puts the MTP module and the n-gram table in FP8. What moves decode on this model is where the 51.2B n-gram table sits. Athena’s install carries it inside the GGUF; the two faster recipes keep it off the hot path.
Two honest caveats, because this is the kind of comparison that goes wrong:
- The DeepSeek side is not comparable at all. Athena’s DeepSeek file is a two-bit-class mix. Quantization that aggressive is a different quality proposition from the FP8 and NVFP4 builds people run elsewhere, so its 19.4 tok/s belongs in a different conversation, not this table.
- This is a speed table, and only that. Each row names what it holds and the depth it was measured at. Which of the three writes better code is a different measurement, and one worth having.
- Each row was measured under different conditions, and here they are. Athena’s figure is 256 generated tokens at temperature 0. @ViC305 ran an actual 240,000-token prompt with an 8-bit KV cache and an MTP draft depth of 5. @redp314’s is a vLLM sweep with MTP 3 and prefix caching on. All three speculate, which is the one thing they do share.
The feature that is actually new
Strip the speed out and one claim in Athena’s README is doing something the others do not: a 141,519-token conversation comes back in 2.1 seconds, against the two minutes and twenty seconds its first reading took. vLLM’s prefix cache saves you the reprocessing inside a session. Athena writes long contexts to disk and restores them across sessions.
For a coding agent that reopens the same repository every morning, that is worth more than 10 tok/s. Prefill at 961 tok/s still means about two and a half minutes to read a 141K conversation from cold. Doing that once instead of every session is the difference between a habit and a chore.
It is also the claim we would most like to see someone reproduce, because it is the one that would change how people work rather than by how much.
What it costs in memory
A GB10 box has 128 GB of unified memory, of which 112 to 119 GiB is usable depending on how far the operating system is tuned, as our DGX Spark guide sets out.
Athena’s Qwen file is 93.68 GB, which is 87.25 GiB. At 256K of context the cache and the recurrent state add about 6.1 GiB at f16 by our engine’s arithmetic, and the draft head and vision projector add another 3.4 GiB, so the set lands near 97 GiB. That leaves 15 GiB on an untuned box and 22 on a tuned one, for the runtime and the operating system.
That is why the engine ships an IQ4_XS build rather than the 111.33 GB UD-Q4_K_XL: the heavier file plus a 256K cache would leave almost nothing. If your box is not a Spark, check what your own hardware holds before picking a quant.
Who should run it
- If you have a GB10 box and want one install that serves both models to your agent, with tool calling, and with images and documents on the Qwen side, this is the least work of any option on this page. Your agent keeps pointing at localhost.
- If you want the fastest decode on Flash-Next, the EXL3 route at 72 tok/s and the NVFP4 vLLM route at 44 are both published, both measured at a real depth, and both in our tool with their settings.
- If you reopen enormous conversations, the disk cache is the reason to try Athena before the faster recipes.
- If your box is not a GB10, none of this applies yet. Athena is built for that one machine.
FAQ
What hardware does the Athena Engine need?
An NVIDIA GB10 system with 128 GB of unified memory, which means a DGX Spark or an ASUS Ascent GX10. The project targets that one platform.
How fast is the Athena Engine on a DGX Spark?
Its own README gives 1,071 tok/s prefill at 8K and 961 at 256K for Qwen3.8 Flash-Next, with decode of 29.9 tok/s at 8K and 32.1 at 256K. DeepSeek V4 Flash reads 1,126 and 948 on prefill, 21.4 and 19.4 on decode.
Is the Athena Engine the fastest way to run Qwen3.8 Flash-Next on one Spark?
On decode, two published recipes are faster at a comparable depth: an ExLlamaV3 route at 72 tok/s at 240,000 tokens and a vLLM NVFP4 route at 44.22 at 258,790. Both keep the model’s 51.2B n-gram table off the GPU, so they read fewer bytes per token than Athena does. Its advantages are the packaging and the cached-context restore.
Why does decode get faster at 256K than at 8K?
Both models keep most of their layers on a fixed-size state, so the part of the cache that grows is small. The README’s own summary gives the Qwen range as 29 to 35 tok/s across the whole window, and 29.9 and 32.1 both sit inside it, which is why we read those two numbers as flat rather than as an improvement.
Is the Athena Engine free?
It is free for personal use, research, study and hobby projects, and commercial use needs a separate licence, per the project’s own terms.