What Inco Splash is, and which Macs can run it
Inco Splash is an open-source inference engine for Apple silicon that runs two model families, Qwen3.8 27B and Qwen3.6 35B-A3B, each with a DFlash 2 speculative decoding draft matched to it. It comes from Inco, the team behind the DFlash 2 draft models, under the Apache 2.0 licence, and LM Studio added it as a backend on launch day, 19 September.
Updated 27 September. A week after launch, Splash loads ordinary Unsloth GGUF files and mlx-community MLX checkpoints of both models, and runs on 24 GB Macs. Inco has also published a test against llama.cpp on the same file. This page is updated from Inco's README.
It needs an Apple M3 or newer and macOS 26.4 or later. How much memory depends on the file you load. Inco's 4-bit examples need 36 GB of unified memory, with 48 GB recommended, and smaller GGUF files run on 24 GB Macs, per the README. An M1 or M2 Mac does not qualify at any memory size, because the chip requirement comes first.
On our Mac catalogue, the 36 GB floor for the 4-bit files is met exactly by four machines: the 36 GB M3 Pro, M3 Max, M4 Max and M5 Max. Every M3, M4 or M5 Pro, Max and Ultra configuration with 48 GB or more clears it too.
At launch, Splash ran only Inco's own packages: Qwen3.8 27B at 17.4 GB to download, holding the model at 4 bits, its DFlash 2 draft and a 0.93 GB vision file, and Qwen3.6 35B-A3B at 20.9 GB. Now it also takes Unsloth's GGUF builds of both models from 1 to 8 bits, including the mixed-precision UD files, except UD-Q8_K_XL and BF16. PrismML's Ternary Bonsai 2 27B runs too, in its 7.2 GB PQ2_0 format.
The numbers Inco wrote down, and the 144
Inco's written benchmark runs on a 48 GB M5 Pro with a 16-core GPU, using coding prompts from NVIDIA's SPEED-Bench, a 1,024-token output limit and reasoning on. On Qwen3.8 27B:
| Measure | Splash | Next fastest engine |
|---|---|---|
| Decode, short prompt | 74 tok/s | 2.0 times slower |
| Decode, 32K prompt | 54 tok/s | oMLX 28, Ollama 19, uzu 15 |
| Prefill, 32K prompt | 363 tok/s | 1.2 times slower |
| First token, 32K prompt already cached | 282 ms | oMLX 2,049 ms |
| Four requests at once, combined decode | 170 tok/s | oMLX 43 |
The Qwen3.6 35B-A3B package, a mixture of experts model with 3B active, is faster again on the same Mac: 210 tok/s on short prompts and 357 tok/s across four requests.
The 144 tok/s in the launch posts is a different Mac. Inco's post on X puts Qwen3.8 27B at 144 tok/s on an M5 Max MacBook Pro, over a demo video, and LM Studio's launch post repeats it as "up to 144". The M5 Max carries 614 GB/s of memory bandwidth in its 40-core GPU version (460 GB/s in the 32-core, 36 GB one) against 307 GB/s for the M5 Pro, per Apple's specifications, so a faster number on the bigger chip is expected. The benchmark table above is the one Inco published with its method, and it is the M5 Pro.
For scale, Inco's Zhijian Liu wrote on launch day that DFlash 2 ran Qwen3.8 27B at 70 tok/s on an M5 Max a month ago, and that Splash runs it at 144 on the same Mac. The same 144 shows up on a very different machine: on our Qwen3.8 27B speed page, @sudoingX measured 66 tok/s with MTP off and 144 with it on, on an RTX 5090. For what Qwen3.8 27B needs on a graphics card, see our Qwen3.8 VRAM page.
The same file, Splash against llama.cpp
Once Splash could load a GGUF, Inco ran the comparison everyone asked for: the same Unsloth UD-Q4_K_M files on Metal, Splash against llama.cpp (commit e6ab7c1). Decode speed in tok/s, from Inco's README:
| Model | Engine | M5 Pro | M3 Max |
|---|---|---|---|
| Qwen3.8 27B | llama.cpp | 16 | 17 |
| Qwen3.8 27B | llama.cpp with MTP | 27 | 20 |
| Qwen3.8 27B | Splash | 74 | 92 |
| Qwen3.6 35B-A3B | llama.cpp | 69 | 66 |
| Qwen3.6 35B-A3B | Splash | 175 | 209 |
On the 27B that is 4.5 to 5.3 times llama.cpp from the same file, and 2.7 to 4.6 times llama.cpp with MTP switched on. The README does not state the M3 Max's memory.
Is it still the same model? Inco checked that too. Over 16,384 positions of prose, code and chat, Splash and llama.cpp rank the same token first 99.30 to 99.45% of the time on the 27B. For comparison, llama.cpp agrees with itself between CPU and Metal 97.8% of the time. Splash ran a BF16 cache for this test; its default is an 8-bit cache.
On a 24 GB Mac. Inco's docs give two runs on a 24 GB M6 with a 12-core GPU, each with its draft:
| Unsloth GGUF | File size | Code decode | Context limit Splash reports |
|---|---|---|---|
| Qwen3.8 27B UD-IQ3_XXS | 10.93 GB | 43.5 tok/s | 102,393 tokens |
| Qwen3.6 35B-A3B UD-Q2_K_XL | 12.29 GB | 145 tok/s | 256K |
The context column is what fits, not the prompt length of the speed run. File sizes are from Unsloth's repositories. All of these are Inco's own measurements, on coding prompts.
Why it is fast
Inco's README gives three reasons, and none of them is a secret setting:
- The draft model is the decode path, not an option. Each package ships its own DFlash 2 draft. The draft proposes a block of tokens and the big model checks the whole block in one pass. Our DFlash 2 guide covers when that pays and when it does little: a draft only helps while enough of its tokens are accepted to pay for its own memory and compute.
- Kernels compiled for one model's exact shapes. Inco writes and tunes Metal kernels for each model and ships them precompiled, so nothing is tuned on your machine.
- A memory plan worked out at startup from what Metal allows on your Mac, minus the weights, the draft and each request's state. If the model does not fit, Splash prints the budget and stops.
That specialisation is by model family, not by file. Splash runs only the families Inco has written kernels for, two at launch, but it now loads them from GGUF or MLX files you may already have.
What it takes in memory
The Qwen3.8 27B package is about 15 GiB of weights and a 1.2 GiB draft, in Inco's launch post, before any conversation is held. The rest of the memory the GPU can use goes to the cache.
macOS does not give the GPU all of a Mac's memory. Our Mac checker works from about 65 to 75 percent of unified memory, which is the range owners report by default. On a 36 GB Mac that is roughly 23 to 27 GB for the GPU, so about 6 to 10 GB is left for the cache after the model and draft load. On a 48 GB Mac the same arithmetic leaves roughly 14 to 19 GB. That is our estimate, not Inco's. Inco's launch post gives its own reason for recommending 48 GB: room for an editor and a browser while a task runs.
On a 24 GB Mac the same range gives the GPU roughly 15.6 to 18 GB, which is why the 24 GB route uses files near 11 GB, such as the 27B at UD-IQ3_XXS.
Splash sizes the context automatically, up to the model's native window, and how much you get depends on that remainder. --max-context caps it by hand, --max-memory caps what Splash may allocate, and --max-cache-disk lets it move cache to the SSD when memory runs short, off by default.

The Local LLM Sizing Guide. Issue 1, 30 September
Choosing a Mac for local models? Know what 128 GB really holds first.
64 models fit a DGX Spark and 59 a 128 GB Mac by default, and Qwen3.8 Flash-Next drops from Q4_K_M to Q2_K on the Mac. Every model on every machine from 8 GB to 128 GB, with the context it holds.
From $9 a month. Card or crypto. The next issue included.
What an independent reviewer flagged
@Youssofal_, who builds the MTPLX engine for Apple silicon, looked at Splash on launch day and called it good work, with three caveats worth knowing:
- The weights are a flat 4 bits. On Qwen3.8 27B, he notes, the GDN layers are very sensitive to quantization, and in his MTPLX Optimized Speed model certain layers are tuned to 8 or even 16 bits. Since then this is a choice rather than a limit: the GGUF route loads Unsloth's UD files, which mix bit widths across layers.
- The KV cache is 8-bit. He runs the same setting on his own RTX 3090s and reports little quality loss, since the cache tolerates quantization better than the weights.
- The published benchmarks are coding prompts, where a DFlash draft is accepted most often. Speed on other kinds of text is the number to watch.
He also wrote that Inco's benchmarks are transparent and that the engine does not take obvious shortcuts.
How to run it
In a terminal, with Homebrew:
brew install incoai/tap/splash
splash serve --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
That is the README's own example. Any supported file goes after --model as OWNER/REPO:VARIANT, so a 24 GB Mac passes a smaller variant such as unsloth/Qwen3.8-27B-GGUF:UD-IQ3_XXS. The first run downloads the model and its matching draft, prepares the weights and serves on 127.0.0.1:8000; leave disk room for both the download and the prepared copy. It speaks the OpenAI and Anthropic APIs, so a coding agent can point at it, and splash opencode, splash claude, splash codex, splash hermes and splash pi launch those agents against it.
In LM Studio, it needs LM Studio Bionic 1.1.5 or newer: open Settings, then Runtime, and download Splash (Metal) under the experimental backends. LM Studio's Splash setup guide has the current steps.
FAQ
What Mac do I need for Inco Splash?
An Apple M3 or newer on macOS 26.4 or later. The 4-bit files need 36 GB of unified memory, and Inco recommends 48 GB; smaller GGUF files run on 24 GB Macs. No M1 or M2 Mac qualifies.
How fast is Inco Splash?
On Inco's written benchmark, a 48 GB M5 Pro, Qwen3.8 27B decodes at 74 tok/s on short coding prompts and 54 tok/s at 32K, about twice oMLX. On the same Unsloth file, Inco measured it at 4.5 to 5.3 times llama.cpp. The 144 tok/s figure comes from Inco's launch video on an M5 Max.
Does Inco Splash work in LM Studio?
Yes. LM Studio Bionic 1.1.5 or newer lists Splash (Metal) under Experimental backends in Settings, Runtime.
Which models does Inco Splash run?
Two model families, Qwen3.8 27B and Qwen3.6 35B-A3B, each with its own DFlash 2 draft, plus Ternary Bonsai 2 27B. Since the week after launch it loads them from Unsloth's GGUF files (1 to 8 bits, except UD-Q8_K_XL and BF16) or mlx-community's 4-bit checkpoints, as well as Inco's launch packages.
Is Inco Splash free?
Yes. The engine is open source under Apache 2.0, and the model weights keep their own licences.