What Bonsai 2 27B needs on your GPU
5.9 GB is the weights. It is not what your card needs. PrismML's Bonsai 2 27B, released on 17 September, is Qwen3.8 27B compressed to ternary weights, and the launch announcement leads with the same footprint. That figure is real: the smaller of the two ternary files PrismML publishes is 5.95 GB. What it leaves out is the KV cache, the memory a model spends holding your conversation, and on this architecture the cache is what decides which card works.
Our GPU checker now carries Bonsai 2 27B, sized from PrismML's own files. Here is what it computes for llama.cpp with the default f16 cache, one card, before the 5% margin every card keeps back:
| Context | PTQ1_0 (5.95 GB file) | PQ2_0 (7.21 GB file) |
|---|---|---|
| 4K | 6.81 GiB | 8.01 GiB |
| 32K | 8.59 GiB | 9.79 GiB |
| 64K | 10.63 GiB | 11.83 GiB |
| 128K | 14.71 GiB | 15.91 GiB |
| 262K (full window) | 22.87 GiB | 24.07 GiB |
Add 0.63 GB if you load the vision file for image input. Text alone never pays for it.
The weights barely move down the table. The cache does all the climbing, from 0.40 GiB at 4K to 16.15 GiB at the full window.
Why the cache is the limit, not the 5.9 GB
Bonsai 2 27B keeps Qwen3.8 27B's architecture unchanged, which PrismML states on the model card and which its published configuration confirms: 64 layers, of which 48 use linear attention and 16 use full attention.
The 48 linear layers hold a fixed state that does not grow with the conversation, about 0.15 GiB in total. The 16 full-attention layers are an ordinary KV cache: 4 key and value heads of 256 dimensions each, in f16, which works out to 64 KiB for every token you keep in context. At 32K tokens that is 2 GiB. At the full 262,144 tokens it is 16 GiB, nearly three times the weights.
PrismML's own framing gets this part right, and their demo notes state the 64 KiB per token outright. Their card says the 262K context is "kept practical by the Qwen3.8-27B hybrid-attention backbone", and it is: a model with every layer on full attention would need four times as much. Practical still means about 16 GiB, and an 8 GB card does not have it.
How much context your card gets
The same calculation turned round: the longest context that fits each card size, with the 5% margin applied, for both files and for three cache types.
| Card memory | PTQ1_0, f16 cache | PTQ1_0, q8_0 | PTQ1_0, q4_0 | PQ2_0, f16 cache | PQ2_0, q8_0 | PQ2_0, q4_0 |
|---|---|---|---|---|---|---|
| 8 GB | 16K | 31K | 59K | does not fit | does not fit | does not fit |
| 10 GB | 46K | 87K | 165K | 27K | 52K | 98K |
| 12 GB | 76K | 143K | full window | 57K | 108K | 205K |
| 16 GB | 136K | full window | full window | 117K | 221K | full window |
| 24 GB | 255K | full window | full window | 236K | full window | full window |
| 32 GB | full window | full window | full window | full window | full window | full window |
At the default f16 cache, an 8 GB card stops at 16K. A q4_0 cache takes it to 59K, and a 12 GB card to the full 262K window. That is the setting in @DogukanUrker's published config for an RTX 3060 12GB, which sets --cache-type-k q4_0 --cache-type-v q4_0 and runs PQ2_0 at 220K, a little past the 5% margin our table keeps, and PTQ1_0 at the full window. @sudoingX reports the full 262K resident on the same card at 11.7 of 12 GB. In our calculation, PQ2_0 does not fit 8 GB with any cache once the 5% margin is kept back.
A quantized cache is not free. In our llama.cpp performance flags guide, a quantized cache cost decode speed in every GPU test we carry, and it switches flash attention on, which has its own cost. Take it when the context matters more than the speed.
The cards that gain the most have 12 and 16 GB
The fair comparison is not against an 8 GB card's usual options. It is against the model Bonsai is made from. Here is the same calculation for the regular Qwen3.8 27B builds in our calculator, which uses Unsloth's current files. The Q2_K and Q3_K_M rows are Unsloth's UD-Q2_K_XL and UD-Q3_K_XL, because Unsloth ships no plain Q2_K or Q3_K_M for this model. Our Qwen3.8 VRAM requirements page measured Unsloth's larger August files, so its limits run lower than these:
| Card memory | Best regular Qwen3.8 27B build that fits | Bonsai 2 27B, PTQ1_0 |
|---|---|---|
| 8 GB | nothing, at any quantization | 16K of context |
| 10 GB | nothing | 46K |
| 12 GB | Q2_K, to 18K | 76K |
| 16 GB | Q3_K_M, to 28K | 136K |
| 24 GB | Q6_K to 16K, Q5_K_M to 49K, or Q4_K_M to 99K | 255K |
That middle of the table is where this release matters. A 12 or 16 GB card could already hold Qwen3.8 27B, but only as a 2 or 3-bit build with a short leash on context. PrismML's own benchmarks put a conventional 2-bit build, IQ2_XXS at 9.4 GB, at 84.1% of the full model's score, against 98.2% for Bonsai 2 27B at 5.9 GB. The first agentic tests on real cards tell a more mixed story, covered below, and they include a better 2-bit build of the base model than IQ2_XXS.
It needs PrismML's llama.cpp build, and your GPU vendor matters
This is the part to read before you download 6 GB.
Stock llama.cpp will not run these files. PrismML's model card says so directly: it rejects PQ2_0 and PTQ1_0 as unknown types. Worse, a third format, Q2_0, which PrismML keeps in a separate testing repo, loads in stock llama.cpp without any warning and produces garbage, because stock llama.cpp cannot apply the Hadamard transform the rotated weights expect on the activations. The ternary kernels live in PrismML's fork of llama.cpp, and the card links prebuilt binaries.
PrismML's model card lists three backends: CUDA, Metal and CPU. That covers NVIDIA cards, Apple Silicon and running on the processor. The fork itself goes further than the card. Its release builds include ROCm and Windows HIP packages for AMD cards, and PrismML's format notes list PQ2_0 kernels for Metal, CUDA, ROCm/HIP and CPU. On any other backend, PrismML says, the file falls back to slow generic code, and an Intel Arc card runs through Vulkan or SYCL, so it takes that path. On AMD, @iotcoi reports about 73 tok/s on a Strix Halo through llama.cpp on ROCm, with a DFlash2 draft model and his own kernel work, so that figure is his tuned setup rather than what a stock fork build gives.
Our GPU checker follows that: it lists Bonsai 2 27B on NVIDIA and AMD cards and not on Intel ones, because "fits" would be the wrong answer for a card with no fast kernel for the format. A Radeon owner who wants to try should use the ROCm build and the PQ2_0 file.
Desktop apps are in the same position until they add PrismML's formats. A llama.cpp feature request filed on 18 September quotes Ollama 0.33.2 stopping on these files with a tensor size overflow, and a user reported the same error on 0.34.0. Check that your app lists PrismML's formats before you download. For older NVIDIA cards, @JakeKAllDay's llamAmpere runtime added Bonsai 2 27B support in v0.3.1, aimed at 8, 10 and 12 GB Ampere cards.
On a Mac, PrismML also publishes an MLX build, 8.60 GB with the vision tower. It needs the loader PrismML ships inside that repo: an ordinary MLX loader returns wrong output rather than an error, the card warns. Our Mac checker sizes the GGUF files.
PTQ1_0 or PQ2_0: which file to download
Both files hold the same ternary weights. PTQ1_0 packs them densely at 1.75 bits per weight and 5.95 GB. PQ2_0 gives each weight a 2-bit slot, 2.13 bits per weight and 7.21 GB, which costs memory and saves the arithmetic of unpacking.
On 8 GB, the choice is made for you: PTQ1_0 is the only one that fits. Above that, PrismML's own measurements split by generation:
| Card | PQ2_0 decode | PTQ1_0 decode |
|---|---|---|
| RTX 5090 | 129.9 tok/s | 120.5 tok/s |
| RTX PRO 6000 Blackwell | 124.8 tok/s | 117.9 tok/s |
| H100 SXM | 113.9 tok/s | 86.9 tok/s |
| RTX 6000 Ada | 82.8 tok/s | 90.4 tok/s |
| RTX 4090 | 81.2 tok/s | 91.1 tok/s |
| L40S | 74.4 tok/s | 81.8 tok/s |
| A100 SXM | 73.9 tok/s | 54.7 tok/s |
| L4 | 29.8 tok/s | 32.1 tok/s |
These are PrismML's own llama-bench numbers: 128 generated tokens at depth 0, so an empty context, batch size 1, no vision tower. The rule their table gives: PTQ1_0 on Ada cards like the 4090 and on the L4, PQ2_0 on Blackwell, Hopper and Ampere datacenter cards. PQ2_0 is also faster at reading the prompt on every card they tested.
One thing the table shows that matters for anyone estimating speed: these kernels are not limited by memory bandwidth on fast cards. On an RTX 5090, 5.95 GB read at the card's 1,792 GB/s would allow roughly 300 tokens a second; PTQ1_0 measures 120.5. PrismML's card explains why: unpacking dense trits costs arithmetic, and on H100, A100 and the Blackwell cards batch-1 decode is limited by instruction throughput and kernel launch overhead, not by memory. So our speed tool does not predict Bonsai 2 27B, because its bandwidth model would promise more than twice what the card delivers. Measured runs are the guide. Testers published these in the first two days, each on their own setup:
| Tester | Machine | File and setup | Context | Decode |
|---|---|---|---|---|
| @DogukanUrker | RTX 3060 12GB | PQ2_0, PrismML fork, q4_0 cache | 220K | about 35 tok/s |
| @DogukanUrker | RTX 3060 12GB | PTQ1_0, same | 262K | 26 tok/s |
| @filicroval | DGX Spark | PQ2_0, PrismML fork | not stated | 29.6 tok/s |
| @iotcoi | AMD Strix Halo | llama.cpp on ROCm, DFlash2 draft, own kernels | not stated | about 73 tok/s |
| @Beamsters1 | M5 Max Mac | MLX, MLXServe | not stated | about 70 to 80 tok/s |
| @geldeki | MacBook M2 Max 32GB | an MLX 4-bit build | 32K | about 10 tok/s |
@DogukanUrker's pair shows the packing trade on one card: PQ2_0 decodes faster, PTQ1_0 holds more context. His words: max context, go PTQ1_0; speed, go PQ2_0.
How good is it: PrismML's numbers, then real work
PrismML ran 14 benchmarks in thinking mode on an H100 and reports these on the model card:
| Build | True bits per weight | Size | Average score | Against full precision |
|---|---|---|---|---|
| Qwen3.8 27B FP16 | 16.0 | 54 GB | 86.32 | 100% |
| Qwen3.8 27B UD-Q4_K_XL | 5.2 | 17.6 GB | 85.18 | 98.7% |
| Qwen3.8 27B IQ2_XXS | 2.8 | 9.4 GB | 72.59 | 84.1% |
| Bonsai 2 27B | 1.72 | 5.9 GB | 84.78 | 98.2% |
The gap they report is concentrated where you would expect it: knowledge and reasoning, and vision. Math is 96.57 against 97.06 for full precision, and coding is level. The conventional 2-bit build fails selectively, by their account, scoring 57.5 on AIME26 against Bonsai's 95.83 while still looking fine on general knowledge questions, which is why a quick chat with it can miss the damage.
PrismML's whitepaper carries two agentic benchmarks the model card's table does not. On Terminal-Bench 2.1, Bonsai 2 27B scores 52.8 against 69.7 for the full Qwen3.8 27B, and on SWE-bench Verified 60.8 against 80.6, which the whitepaper describes as "retaining roughly three quarters of the full-precision performance on both benchmarks". On PrismML's own numbers, long multi-step agent work is where the 98.2% average does not carry over.
Testers ran it on their own machines in the first two days, with mixed results:
- @DogukanUrker ran his own agentic coding benchmark on an RTX 3060, 20 multi-step bug fixes of about 55 turns each: Bonsai 2 27B solved 13 of 20. On the same tasks, a GSQ-RCO 2-bit build of the regular Qwen3.8 27B (8.4 GB) solved 16, the official Qwen API 16, a ByteShape 2-bit build 10 and an AtomicChat 2-bit build 7. On HumanEval+, a single-function coding test, Bonsai scored 95.7 against 92.7 for a 2-bit Qwen3.8 27B.
- @superalesha gave it his standard agent task on an RTX 3090, a three.js game. It spent its first 32K tokens planning without writing a file, then produced a black screen and shaders that did not compile, and reported the work as verified. His simpler test, a voxel pagoda in one HTML file, took three hours.
- @MiaAI_lab called it "a hard pass" after it could not complete what she described as a not-so-difficult HTML task.
- @iotcoi reports using it all day for agentic coding on his Strix Halo, and @nisten replied to @superalesha that others running his prompt got better results, so outcomes vary with the harness and the setup.
The 12 GB alternative both testers ranked above it. @superalesha's pick for 12 GB cards is ISTA-DASLab's GSQ-RCO IQ2_XS of the regular Qwen3.8 27B: 8.4 GB on disk, 2.50 bits per weight in his words, 131,072 tokens of context with a q4_0 cache, and 47 tok/s at 128K on an RTX 3080 Ti. It is the same build that scored 16 of 20 in @DogukanUrker's benchmark.
What to do
- Check your card first. Pick it in the GPU checker at the context you actually use. If Bonsai 2 27B does not appear on an Intel card, that is the missing backend support, not the memory.
- On 8 GB, download PTQ1_0. That gives 16K of context at the default cache, or up to 59K with a q4_0 cache if you accept the speed cost.
- On 12 or 16 GB, choose by the work. Bonsai 2 27B's PTQ1_0 file holds the full window on 12 GB with a q4_0 cache. For multi-step agentic coding, test it against the GSQ-RCO 2-bit build of the regular Qwen3.8 27B, which scored higher in @DogukanUrker's agentic benchmark on a 12 GB card and which @superalesha runs at 128K on a 12 GB RTX 3080 Ti with a q4_0 cache.
- On 24 GB, 255K fits with an f16 cache and the full window with q8_0, and you are choosing between Bonsai and a Q4_K_M or Q5_K_M of the base model. Check the speed page for Qwen3.8 27B before you trade.
- Use PrismML's llama.cpp build, not your usual one, and never run the Q2_0 testing file with stock llama.cpp.
FAQ
Can Bonsai 2 27B run on an 8 GB GPU?
Yes, with the PTQ1_0 file: about 16K tokens of context with the default f16 cache, 31K with a q8_0 cache and 59K with a q4_0 cache. The 7.21 GB PQ2_0 file does not fit 8 GB.
How much VRAM does Bonsai 2 27B need?
The PTQ1_0 file is 5.95 GB, but the total depends on context: about 6.8 GiB at 4K, 8.6 GiB at 32K and 22.9 GiB at the full 262,144-token window with an f16 cache in llama.cpp. The KV cache is what grows.
Does Bonsai 2 27B work with regular llama.cpp, Ollama or LM Studio?
Not with stock llama.cpp, which rejects its formats. It needs PrismML's llama.cpp fork. Ollama failed to load it in two public reports on 18 September. Check that your app lists PrismML's PQ2_0 and PTQ1_0 formats before you download.
Can I run Bonsai 2 27B on an AMD or Intel GPU?
On AMD, yes: the fork ships ROCm builds, PrismML lists PQ2_0 kernels for ROCm/HIP, and @iotcoi runs it on a Strix Halo at about 73 tok/s with his own tuning. On Intel, not at speed: PrismML's format notes list PQ2_0 kernels for Metal, CUDA, ROCm/HIP and CPU, and an Arc card runs through Vulkan or SYCL, where the file falls back to slow generic code.
Which Bonsai 2 27B file should I download, PTQ1_0 or PQ2_0?
On 8 GB, PTQ1_0, because it is the only one that fits. Above that, PrismML's measurements favour PTQ1_0 on Ada cards such as the RTX 4090 and PQ2_0 on Blackwell, Hopper and Ampere datacenter cards.
Is Bonsai 2 27B as good as the full Qwen3.8 27B?
On PrismML's model card, 98.2% of the full model's average across 14 benchmarks, 84.78 against 86.32. On agentic work it drops further: PrismML's own whitepaper gives about three quarters of the full model on Terminal-Bench 2.1 and SWE-bench Verified, and in @DogukanUrker's agentic coding benchmark it solved 13 of 20 tasks where Qwen's official API solved 16.