Bonsai 2 27B: The 5.9 GB Is Not What Your GPU Needs

What Bonsai 2 27B needs on your GPU

5.9 GB is the weights. It is not what your card needs. PrismML's Bonsai 2 27B, released on 17 September, is Qwen3.8 27B compressed to ternary weights, and the launch announcement leads with the same footprint. That figure is real: the smaller of the two ternary files PrismML publishes is 5.95 GB. What it leaves out is the KV cache, the memory a model spends holding your conversation, and on this architecture the cache is what decides which card works.

Our GPU checker now carries Bonsai 2 27B, sized from PrismML's own files. Here is what it computes for llama.cpp with the default f16 cache, one card, before the 5% margin every card keeps back:

Context PTQ1_0 (5.95 GB file) PQ2_0 (7.21 GB file)
4K 6.81 GiB 8.01 GiB
32K 8.59 GiB 9.79 GiB
64K 10.63 GiB 11.83 GiB
128K 14.71 GiB 15.91 GiB
262K (full window) 22.87 GiB 24.07 GiB

Add 0.63 GB if you load the vision file for image input. Text alone never pays for it.

The weights barely move down the table. The cache does all the climbing, from 0.40 GiB at 4K to 16.15 GiB at the full window.

Why the cache is the limit, not the 5.9 GB

Bonsai 2 27B keeps Qwen3.8 27B's architecture unchanged, which PrismML states on the model card and which its published configuration confirms: 64 layers, of which 48 use linear attention and 16 use full attention.

The 48 linear layers hold a fixed state that does not grow with the conversation, about 0.15 GiB in total. The 16 full-attention layers are an ordinary KV cache: 4 key and value heads of 256 dimensions each, in f16, which works out to 64 KiB for every token you keep in context. At 32K tokens that is 2 GiB. At the full 262,144 tokens it is 16 GiB, nearly three times the weights.

PrismML's own framing gets this part right, and their demo notes state the 64 KiB per token outright. Their card says the 262K context is "kept practical by the Qwen3.8-27B hybrid-attention backbone", and it is: a model with every layer on full attention would need four times as much. Practical still means about 16 GiB, and an 8 GB card does not have it.

How much context your card gets

The same calculation turned round: the longest context that fits each card size, with the 5% margin applied, for both files and for three cache types.

Card memory PTQ1_0, f16 cache PTQ1_0, q8_0 PTQ1_0, q4_0 PQ2_0, f16 cache PQ2_0, q8_0 PQ2_0, q4_0
8 GB 16K 31K 59K does not fit does not fit does not fit
10 GB 46K 87K 165K 27K 52K 98K
12 GB 76K 143K full window 57K 108K 205K
16 GB 136K full window full window 117K 221K full window
24 GB 255K full window full window 236K full window full window
32 GB full window full window full window full window full window full window

At the default f16 cache, an 8 GB card stops at 16K. A q4_0 cache takes it to 59K, and a 12 GB card to the full 262K window. That is the setting in @DogukanUrker's published config for an RTX 3060 12GB, which sets --cache-type-k q4_0 --cache-type-v q4_0 and runs PQ2_0 at 220K, a little past the 5% margin our table keeps, and PTQ1_0 at the full window. @sudoingX reports the full 262K resident on the same card at 11.7 of 12 GB. In our calculation, PQ2_0 does not fit 8 GB with any cache once the 5% margin is kept back.

A quantized cache is not free. In our llama.cpp performance flags guide, a quantized cache cost decode speed in every GPU test we carry, and it switches flash attention on, which has its own cost. Take it when the context matters more than the speed.

The cards that gain the most have 12 and 16 GB

The fair comparison is not against an 8 GB card's usual options. It is against the model Bonsai is made from. Here is the same calculation for the regular Qwen3.8 27B builds in our calculator, which uses Unsloth's current files. The Q2_K and Q3_K_M rows are Unsloth's UD-Q2_K_XL and UD-Q3_K_XL, because Unsloth ships no plain Q2_K or Q3_K_M for this model. Our Qwen3.8 VRAM requirements page measured Unsloth's larger August files, so its limits run lower than these:

Card memory Best regular Qwen3.8 27B build that fits Bonsai 2 27B, PTQ1_0
8 GB nothing, at any quantization 16K of context
10 GB nothing 46K
12 GB Q2_K, to 18K 76K
16 GB Q3_K_M, to 28K 136K
24 GB Q6_K to 16K, Q5_K_M to 49K, or Q4_K_M to 99K 255K

That middle of the table is where this release matters. A 12 or 16 GB card could already hold Qwen3.8 27B, but only as a 2 or 3-bit build with a short leash on context. PrismML's own benchmarks put a conventional 2-bit build, IQ2_XXS at 9.4 GB, at 84.1% of the full model's score, against 98.2% for Bonsai 2 27B at 5.9 GB. The first agentic tests on real cards tell a more mixed story, covered below, and they include a better 2-bit build of the base model than IQ2_XXS.

It needs PrismML's llama.cpp build, and your GPU vendor matters

This is the part to read before you download 6 GB.

Stock llama.cpp will not run these files. PrismML's model card says so directly: it rejects PQ2_0 and PTQ1_0 as unknown types. Worse, a third format, Q2_0, which PrismML keeps in a separate testing repo, loads in stock llama.cpp without any warning and produces garbage, because stock llama.cpp cannot apply the Hadamard transform the rotated weights expect on the activations. The ternary kernels live in PrismML's fork of llama.cpp, and the card links prebuilt binaries.

PrismML's model card lists three backends: CUDA, Metal and CPU. That covers NVIDIA cards, Apple Silicon and running on the processor. The fork itself goes further than the card. Its release builds include ROCm and Windows HIP packages for AMD cards, and PrismML's format notes list PQ2_0 kernels for Metal, CUDA, ROCm/HIP and CPU. On any other backend, PrismML says, the file falls back to slow generic code, and an Intel Arc card runs through Vulkan or SYCL, so it takes that path. On AMD, @iotcoi reports about 73 tok/s on a Strix Halo through llama.cpp on ROCm, with a DFlash2 draft model and his own kernel work, so that figure is his tuned setup rather than what a stock fork build gives.

Our GPU checker follows that: it lists Bonsai 2 27B on NVIDIA and AMD cards and not on Intel ones, because "fits" would be the wrong answer for a card with no fast kernel for the format. A Radeon owner who wants to try should use the ROCm build and the PQ2_0 file.

Desktop apps are in the same position until they add PrismML's formats. A llama.cpp feature request filed on 18 September quotes Ollama 0.33.2 stopping on these files with a tensor size overflow, and a user reported the same error on 0.34.0. Check that your app lists PrismML's formats before you download. For older NVIDIA cards, @JakeKAllDay's llamAmpere runtime added Bonsai 2 27B support in v0.3.1, aimed at 8, 10 and 12 GB Ampere cards.

On a Mac, PrismML also publishes an MLX build, 8.60 GB with the vision tower. It needs the loader PrismML ships inside that repo: an ordinary MLX loader returns wrong output rather than an error, the card warns. Our Mac checker sizes the GGUF files.

PTQ1_0 or PQ2_0: which file to download

Both files hold the same ternary weights. PTQ1_0 packs them densely at 1.75 bits per weight and 5.95 GB. PQ2_0 gives each weight a 2-bit slot, 2.13 bits per weight and 7.21 GB, which costs memory and saves the arithmetic of unpacking.

On 8 GB, the choice is made for you: PTQ1_0 is the only one that fits. Above that, PrismML's own measurements split by generation:

Card PQ2_0 decode PTQ1_0 decode
RTX 5090 129.9 tok/s 120.5 tok/s
RTX PRO 6000 Blackwell 124.8 tok/s 117.9 tok/s
H100 SXM 113.9 tok/s 86.9 tok/s
RTX 6000 Ada 82.8 tok/s 90.4 tok/s
RTX 4090 81.2 tok/s 91.1 tok/s
L40S 74.4 tok/s 81.8 tok/s
A100 SXM 73.9 tok/s 54.7 tok/s
L4 29.8 tok/s 32.1 tok/s

These are PrismML's own llama-bench numbers: 128 generated tokens at depth 0, so an empty context, batch size 1, no vision tower. The rule their table gives: PTQ1_0 on Ada cards like the 4090 and on the L4, PQ2_0 on Blackwell, Hopper and Ampere datacenter cards. PQ2_0 is also faster at reading the prompt on every card they tested.

One thing the table shows that matters for anyone estimating speed: these kernels are not limited by memory bandwidth on fast cards. On an RTX 5090, 5.95 GB read at the card's 1,792 GB/s would allow roughly 300 tokens a second; PTQ1_0 measures 120.5. PrismML's card explains why: unpacking dense trits costs arithmetic, and on H100, A100 and the Blackwell cards batch-1 decode is limited by instruction throughput and kernel launch overhead, not by memory. So our speed tool does not predict Bonsai 2 27B, because its bandwidth model would promise more than twice what the card delivers. Measured runs are the guide. Testers published these in the first two days, each on their own setup:

Tester Machine File and setup Context Decode
@DogukanUrker RTX 3060 12GB PQ2_0, PrismML fork, q4_0 cache 220K about 35 tok/s
@DogukanUrker RTX 3060 12GB PTQ1_0, same 262K 26 tok/s
@filicroval DGX Spark PQ2_0, PrismML fork not stated 29.6 tok/s
@iotcoi AMD Strix Halo llama.cpp on ROCm, DFlash2 draft, own kernels not stated about 73 tok/s
@Beamsters1 M5 Max Mac MLX, MLXServe not stated about 70 to 80 tok/s
@geldeki MacBook M2 Max 32GB an MLX 4-bit build 32K about 10 tok/s

@DogukanUrker's pair shows the packing trade on one card: PQ2_0 decodes faster, PTQ1_0 holds more context. His words: max context, go PTQ1_0; speed, go PQ2_0.

How good is it: PrismML's numbers, then real work

PrismML ran 14 benchmarks in thinking mode on an H100 and reports these on the model card:

Build True bits per weight Size Average score Against full precision
Qwen3.8 27B FP16 16.0 54 GB 86.32 100%
Qwen3.8 27B UD-Q4_K_XL 5.2 17.6 GB 85.18 98.7%
Qwen3.8 27B IQ2_XXS 2.8 9.4 GB 72.59 84.1%
Bonsai 2 27B 1.72 5.9 GB 84.78 98.2%

The gap they report is concentrated where you would expect it: knowledge and reasoning, and vision. Math is 96.57 against 97.06 for full precision, and coding is level. The conventional 2-bit build fails selectively, by their account, scoring 57.5 on AIME26 against Bonsai's 95.83 while still looking fine on general knowledge questions, which is why a quick chat with it can miss the damage.

PrismML's whitepaper carries two agentic benchmarks the model card's table does not. On Terminal-Bench 2.1, Bonsai 2 27B scores 52.8 against 69.7 for the full Qwen3.8 27B, and on SWE-bench Verified 60.8 against 80.6, which the whitepaper describes as "retaining roughly three quarters of the full-precision performance on both benchmarks". On PrismML's own numbers, long multi-step agent work is where the 98.2% average does not carry over.

Testers ran it on their own machines in the first two days, with mixed results:

  • @DogukanUrker ran his own agentic coding benchmark on an RTX 3060, 20 multi-step bug fixes of about 55 turns each: Bonsai 2 27B solved 13 of 20. On the same tasks, a GSQ-RCO 2-bit build of the regular Qwen3.8 27B (8.4 GB) solved 16, the official Qwen API 16, a ByteShape 2-bit build 10 and an AtomicChat 2-bit build 7. On HumanEval+, a single-function coding test, Bonsai scored 95.7 against 92.7 for a 2-bit Qwen3.8 27B.
  • @superalesha gave it his standard agent task on an RTX 3090, a three.js game. It spent its first 32K tokens planning without writing a file, then produced a black screen and shaders that did not compile, and reported the work as verified. His simpler test, a voxel pagoda in one HTML file, took three hours.
  • @MiaAI_lab called it "a hard pass" after it could not complete what she described as a not-so-difficult HTML task.
  • @iotcoi reports using it all day for agentic coding on his Strix Halo, and @nisten replied to @superalesha that others running his prompt got better results, so outcomes vary with the harness and the setup.

The 12 GB alternative both testers ranked above it. @superalesha's pick for 12 GB cards is ISTA-DASLab's GSQ-RCO IQ2_XS of the regular Qwen3.8 27B: 8.4 GB on disk, 2.50 bits per weight in his words, 131,072 tokens of context with a q4_0 cache, and 47 tok/s at 128K on an RTX 3080 Ti. It is the same build that scored 16 of 20 in @DogukanUrker's benchmark.

What to do

  1. Check your card first. Pick it in the GPU checker at the context you actually use. If Bonsai 2 27B does not appear on an Intel card, that is the missing backend support, not the memory.
  2. On 8 GB, download PTQ1_0. That gives 16K of context at the default cache, or up to 59K with a q4_0 cache if you accept the speed cost.
  3. On 12 or 16 GB, choose by the work. Bonsai 2 27B's PTQ1_0 file holds the full window on 12 GB with a q4_0 cache. For multi-step agentic coding, test it against the GSQ-RCO 2-bit build of the regular Qwen3.8 27B, which scored higher in @DogukanUrker's agentic benchmark on a 12 GB card and which @superalesha runs at 128K on a 12 GB RTX 3080 Ti with a q4_0 cache.
  4. On 24 GB, 255K fits with an f16 cache and the full window with q8_0, and you are choosing between Bonsai and a Q4_K_M or Q5_K_M of the base model. Check the speed page for Qwen3.8 27B before you trade.
  5. Use PrismML's llama.cpp build, not your usual one, and never run the Q2_0 testing file with stock llama.cpp.

FAQ

Can Bonsai 2 27B run on an 8 GB GPU?

Yes, with the PTQ1_0 file: about 16K tokens of context with the default f16 cache, 31K with a q8_0 cache and 59K with a q4_0 cache. The 7.21 GB PQ2_0 file does not fit 8 GB.

How much VRAM does Bonsai 2 27B need?

The PTQ1_0 file is 5.95 GB, but the total depends on context: about 6.8 GiB at 4K, 8.6 GiB at 32K and 22.9 GiB at the full 262,144-token window with an f16 cache in llama.cpp. The KV cache is what grows.

Does Bonsai 2 27B work with regular llama.cpp, Ollama or LM Studio?

Not with stock llama.cpp, which rejects its formats. It needs PrismML's llama.cpp fork. Ollama failed to load it in two public reports on 18 September. Check that your app lists PrismML's PQ2_0 and PTQ1_0 formats before you download.

Can I run Bonsai 2 27B on an AMD or Intel GPU?

On AMD, yes: the fork ships ROCm builds, PrismML lists PQ2_0 kernels for ROCm/HIP, and @iotcoi runs it on a Strix Halo at about 73 tok/s with his own tuning. On Intel, not at speed: PrismML's format notes list PQ2_0 kernels for Metal, CUDA, ROCm/HIP and CPU, and an Arc card runs through Vulkan or SYCL, where the file falls back to slow generic code.

Which Bonsai 2 27B file should I download, PTQ1_0 or PQ2_0?

On 8 GB, PTQ1_0, because it is the only one that fits. Above that, PrismML's measurements favour PTQ1_0 on Ada cards such as the RTX 4090 and PQ2_0 on Blackwell, Hopper and Ampere datacenter cards.

Is Bonsai 2 27B as good as the full Qwen3.8 27B?

On PrismML's model card, 98.2% of the full model's average across 14 benchmarks, 84.78 against 86.32. On agentic work it drops further: PrismML's own whitepaper gives about three quarters of the full model on Terminal-Bench 2.1 and SWE-bench Verified, and in @DogukanUrker's agentic coding benchmark it solved 13 of 20 tasks where Qwen's official API solved 16.