M5 Ultra Mac Studio: What 512 GB at 1.2 TB/s Runs
Apple announced the M5 Ultra Mac Studio and the M6 Mac mini today, with pre-orders open now and machines arriving 22 September. For local models the number that matters is not the 512 GB. It is the 1.2 TB/s.
Here is why, and it is the part the launch coverage will not mention:
| Chip | Memory bandwidth |
|---|---|
| M1 Ultra | 800 GB/s |
| M2 Ultra | 800 GB/s |
| M3 Ultra | 819 GB/s |
| M5 Ultra | 1,200 GB/s |
There is no M4 Ultra in that list because Apple never built one: the M4 Max shipped without an UltraFusion connector, so the 2025 Mac Studio paired an M4 Max against a year-old M3 Ultra.
The Ultra line sat still for three generations and just moved 47%. Apple’s own figure is 50 percent, which is the jump measured against the round 800 GB/s of the M1 and M2 Ultra. Against the M3 Ultra’s actual 819 GB/s the number is 47. Single-stream decode is bandwidth-bound, which means tokens per second on a Mac tracks that column almost directly. This is the first Apple release in four years that makes local models meaningfully faster rather than just larger.
Below: every new machine with its real numbers, what our calculator says actually fits, the model that does not fit in 512 GB, and the figure Apple has never published that decides all of it.
Every Machine Apple Launched Today
From Apple’s own spec pages, not the announcement posts, and that distinction earns its place three lines down. One timing note before the table: the 512 GB configuration is the one machine here you cannot order today. Apple has it arriving in late October, while everything else pre-orders now and lands 22 September.
| Machine | Chip | Memory | Bandwidth |
|---|---|---|---|
| Mac mini | M6, 12-core CPU and GPU | 16 GB | 153 GB/s |
| Mac mini | M6, 12-core CPU and GPU | 24 or 32 GB | 170 GB/s |
| Mac mini | M5 Pro, 15-core CPU, 16-core GPU | 24, 48, 64 GB | 307 GB/s |
| Mac Studio | M5 Max, 32-core GPU | 36 GB | 460 GB/s |
| Mac Studio | M5 Max, 40-core GPU | 48, 64, 128 GB | 614 GB/s |
| Mac Studio | M5 Ultra, 64-core GPU | 96 GB | 1.2 TB/s |
| Mac Studio | M5 Ultra, 80-core GPU | 256 or 512 GB | 1.2 TB/s |
Note the first two rows. The M6 is being reported everywhere as a 170 GB/s chip. Apple lists 153 GB/s at 16 GB, with 170 appearing only on the 24 and 32 GB builds. Bandwidth moves with capacity on this chip, which is not how a spec sheet usually reads, and on a machine bought for local models it is an 11% difference in decode speed for the price of a memory upgrade.
The other thing worth seeing in one table: the mini lineup now spans a 2x bandwidth gap, 153 to 307 GB/s, between two machines that look nearly identical on a shelf.
What Actually Fits
Every figure below is our own calculator, MLX backend, 32,000 token context, and it is capacity rather than measured speed. The pool is not the memory. macOS hands the GPU a fraction of unified memory, so a 512 GB machine does not offer 512 GB to a model. More on that at the end, because it is the caveat that decides everything else.
| Model at Q4 | Needs | M6 32 GB | M5 Max 128 GB | M5 Ultra 256 GB | M5 Ultra 512 GB |
|---|---|---|---|---|---|
| Qwen3.8 27B | 19.1 GiB | tight | yes | yes | yes |
| Qwen3.8 27B at Q8 | 30.2 GiB | no | yes | yes | yes |
| DeepSeek-V4-Flash 284B | 156.5 GiB | no | no | yes | yes |
| Llama 4 Maverick 400B | 232.8 GiB | no | no | no | yes |
| GLM-5.2 753B | 429.1 GiB | no | no | no | no |
| Kimi K2.6 1T | 567.8 GiB | no | no | no | no |
Two rows in that table are the whole article.
The 512 GB machine does not run GLM-5.2 at Q4. It needs 429.1 GiB and the pool on a 512 GB Mac is roughly 333 to 384 GiB. “Half a terabyte runs anything” is the obvious assumption and it is wrong at the top of the range. You would need a lower quantization tier, and at that size that is a real quality decision rather than a free one.
The 32 GB M6 mini runs Qwen3.8 27B at Q4, and only just. 19.1 GiB against a pool of about 21 to 24 GiB. That is the cheapest new Mac Apple sells running the model this site’s readers ask about more than any other. It is tight, it depends on where your machine’s limit actually lands, and it works.
The Upgrade Question, Answered Honestly
If you own an M3 Ultra with 512 GB, the new machine holds exactly the same models. Same capacity, same pool, same fit table. Nothing that failed before will start working.
What you get is 819 GB/s becoming 1,200. On Qwen3.8 27B at Q4, whose published flat Q4_K_M file is 17.77 GB, that moves the arithmetic ceiling from roughly 46 tok/s to roughly 68. Call those ceilings rather than measurements: they are bandwidth divided by weights, they ignore every real overhead, and nobody has run a model on an M5 Ultra yet, because outside Apple nobody has one until 22 September.
So the upgrade is not about what you can run. It is about whether a 47% speed increase on what you already run is worth the price of a Mac Studio. For an agentic workload where you sit and wait on decode, that is a real argument. For occasional use it is not.
Going the other way is more interesting. If you are choosing between a 256 GB M5 Ultra and a 512 GB one, capacity is the only difference: both run at 1.2 TB/s. The 256 GB machine covers everything up to the 284B class. The 512 GB machine adds the 400B class and stops short of the 753B class anyway.
The Number Apple Has Never Published
Everything above depends on one figure that Apple does not document: how much of unified memory the GPU is actually allowed to use.
macOS reserves a share for the system. In practice the GPU is offered somewhere around 75% of the pool, and reported values run from about 65% to 75% depending on the machine. Metal exposes it as recommendedMaxWorkingSetSize, no browser can read it, and Apple publishes no formula.
That is why the fit table above is built on a band rather than a number, and why our Mac Checker asks you to paste what your own runtime prints. On llama.cpp it is the ggml_metal_init: recommendedMaxWorkingSetSize line. On MLX it is one call to mx.metal.device_info().
On a 512 GB machine that band is 64 GiB wide. It is the difference between Llama 4 Maverick fitting comfortably and fitting at all, and it is the single number worth checking before spending five figures.
Which One To Buy
- Running 27B class models and nothing bigger: the M6 mini at 32 GB does it, tightly. The M5 Pro mini at 48 GB does it with room and twice the bandwidth, and for local models that is the better mini.
- Running up to the 284B class: M5 Ultra at 256 GB. Same speed as the 512, and everything below 200 GiB fits.
- Running the 400B class: 512 GB, and it is the only machine here that does. It is also the one you have to wait for, since Apple has that configuration arriving in late October.
- Running the frontier 700B and 1T class at Q4: none of them. That tier needs a lower quantization tier on a Mac, or an expert-offload engine on a PC where system RAM does the holding. If that is the direction you are heading, size the card first with the GPU checker.
- Already own an M3 Ultra: you are buying speed, not capability. 47% more bandwidth, which is a ceiling on decode rather than a measured gain.
Sources
- Apple, Mac Studio technical specifications and Mac mini technical specifications, read 2026-08-25. Every chip, memory and bandwidth figure comes from these
- Fit figures are our own VRAM calculator, MLX backend at 32,000 token context, and are capacity rather than measured speed
- Decode ceilings are bandwidth divided by weight size, labelled as ceilings throughout. We have not run a model on any of these machines, and as of 2026-08-25 nobody has published a measured tokens per second figure for the M5 Ultra, since the hardware does not reach customers until 22 September
FAQ
How much bandwidth does the M5 Ultra Mac Studio have?
1.2 TB/s, on both the 64-core and 80-core GPU configurations. That is up from 819 GB/s on the M3 Ultra and 800 GB/s on the M2 and M1 Ultra, so it is the first significant increase the Ultra line has had in four years.
Can a 512 GB Mac Studio run any local model?
No. At Q4, GLM-5.2 at 753B needs about 429 GiB and Kimi K2.6 at 1T needs about 568, while the usable pool on a 512 GB machine is roughly 333 to 384 GiB. The 400B class fits. The 700B class and above does not, at that quantization tier.
Is the M6 Mac mini 153 or 170 GB/s?
Both, depending on memory. Apple lists 153 GB/s on the 16 GB configuration and 170 GB/s on the 24 and 32 GB builds. Bandwidth moves with capacity on this chip.
Can the M6 Mac mini run Qwen3.8 27B?
At Q4 and 32 GB of memory, yes, with very little to spare: about 19.1 GiB against a usable pool of roughly 21 to 24 GiB. At Q8 it does not, needing about 30.2 GiB.
Is the M5 Ultra worth upgrading to from an M3 Ultra?
It runs the same models, because the memory is the same. What changes is speed, by about 47% on the arithmetic. Whether that is worth a new machine depends on whether you spend your day waiting on decode.
Why does my Mac not give the GPU all its memory?
macOS reserves a share for the system and offers the GPU roughly 65% to 75% of unified memory, varying by machine. Apple publishes no formula for it. Our Mac Checker takes a paste of what your own runtime reports so you can size against your real limit.