RTX Spark’s Real Rival Is AMD’s 192 GB Ryzen AI Max

Two Windows machines for local models, one month

NVIDIA's RTX Spark and AMD's Ryzen AI Max+ PRO 495 both reach Windows PCs in October, and for anyone running models the difference between them is one number: how much memory the GPU gets. Both are a CPU and a GPU on one package sharing one pool of LPDDR5X memory on a 256-bit bus. RTX Spark tops out at 128 GB. The 495 goes to 192 GB, and AMD says 160 GB of it can be dedicated to graphics.

Check which RTX Spark it is before anything else. NVIDIA's spec page lists two N1X chips. The desktop, and the larger of the two laptop chips, pair a 20-core Grace CPU with a 6,144-core Blackwell GPU and take up to 128 GB. The second is for laptops only: an 18-core CPU, a 5,120-core GPU and up to 64 GB. Everything after this paragraph is the 128 GB chip. On a 64 GB machine, NVIDIA's rule gives the GPU 48 GB with any carveout up to 32 GB and 56 GB with a 48 GB one, and our engine fits 54 to 55 models there: Qwen 3.8 27B at Q8_0 with its full 256K window, or Llama 3.3 70B at Q4_K_M or Q5_K_M with 16K. The 180B class that makes the 128 GB machines interesting is out: Qwen 3.8 Flash Next does not fit, nor does gpt-oss 120B, whose GGUF files stay near 63 GB at every quantization because its experts ship in MXFP4, and Nemotron 3 Super fits only at Q2_K. On a laptop spec sheet, the GPU core count and the memory tell the two apart.

The AMD chip in this class already on sale is the Ryzen AI Max+ 395, AMD's 128 GB part, shipping in laptops and mini PCs since 2025 and the heart of AMD's own Ryzen AI Halo developer box. The 495 is its successor: same 16 Zen 5 cores and 40 graphics compute units, faster memory, half as much again of it.

NVIDIA RTX Spark N1X, 6,144-core GPU AMD Ryzen AI Max+ 395 AMD Ryzen AI Max+ PRO 495
CPU 20 Arm cores (Grace) 16 Zen 5 cores, x86 16 Zen 5 cores, x86
GPU Blackwell, 6,144 cores, 1 petaflop FP4 by NVIDIA's figure Radeon 8060S, 40 CUs Radeon 8065S, 40 CUs
Memory, most 128 GB LPDDR5X 128 GB LPDDR5x-8000 192 GB LPDDR5x-8533
Bandwidth 300 GB/s 256 GB/s up to 273 GB/s
GPU memory, by the vendor 102.4 to 112 GB on NVIDIA's rule 96 GB dedicated, 112 reachable 160 GB dedicated
Windows Windows on Arm Windows, x86 Windows, x86
Available October now October

NVIDIA publishes RTX Spark's bandwidth in its porting guide, and AMD's own FAQ gives 256 GB/s for the Ryzen AI Max+ 395. For the 495, AMD lists bus width and memory speed, so that figure is derived: 256 bits at 8,533 MT/s is 273 GB/s. That is the chip's ceiling, and a machine can run slower memory. HP's product page lists up to 8,000 MT/s for the 128 and 192 GB ZBook Ultra G3a, which works out to 256 GB/s, while Minisforum's MS-S1 MAX-P495 lists 8,533. AMD dates the 400 series to the third quarter of 2026 from HP and Lenovo; HP expects its ZBook Ultra G3a 16 with the 495 and 192 GB to ship in October, according to StorageReview, which has a pre-production unit.

How each one splits memory with the GPU

NVIDIA publishes a rule. Its porting guide describes a dedicated "carveout" for the GPU, reported by Windows as dedicated GPU memory, plus shared system memory that both sides use: "The nominal shared system memory size is the post-carveout capacity minus 16 GB. That value is clamped to a minimum of 50% and a maximum of 80% of the post-carveout capacity." The GPU budget is the two added together. On 128 GB that gives:

Carveout None 16 GB 32 GB 48 to 96 GB
GPU budget, 128 GB machine 102.4 GB 105.6 GB 108.8 GB 112.0 GB

The guide leaves the carveout to the machine and tells developers to read it at run time, so it is the number to find in a review.

AMD publishes a setting. Variable Graphics Memory reserves part of the memory as dedicated graphics memory, and AMD's own figure for a 128 GB Ryzen AI Max+ 395 on Windows is 96 GB. AMD's FAQ adds the nuance: "The highest performance can be unlocked by limiting workloads to 96GB", but the GPU set to 96 GB "can technically access a total graphics memory size of 112GB". For the 495, AMD's footnote reads "up to 160 GB dedicated graphics memory", which it says is enough for models of 300 billion parameters and more at 4-bit.

So on a 128 GB machine the two vendors meet at 112 GB by different roads. NVIDIA reaches it with a carveout of 48 GB, and any carveout up to 96 GB gives the same 112; AMD reaches it by spilling past the 96 GB it reserves. Both warn about the last stretch. NVIDIA says allocations past the carveout "are backed by smaller pages than allocations in the dedicated segment, so access can be slower", and AMD recommends staying inside the 96.

The 495 is not in that contest. 160 GB of dedicated graphics memory is 48 GB more than either 128 GB machine gives its GPU by its vendor's own figures.

On Linux, AMD's ROCm documentation takes the other route on these chips: keep the dedicated reservation small, around 0.5 GB, and let the GPU map system memory through GTT, with a kernel limit the user sets. Its worked example is a 128 GB machine set to 100 GB.

What fits on each

We ran every model we size through the same engine as our calculator: llama.cpp, one machine, 8K of context with an F16 cache, the largest quantization that fits, at each machine's GPU budget.

Machine Memory The GPU gets Bandwidth System Models that fit On sale
Mac with M5 Max 128 GB 83 to 96 GB (macOS, 65 to 75%) 614 GB/s macOS 60 to 62 now
Ryzen AI Max+ 395 128 GB 96 GB, 112 reachable 256 GB/s Windows or Linux, x86 62, or 64 at 112 now
RTX Spark N1X 128 GB 102.4 to 112 GB 300 GB/s Windows on Arm 64 October
DGX Spark 128 GB 112 to 119 GB, reported 273 GB/s DGX OS, Linux 64 now
Ryzen AI Max+ PRO 495 192 GB 160 GB up to 273 GB/s Windows or Linux, x86 70 October
Mac Studio with M5 Ultra 256 GB 166 to 192 GB (macOS, 65 to 75%) 1,200 GB/s macOS 70 to 71 now

At 128 GB the four machines give their GPUs between 83 and 119 GB, and at their best the three that are not Macs hold the same 64 models. Above 128 GB the choice is between AMD's 192 GB chip and a Mac Studio, whose 512 GB configuration Apple dates to late October.

The full table is in the paid guide. This page sizes the headline models. Issue 1 of our Local LLM Sizing Guide, out 30 September, sets every unified-memory machine above side by side against every model we size, with the quantization and the context each one holds.

The two that separate RTX Spark from the 395 at 96 GB are the 300B class at Q2_K: GLM-5.3-Flash and Hy3 need 96 to 101 GB, just over AMD's recommended line and inside even RTX Spark's smallest budget. MiMo-V2.6-Flash is not among them: its published Q2_K is a 126 GB file, too big for any 128 GB machine here. Naive-N0.5-Flash, built on MiMo's frame, has no GGUF yet, so we size it only at its BF16 release until one appears. How the headline models land, context being the longest step that fits with an F16 cache, stopped at the model's own window:

Model 96 GB: Max+ 395, or a 128 GB Mac at 75% 102.4 GB: RTX Spark, no carveout 112 GB: RTX Spark, Max+ 395 or DGX Spark 160 GB: Max+ PRO 495 192 GB: M5 Ultra 256 GB at 75%
Qwen 3.8 Flash Next 180B / 6B active Q3_K_M, 256K Q3_K_M, 256K Q4_K_M, 128K Q5_K_M, 256K Q8_0, 256K
GLM-5.3-Flash 320B / 18B active does not fit Q2_K, 64K Q2_K, 512K Q3_K_M, 1M Q4_K_M, 256K
MiMo-V2.6-Flash 309B / 15B active does not fit does not fit does not fit Q2_K, 1M native MXFP4, 1M
DeepSeek-V4-Flash 284B / 13B active Q2_K, 512K Q2_K, 1M Q2_K, 1M native FP4 and FP8, 128K native FP4 and FP8, 1M
MiniMax M2.7 229B Q3_K_M, 16K Q3_K_M, 32K Q3_K_M, 64K Q5_K_M, 16K Q6_K, 32K
Nemotron 3 Super 120B / 12B active Q4_K_M, 256K Q4_K_M, 256K Q6_K, 128K Q8_0, 256K Q8_0, 256K
Llama 3.3 70B Q8_0, 64K Q8_0, 64K Q8_0, 64K FP16, 64K FP16, 128K
gpt-oss 120B / 5.1B active its native MXFP4, full 128K the same the same the same the same

The DeepSeek row is the one to read twice. At 159.3 GiB on our engine with its cache at 8K, V4-Flash at DeepSeek's own precision (FP4 experts, FP8 for most of the rest) fits on the 495 and on a 256 GB Mac Studio; in llama.cpp that is a GGUF that keeps those weights as they are, such as Unsloth's UD-Q8_K_XL, which Unsloth reports as lossless. Every 128 GB machine here runs a community 2-bit GGUF of it instead. A DGX Spark's reported 119 GB adds context rather than quality: the same quantizations as at 112, with Qwen 3.8 Flash Next at 256K and GLM-5.3-Flash at 1M.

What 192 GB adds

Six models fit on the 495 and on none of the 128 GB machines, all at the low end of the ladder:

Model Quantization Size at 8K
MiMo-V2.6-Flash 309B / 15B active Q2_K 120.9 GiB
Llama 4 Maverick 400B / 17B active Q2_K 127.0 GiB
Llama 3.1 405B Q2_K 131.0 GiB
MiniMax M3 427B Q2_K 134.8 GiB
Qwen3 Coder 480B / 35B active Q2_K 152.4 GiB
Ornith 1.5 397B Q3_K_M 159.0 GiB

Ornith at Q3_K_M leaves 1 GiB of the 160, so treat it as the edge rather than a working setup. A Mac Studio with M5 Ultra and 256 GB holds all six as well, even at the low end of macOS's GPU share. The bigger change is up the ladder. On the 495, GLM-5.3-Flash moves from Q2_K to Q3_K_M with its full 1M window, MiMo-V2.6-Flash runs its Q2_K at 1M too (its native MXFP4 file fits only at 8K), Qwen 3.8 Flash Next goes to Q5_K_M with 256K, and a 70B runs unquantized.

Speed: a ceiling, not a measurement

Bandwidth sets the ceiling on decode speed, because every generated token reads the active weights from memory. RTX Spark's 300 GB/s is about 10% above a 495 running 8,533 MT/s memory (273 GB/s) and 17% above the 395's 256, which is also what a 495 gets at the 8,000 MT/s HP lists for its 192 GB ZBook. On the same model at the same quantization, RTX Spark's ceiling is the highest of the Windows machines and above DGX Spark's 273. The Macs sit in another range: 614 GB/s on the M5 Max and 1,200 on the M5 Ultra, two and four times RTX Spark's.

That comparison holds only while the model is the same. The 495's advantage is that it runs a bigger quantization or a bigger model, and a bigger file reads more bytes per token: Qwen 3.8 Flash Next at Q5_K_M decodes slower than at Q3_K_M on any machine. Everything in this section is a ceiling from the specifications, not a measured speed. For the 395 and every card in its list, the tokens per second checker estimates decode speed from memory bandwidth and shows the measured runs beside the estimate.

Software: Arm against x86

RTX Spark runs Windows on Arm. The model runs natively on the GPU either way, but the program around it may not: NVIDIA's guide notes that "The emulation layer translates CPU instructions only. GPU device code runs natively on the iGPU." An app without an Arm build runs its CPU side under emulation. On the GPU side it is CUDA.

The Ryzen AI Max chips run ordinary x86 Windows, so every Windows program runs as it does on any PC. AMD ran its own 96 GB demonstrations on llama.cpp over Vulkan in LM Studio, and its ROCm stack now reaches Windows too: AMD lists PyTorch on ROCm 7.2.1 for the Ryzen AI Max 300 series on Windows, and its Lemonade server recommends ROCm first on these chips, with Vulkan as the fallback. On Linux, AMD's ROCm documentation covers the 395 directly.

DGX Spark runs DGX OS, NVIDIA's Linux, with the same CUDA stack as RTX Spark and none of the Windows questions. For what a Spark decodes today, see DeepSeek V4 Flash on DGX Spark.

The Macs run macOS, where the GPU's share of memory is a system limit rather than a vendor setting: our figures take the range from 65% to about 75%, and our Mac checker lets you enter your own.

Before you buy one for models

  • At 128 GB, the GPU budget decides it, and both vendors' figures meet at 112. On RTX Spark, ask for the carveout: none gives 102.4 GB, and anything from 48 to 96 GB gives 112. On a Ryzen AI Max+ 395, set Variable Graphics Memory to 96 GB and treat the last 16 as spill.
  • If the models you want are 300B-class, the 495 is the machine. At 160 GB it moves them off Q2_K, and it is the only one of the three that fits a 400B model at all.
  • Check your app has a native Arm build before buying RTX Spark. Emulation does not touch the GPU, but a slow front end is still slow.
  • Leave headroom. Every fit above is sized to the full budget, so read each one as the edge, not the target. NVIDIA's own guide warns that allocating the whole budget "can leave too little host memory and can make the system unresponsive."

FAQ

Is there more than one RTX Spark?

Yes. NVIDIA lists two N1X chips. The desktop and the larger laptop chip have a 20-core CPU, a 6,144-core GPU and up to 128 GB. A laptop-only chip has an 18-core CPU, a 5,120-core GPU and up to 64 GB, which by NVIDIA's rule gives its GPU 48 to 56 GB.

How much memory does RTX Spark give the GPU?

By NVIDIA's porting guide, the GPU gets a dedicated carveout plus a shared region equal to the rest minus 16 GB, clamped between 50% and 80% of the rest. On a 128 GB machine that is 102.4 GB with no carveout, rising to 112 GB with a carveout of 48 to 96 GB.

How much VRAM does the Ryzen AI Max+ 395 have?

AMD's figure for a 128 GB Ryzen AI Max+ 395 on Windows is 96 GB of dedicated graphics memory through Variable Graphics Memory. AMD says the GPU can technically reach 112 GB in total, with the best performance inside the 96.

How much VRAM does the Ryzen AI Max+ PRO 495 have?

AMD says the Ryzen AI Max+ PRO 495 supports up to 160 GB of dedicated graphics memory on a 192 GB machine. On our engine that fits 70 models at 8K of context, against 64 on a 128 GB RTX Spark.

How does RTX Spark compare with DGX Spark and a Mac?

A DGX Spark gives its GPU 112 to 119 GB on Linux and holds the same 64 models as RTX Spark on our engine. A 128 GB Mac with M5 Max gives 83 to 96 GB and holds 60 to 62, with twice the bandwidth. A Mac Studio with M5 Ultra and 256 GB gives 166 to 192 GB and holds 70 to 71, at 1,200 GB/s.

Is RTX Spark faster than Ryzen AI Max?

Its memory bandwidth is higher: 300 GB/s against up to 273 on the Max+ PRO 495 and 256 on the Max+ 395, so on the same model and quantization its decode ceiling is the highest. That is a ceiling from the specifications, not a measured speed.

When do RTX Spark and the Ryzen AI Max+ PRO 495 ship?

NVIDIA says RTX Spark Windows PCs arrive in October 2026. AMD dates the Ryzen AI Max PRO 400 series to the third quarter of 2026 from HP and Lenovo, and HP expects its ZBook Ultra G3a 16 with the 495 to ship in October.