Is the NVIDIA DGX Spark 64GB worth $4,999? The local AI box and a cloud GPU are not billed the same way

Hardware  ·  2026.10.08  ·  About 12 min read

A small desktop computer on a desk, standing in for a local AI workstation beside remote GPU rental

A local AI box beside the monitor and a cloud GPU billed by the hour do not spend the same kind of money. The sections below separate this 64GB machine’s memory, bandwidth, and software stack, then put three years of use — four hours a day, eight hours a day, and always on — next to RTX 4090, L40S, A100, and H100 rental. Which models fit, where generation speed stalls, and where a Mac workstation still belongs come after that.

$4,999
64GB starting price
64 GB
Unified memory
273
Bandwidth, GB/s

What the 64GB configuration actually is

In October 2026 NVIDIA added a 64GB DGX Spark. The chip is still the GB10 Grace Blackwell superchip: a 20-core Arm CPU (10 Cortex-X925 plus 10 Cortex-A725, co-designed with MediaTek), a Blackwell GPU, fifth-generation Tensor Cores, and an onboard ConnectX-7 at 200 Gbps. System memory is 64GB of coherent LPDDR5X shared by CPU and GPU, at 273 GB/s. The chassis supply is 240W, GB10 TDP is listed at 140W, the box is about 150 × 150 × 50.5 mm, and it weighs 1.2 kg. The OS is DGX OS with NVIDIA’s CUDA stack. It is not a Windows desktop and it is not macOS.

From 23 October 2026 the configuration is sold by Acer, ASUS, Dell, Gigabyte, HP, and MSI. NVIDIA’s starting price is $4,999. There is no NVIDIA-branded 64GB Founders unit. The 128GB model remains on sale; The Register reported on 2 October 2026 that its price moved to $6,950, well above the original $3,999 launch tag. Several reports also say storage on the 64GB model is cut in half. Drive count and capacity follow each OEM’s sheet. “Up to 4TB” is a family maximum, not a promise that every unit ships full.

Machine Memory Bandwidth What you are actually buying
DGX Spark 64GB 64GB unified 273 GB/s A large model that stays on the desk, with data on the desk
RTX 4090 cloud 24GB about 1 TB/s Models that fit in 24GB, switched on by the hour
L40S 48GB about 864 GB/s One step above a 4090, still hourly
A100 80GB / H100 80GB 80GB HBM about 2 TB/s / about 3.35 TB/s Training, concurrency, and larger models at higher precision

1 PFLOP is not how you compare it with an H100

NVIDIA lists up to 1 petaFLOP of FP4 AI performance. That peak is useful inside the Spark family. It is not H100 BF16 training throughput. Single-stream speed is the bandwidth story in the next section.

Which models actually fit in 64GB

NVIDIA’s claim is that one 64GB unit runs models up to about 100 billion parameters locally, for inference, fine-tuning, and agents, and that two units linked over ConnectX-7 pool 128GB and reach about 200 billion. Read that next to precision. Weight size is roughly parameter count times bytes per parameter. GGUF Q4 is about 0.55–0.6 bytes, FP8 or Q8 about 1 byte, BF16 about 2 bytes. Leave another 8–12GB for the OS, the runtime, and KV cache. The table is a capacity check, not a measured tokens-per-second score from one repo.

Model class Weights, roughly Fits in 64GB? The practical edge
8B, Q4 about 5GB Easily Room for a long context and tool calls
32B, Q4 about 18GB Easily The usual local coding-assistant size
32B, Q8 / FP8 about 32–35GB Yes You have to cap context length yourself
70B, Q4 about 38–42GB Yes, tight A longer context exhausts KV before it exhausts the weights
70B, FP8 about 70GB No 128GB Spark, or an 80GB cloud GPU
About 100B, Q4 about 55GB Against the stated ceiling Almost no KV budget. Not a daily driver
405B at a usable precision Far past 64GB No A cluster or the cloud, not this box

Fine-tuning spends memory faster than inference. QLoRA from 8B to 14B is the comfortable band. 32B QLoRA is possible with a very small batch. Full fine-tuning or a high-rank adapter at 70B does not fit in 64GB. NVIDIA does list fine-tuning as a 64GB workload, meaning the models that already fit, not “train a 70B however you like.”

Measure the weights before the price

Size the model you actually want resident, once at Q4 and once at FP8. If the answer lands between 20GB and 45GB, this machine is in the conversation. Eight gigabytes, or 70B at FP8, moves you to a different column in the cost table.

Generation speed hits bandwidth first

With a batch of one, each new token roughly reads the weights once. A ceiling is memory bandwidth divided by weight size. It is a ceiling, not a benchmark: kernel overhead, the KV cache, and the CPU all land below it.

  • 8B Q4, about 5GB — 273 ÷ 5 ≈ 55 tokens/s of bandwidth headroom
  • 32B Q4, about 18GB — 273 ÷ 18 ≈ 15 tokens/s
  • 70B Q4, about 40GB — 273 ÷ 40 ≈ 6.8 tokens/s

The same sketch on an H100 SXM at about 3.35 TB/s puts a 40GB weight near an 80 token/s ceiling. A 4090 is around 1 TB/s, and its 24GB cannot hold a 70B Q4. That gap is the Spark’s job: the model is larger than a consumer GPU, and you do not want to pay H100 rates for the throughput. You are buying “it fits, and the data stays here,” not “it matches a data-center GPU.”

Three years: a purchase and an hourly meter are different units

The $4,999 tag is only the first line. Electricity here assumes 160W while a model is generating and $0.15 per kWh. 160W sits between the 140W chip TDP and the 240W supply. It is a planning assumption, not a measured power curve. Over 365.25 days, three years is 4,383 hours at four hours a day, 8,766 hours at eight hours a day, and 26,298 hours if the machine never sleeps.

Cloud rates are planning figures taken from RunPod’s public on-demand sheet as it still read in late September 2026, not marketplace floors. RTX 4090 at $0.50/hour sits between community cloud at $0.34 and secure cloud at $0.74. L40S at $0.90 sits between $0.79 and $1.09. A100 80GB is $1.39. H100 SXM is the community-cloud $2.69. Storage, egress, and queue time are not in the table. Those rates move every week. Recalculate with the price on the day you order.

Three-year spend (drop in your own hours)
# hours: time the GPU is actually generating, across three years
hours = hours_per_day * 365.25 * 3
cloud = hours * hourly_usd
# 0.160 kW and $0.15/kWh are assumptions
spark = 4999 + (0.160 * hours * 0.15)
Three-year pattern Spark 64GB 4090 · $0.50 L40S · $0.90 A100 · $1.39 H100 · $2.69
4 hours a day $5,104 $2,192 $3,945 $6,092 $11,790
8 hours a day $5,209 $4,383 $7,889 $12,185 $23,581
Always on $5,630 $13,149 $23,668 $36,554 $70,742

Divide $4,999 by the hourly rate and the rough crossover is about nine hours a day against a 4090, five against an L40S, 3.3 against an A100, and 1.7 against an H100. Power and whatever the box is worth after three years move that line. There is no resale forecast here: the machine still exists, the rental hours do not. If a later sale returned $1,500, sunk hardware would be $3,499 and the crossover would arrive sooner. That is a sensitivity, not a promised recovery price.

64GB, 128GB, and two boxes on a cable

The 64GB and 128GB models share the chip, the bandwidth, DGX OS, and ConnectX-7. Capacity differs, and so does the storage that reports describe as halved. At $6,950, the extra $1,951 on the 128GB model buys room for 70B at FP8, KV headroom on a Q4 70B, and a larger fine-tune batch. If your weight table is already near 55GB, do not buy 64GB and hope a harsher quant will save you later.

Two 64GB units link with a QSFP cable and no switch. NVIDIA’s cluster assistant handles discovery and the ConnectX-7 setup. The pool is 128GB, compute and bandwidth roughly double, and the starting bill is $9,998. Next to one 128GB system at $6,950, a pair of the smaller machines is not how you save money. It is how you add a second unit after a single 32B–70B Q4 box has earned it, not how you assemble 128GB on day one.

When the $4,999 is the right spend

Buy it when four things are true together. The resident weights sit around 20–45GB, so a 24GB card cannot hold them. The machine will be on most workdays, not for a few nights a month. The data cannot leave the room. And you can live with 70B Q4 generation near single-digit tokens per second. In that case you are paying for residency and a data boundary. Three years of A100 or H100 rental sits an order of magnitude higher.

Skip it in equally concrete cases. An 8B–32B model used a few hours a day is cheaper for three years on a 4090, and the bandwidth is higher. Training, heavy concurrency, 70B at FP8, or a long context outgrow both 64GB and 273 GB/s. Rent an 80GB GPU by the hour, or look at the $6,950 128GB Spark. If the software must be macOS, Xcode, or Core ML, an Arm Linux box is the wrong system before price even enters.

Write down last week’s real hours

Add the hours a model actually occupied the GPU over the past two weeks, then multiply by 26 for a year. Do not fill the table with “we might leave it on all night later.” Inflated hours are how that three-year sheet talks you into a machine you will not fill.

It does not stand in for a Mac

DGX Spark answers a CUDA problem: the model should stay local, and it is larger than a consumer GPU. Claude Code and Cursor’s execution environment, Xcode builds, the simulator, and Core ML are a different problem. They need real macOS on Apple Silicon, not a GB10. Treating a $4,999 Spark as a stronger Mac mini misses both jobs. The CUDA stack is underused, and the iOS pipeline never starts.

When the heavy work is overnight CI, repo indexing, and long agent runs, a dedicated remote Mac usually fits better than another desktop GPU. On ZavCloud’s pricing page the Mac mini M4 16GB monthly reference is about $99.3 and the 24GB tier about $199.3. Thirty-six continuous months is about $3,575 and $7,175, and the price at checkout is the one that counts. Neither tier runs a 70B, and neither should be compared with Spark on CUDA throughput. They line up with the three-year Mac mini purchase versus Cloud Mac note, and with how a local Mac and a remote macOS node split the work. Keep the model on Spark or a cloud GPU, and keep packaging and signing on a Mac. That split is cleaner than asking one machine to do both.

Questions

Can the DGX Spark 64GB run a 70B?
Q4 is about 40GB, so yes, with a narrow context budget. FP8 is about 70GB, so no.

Which is cheaper over three years?
If the model fits in 24GB and you only need a few hours a day, a 4090 rental is usually cheaper. If you need around 40GB on most workdays, buying the Spark stays under hourly A100 or H100 rental. Leaving an H100 on for three years lands near seventy thousand dollars. That is a different question from this box.

Does 1 PFLOP mean it is faster than a data-center GPU?
No. That is an FP4 peak. A single stream is capped by 273 GB/s, about 7 tokens/s of bandwidth headroom on a 70B Q4.

Do two 64GB units make a cheap 128GB?
Not if the goal is memory. $9,998 against one 128GB system at $6,950 means you are paying for a second set of compute. Buy the second unit when you need the compute, not to assemble capacity on day one.

Can it host Xcode or Claude Code on macOS?
No. The system is DGX OS. The macOS toolchain belongs on a Mac, or on a remote dedicated Mac that takes the build and the agent loop.

  • Specs and starting price — NVIDIA DGX Spark, with 64GB sold through OEM partners
  • Price and on-sale date — Tom's Hardware, October 2026
  • 128GB price change and halved storage — The Register, 2 October 2026
  • Cloud GPU rates — RunPod pricing; the table uses public tiers still visible in late September 2026

ZavCloud Cloud Mac

Leave the model where it fits. Put the build on a Mac.

A dedicated Mac mini M4, real macOS, over SSH or VNC. It is for Xcode, Core ML, and an agent host you can leave running. It is not a 64GB CUDA box.

View Cloud Mac pricing
Cloud Mac View Cloud Mac pricing