Dell's Version of the Dgx Spark Fixes Pain Points

Posted6d agoActive4d ago

thomasjb

145 points

87 comments

jeffgeerling.comTech DiscussionstoryHigh profile

informativepositive

Debate

20/100

Dgx SparkDell SolutionsAI Bubble

Key topics

Dgx Spark

Dell Solutions

AI Bubble

Dell's tweaked version of Nvidia's DGX Spark is sparking debate, with some commenters praising the company's efforts to fix existing pain points, while others remain skeptical about dealing with Dell's firmware updates. The discussion highlights a divide between those who value the DGX Spark's niche role in AI research and development, and those who argue that alternative hardware like Apple's Mac or AMD's Strix Halo offer better performance or value. Notably, owners of the DGX Spark device defend its unique strengths, pointing out that its capabilities are tailored to specific use cases that aren't directly comparable to other hardware. As AI research continues to gain momentum, this conversation feels particularly relevant, shedding light on the trade-offs and specialized needs of this rapidly evolving field.

Snapshot generated from the HN discussion

Discussion Activity

Active discussion

First comment

Peak period

6-9h

Avg / period

Comment distribution88 data points

Loading chart...

Based on 88 loaded comments

Key moments

01Story posted
Jan 1, 2026 at 2:11 PM EST
6d ago
Step 01
02First comment
Jan 1, 2026 at 3:14 PM EST
1h after posting
Step 02
03Peak activity
20 comments in 6-9h
Hottest window of the conversation
Step 03
04Latest activity
Jan 3, 2026 at 4:45 AM EST
4d ago
Step 04

Generating AI Summary...

Analyzing up to 500 comments to identify key contributors and discussion patterns

Discussion (87 comments)

Showing 88 comments

kachapopopow

6d ago

1 reply

Dell fixing issues instead of creating new ones? That's a new one for me. Would rather still not deal with their firmware updaters thought.

cjbgkagh

6d ago

1 reply

Give them a chance, I’m sure they’ll add new issues in one of their monthly bios updates.

kachapopopow

6d ago

1 reply

nothing beats perfectly good vendor firmware updates packaged in an obscenely complicated bash file that just extracts the tool and runs it while performing unnecessary and often broken validation that only runs on hardware that is part of their ecosystem (ex: dell nic on non dell chassis).

BadBadJellyBean

5d ago

2 replies

On linux I use fwupdmgr to upgrade the firmware on my dell laptop. Not sure if that works for servers though.

kachapopopow

5d ago

tends to be hit or miss when you use dell parts on non dell hardware (but the cost savings are worth it since typically nobody wants to touch dell hardware due to these issues)

buildbot

5d ago

Indeed, same process here: https://www.dell.com/support/kbdoc/en-us/000379162/how-to-up...

Tepix

6d ago

3 replies

You can get two Strix Halo PCs with similar specs for that $4000 price. I just hope that prompt preprocessing speeds will continue to improve, because Strix Halo is still quite slow in that regard.

Then there is the networking. While Strix Halo systems come with two USB4 40Gbit/s ports, it's difficult to

a) connect more than 3 devices with two ports each

b) get more than 23GBit/s or so, if you're lucky.

Something like Apple's RDMA via Thunderbolt would be great to have on Strix Halo…

Aurornis

6d ago

2 replies

The primary advantage of the DGX box is that it gives you access to the nVidia ecosystem. You can develop against it almost like a mini version of the big servers you're targeting.

It's not really intended to be a great value box for running LLMs at home. Jeff Geerling talks about this in the article.

cmrdporcupine

6d ago

2 replies

Exactly this. I'm not sure why people keep drumming the "a Mac or Strix Halo is faster/cheaper" drum. Different market.

If I want to do hobby / amateur AI research or do stuff with fine tuning models etc, learn the tooling. I'm better off with the DG10 than AMD or Apple's systems.

The Strix Halo machines look nice. I'd like one of those too. Especially if/when they ever get around to getting it into a compelling laptop.

But I ordered the ASUS Ascent DG10 machine (since it was more easily available for me than the other versions of these) because I want to play around with fine tuning open weight models, learning tooling, etc.

That and I like the idea of having a (non-Apple) Aarch64 linux workstation at home.

Now if the courier would just get their shit together and actually deliver the thing...

lostmsu

5d ago

1 reply

[delayed]

cmrdporcupine

5d ago

I had my finger over the buy button for various Strix Halo machines for weeks.

I ended up going with the Asus DG10 because if the goal is to "learn me some AI tooling" I didn't want to have to add "learn me some only recently and shallowly supported-in-linux AMD tooling" to the mix.

I hate NVIDIA -- the company -- but in this case it comes down to pure self-interest in that I want to add some of this stuff to my employable skill set, and NVIDIA ships the machine with all the pieces I need right in the OS distribution.

Plus I have a bias for ARM over x86.

Long run I'm sure I'll end up with a Strix Halo type machine in my collection at some point.

But I also expect those machines to not drop in price, and perhaps even go up, as right now the 128GB of RAM in them is worth the price of the whole machine.

mapontosevenths

6d ago

I have this device, it's exactly as you say. This is a device for AI research and development. My buddies mac ultra beats it squarely for inference workloads, but for real tinkering it can't be beat.

I've used it to fine tune 20+ models in the last couple of weeks. Neither a Mac or Strix Halo even try to compete.

saagarjha

5d ago

DGX Spark has a different compute capability, so no, you really aren’t.

coder543

6d ago

3 replies

[delayed]

plagiarist

5d ago

2 replies

Could I get your thoughts on the Asus GX10 vs. spending on GPU compute? It seems like one could get a lot of total VRAM with better memory bandwidth and make PCIe the bottleneck. Especially if you already have a motherboard with spare slots.

I'm trying to better understand the trade offs, or if it depends on the workload.

yowlingcat

5d ago

2 replies

[delayed]

fragmede

5d ago

1 reply

Thank you for sharing your painful learning experiences! What do you have now?

yowlingcat

5d ago

[delayed]

Y_Y

5d ago

I choose fast and cheap

coder543

5d ago

[delayed]

zozbot234

5d ago

1 reply

Prompt processing could be sped up with NPU inference. The Strix Halo NPU is a bit weird, but it's there.

EnPissant

5d ago

I’ve seen this claim a lot, but I’m skeptical. Has anyone actually published benchmarks showing a big speedup from using the NPU for prefill?

AMD’s own marketing numbers suggest the NPU is about 50 TOPS out of 126 TOPS total compute for the platform. Even if you hand-wave everything else away, that caps the theoretical upside at around ~1.6×.

But that assumes:

1. Your workload maps cleanly onto the NPU’s 8-bit fast path.

2. There’s no overhead coordinating the iGPU + NPU (which seems... optimistic).

My expectation is the real-world gain won't be very significant, but I'd love to be proven wrong!

EnPissant

5d ago

Then again, I have a RTX 5090 + 96GB DDR5-6000 that crushes the spark on prompt processing (something like 2-3x faster), while token generation is pretty close. The cost I paid was ~$3200 for the entire computer. With the currently inflated RAM prices, it would probably be closer to the dell.

So while I think the Strix Halo is a mostly useless machine for any kind of AI, and I think the spark is actually useful, I don't think pure inference is a good use case for them.

It probably only makes sense as a dev kit for larger cloud hardware.

benreesman

5d ago

NVFP4 (and to a lesser extent, MXFP8) work, in general. In terms of usable FLOPS the DGX Spark and the GMTek EVO-X2 both lose to the 5090, with NCCL and OpenMPI set up the DGX is still the nicest way to dev for our SBSA future. Working on that too, harder problem.

jasoneckert

6d ago

2 replies

I've got the Dell version of the DGX Spark as well, and was very impressed with the build quality overall. Like Jeff Geerling noted, the fans are super quiet. And since I don't keep it powered on continuously and mainly connect it remotely, the LED is a nice quick check for power.

But the nicest addition Dell made in my opinion is the retro 90's UNIX workstation-style wallpaper: https://jasoneckert.github.io/myblog/grace-blackwell/

ranger_danger

6d ago

1 reply

I just want a standard, affordable mini PC that looks like this one. Or better yet, with the brown accents normally found on recent PowerEdge systems.

https://www.fsi-embedded.jp/contents/uploads/2018/11/DELLEMC...

storus

6d ago

Zotac has a bunch of x64 mini PCs that use a similar hexagonal styling.

mapontosevenths

5d ago

I've had mine for a while now, and never actually connected a monitor to it. Now I'll have to. Thanks. :)

alecco

6d ago

1 reply

DGX Spark at $4,000 is a bad deal with only 273 GB/s bandwidth and the compute capacity between a 5070 and a 5070 TI. And with PCIe 5.0 at 64 GB/s it's not such a big difference.

I liked the idea until the final specs came out.

BadBadJellyBean

5d ago

1 reply

I think the selling point is the 128GB of unified system memory. With that you can run some interesting models. The 5090 maxes out at 32GB. And they cost about $3000 and more at the moment.

alecco

5d ago

2 replies

At launch date for that money you could've bought a decent rig with a 5090 32 GB DDR7 + 128 GB of DDR5 RAM. Notice also L2 96 GB vs 24 GB.

On one hand, /r/localllama doesn't like the Spark for running models (too low for compute and bandwidth). And I am a CUDA developer an find it overpriced.

Finally, 128 GB DDR5 was $1000 and now it's $3200. So I bet the DGX Spark will double or triple the price soon, too.

BadBadJellyBean

5d ago

1 reply

I am not on reddit. What are they saying?

mapontosevenths

5d ago

2 replies

It isn't for "running models." Inference workloads like that are faster on a mac studio, if that's the goal. Apple has faster memory.

These devices are for AI R&D. If you need to build models or fine tune them locally they're great.

That said, I run GPT-OSS 120B on mine and it's 'fine'. I spend some time waiting on it, but the fact that I can run such a large model locally at a "reasonable" speed is still kind of impressive to me.

It's REALLY fast for diffusion as well. If you're into image/video generation it's kind of awesome. All that compute really shines when for workloads that aren't memory speed bound.

lostmsu

5d ago

1 reply

[delayed]

mapontosevenths

5d ago

2 replies

Yeah, that's mostly fair, but it kind of misses the point. This is a professional tool for AI R&D. Not something that strives to be the cheapest possible option for the homelab. It's fine to use them in the lab, but that's not who they built it for.

If I wanted to I could go on ebay, buy a bunch of parts, build my own system, install my own OS, compile a bunch of junk, tinker with config files for days, and then fire up an extra generator to cope with the 2-4x higher power requirements. For all that work I might save a couple of grand and will be able to actually do less with it. Or... I could just buy a GB10 device and turn it on.

It comes preconfigured to run headless and use the NVIDIA ecosystem. Mine has literally never had a monitor attached to it. NVIDIA has guides and playbooks, preconfigured docker containers, and documentation to get me up and developing in minutes to hours instead of days or weeks. If it breaks I just factory reset it. On top of that it has the added benefit of 200Gbe QSFP networking that would cost $1,500 on it's own. If I decide I need more oomph and want a cluster I just buy another one and connect them, then copy/paste the instructions from NVIDIA.

saagarjha

5d ago

1 reply

You could also pay someone $5 an hour and they’ll give you a better machine for similar hassle.

mapontosevenths

5d ago

1 reply

But how much is MY time worth? Every hour I spend fixing some goof up Jimmy in I.T. made or Googling obscure incompatibilities is another hour I could have been productive.

Sometimes a penny saved is a dollar lost.

saagarjha

5d ago

I guarantee you the $5 a month option is easier than what you're setting up on a DGX Spark. Which should make sense, because you can buy server hardware for cheaper in the long run.

kouteiheika

5d ago

1 reply

> This is a professional tool for AI R&D.

Not really, not it isn't, because it's deliberately gimped and doesn't support the same feature-set as the datacenter GPUs[1]. So as a professional development box to e.g. write CUDA kernels before you burn valuable B200 time it's completely useless. You're much better off getting an RTX 6000 or two, which is also gimped, but at least is much faster.

[1] -- https://github.com/NVIDIA/dgx-spark-playbooks/issues/22

mapontosevenths

5d ago

Fair enough if that's your use case. I have to be honest with you though, I've never written cuda code in my life and wouldn't know sm_121 from LMNOPO. :)

It does seem really shady that they'd claim it to be 5th gen tensor cores and then not support the full feature set. I searched through the spark forums, and as that poster said nobody is answering the question.

nickthegreek

5d ago

what workflow/models are you using for media generation?

mi_lk

5d ago

2 replies

What’s GH and GB server?

alecco

5d ago

Grace-Hopper and Grace-Blackwell. "Grace" is the integrated CPU+GPU architecture. DGX Spark is GB10 and it's allegedly like a small version of the server GB200.

saagarjha

5d ago

GH200/GB200, Nvidia’s server hardware

colordrops

6d ago

2 replies

I assume they didn't fix the memory bandwidth pain point though.

llm_nerd

6d ago

2 replies

The memory bandwidth limitation is baked into the GB10, and every vendor is going to be very similar there.

I'm really curious to see how things shift when the M5 Ultra with "tensor" matmul functionality in the GPU cores rolls out. This should be a multiples speed up of that platform.

storus

6d ago

1 reply

My guess is M5 Ultra will be like DGX Spark for token prefill and M3 Ultra for token generation, i.e. the best of both worlds, at FP4. Right now you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part; with M5U that should no longer be necessary. However given RAM prices situation I am wondering if M5U will ever get close to the price/performance of Spark + M3U we have right now.

echion

5d ago

1 reply

> you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part

Are you doing this with vLLM, or some other model-running library/setup?

coder543

5d ago

[delayed]

kristianp

5d ago

1 reply

The M3 ultra was released about 18 months after the original M3, so you could be waiting a while for the M5 Ultra.

llm_nerd

5d ago

The M3 Ultra was oddly delayed, though rumours are that the M5 Ultra should arrive much quicker. Most are estimating March-ish. We'll see. I think Apple has a much higher motivation to get the M5 higher end variants out given the enormous benefits the new matmul functionality offers.

cat_plus_plus

5d ago

At least for transformers, it can be kind of fixed with MOE + NVFP4 for small working set despite large resident size.

dagaci

5d ago

1 reply

A nice little AI review with comparison of the CPU/Power Draw & Networking would be interested in seeing a fine-tuning comparison too. I think pricing was missing also.

geerlingguy

5d ago

I've been working on fine tuning testing, it's something I hope to set up for comparison against the Mac Studio and Framework Desktop clusters soon.

kristianp

5d ago

2 replies

I know it's just a quick test, but llama 3.1 is getting a bit old. I would have liked to see a newer model that can fit, such as gpt-oss-120, (gpt-oss-120b-mxfp4.gguf), which is about 60gb of weights (1).

(1) https://github.com/ggml-org/llama.cpp/discussions/15396

eurekin

5d ago

2 replies

Correct, most of r/LocalLlama moved onto next gen MoE models mostly. Deepseek introduced few good optimizations that every new model seems to use now too. Llama 4 was generally seen as a fiasco and Meta haven't made a release since

fragmede

5d ago

1 reply

What are some of the models people are using? (Rather than naming the ones they aren't.)

eurekin

5d ago

2 replies

GLM 4.7 is new and promising. Minimax 2.1 is good for agents. Of course the qwen3 family, vl versions are spectacular. NVIDIA Nemotron Nano 3 excels at long context and the unsloth variant has been extended to 1m tokens.

I thought the last one was a toy, until I tried with a full 1.2 megabyte repomix project dump. It actually works quite well for general code comprehension across the whole codebase, CI scripts included.

nightski

5d ago

Does GLM 4.7 run well on the spark? I thought I read it didn’t but it wasn’t clear.

magicalhippo

5d ago

[delayed]

kouteiheika

5d ago

2 replies

Llama 4 isn't that bad, but it was overhyped, and people in generally "hold it wrong".

I recently needed an LLM to batch process me some queries. I ran an ablation on 20+ models from Open Router to find the best one. Guess which ones got 100% accuracy? GPT-5-mini, Grok-4.1-fast and... Llama4 Scout. For comparison, DeepSeek v3.2 got 90%, and the community darling GLM-4.5-Air got 50%. Even the newest GLM-4.7 only got 70%.

Of course, this is just an anecdotal single datapoint which doesn't mean anything, but it shows that Llama 4 is probably underrated.

coder543

5d ago

1 reply

[delayed]

zozbot234

5d ago

1 reply

Large MoE models are more socially accepted because medium/large sized MoE models can still be quite small wrt. expert size (which is what sets the amount of required VRAM). But a large dense model is still challenging to get to run.

coder543

5d ago

[delayed]

eurekin

5d ago

Oh, this is very interesting. Will have to test it out on coding too. Very good point about testing. Had I only followed benchmarks, I'd miss few gems completely (long context models and 4b vision models that are unbelievably capable for their size). I'd encourage anyone to test the models on actual problems you're working on.

geerlingguy

5d ago

2 replies

That and more in https://github.com/geerlingguy/ai-benchmarks/issues/34

kristianp

5d ago

Thanks!

coder543

5d ago

[delayed]

barelysapient

5d ago

1 reply

Great article but would be nice to see how larger models work.

geerlingguy

5d ago

See: https://github.com/geerlingguy/ai-benchmarks/issues/34

cat_plus_plus

5d ago

1 reply

I have a slightly cheaper similar box, NVIDIA Thor Dev Kit. The point is exactly to avoid deploying code to servers that cost half a million dollars each. It's quite capable in running or training smart LLMs like Qwen3-Next-80B-A3B-Instruct-NVFP4. So long as you don't tear your hair out first figuring out pecularities and fighting with bleeding edge nightly vLLM builds.

echion

5d ago

> training smart LLMs like Qwen3-Next-80B-A3B-Instruct-NVFP4

Sounds interesting; can you suggest any good discussions of this (on the web)?

postalrat

5d ago

1 reply

Spark's biggest paint point is the price. Does it fix that?

bigyabai

5d ago

There's an entire line of Linux-supported Jetson products available for your perusal, in addition to all of the GTX and RTX cards that have native ARM64 support.

mmaunder

5d ago

2 replies

For those of you wondering if this fits your use case vs the RTX 5090 the short answer is this:

The desktop RTX 5090 has 1792 GB/s of memory bandwidth partially due to the 512 bit bus width, compared to the DGX Spark with a 256 bit bus and 273 GB/s memory bandwidth.

The RTX 5090 has 32G of VRAM vs the 128G of “VRAM” in the DGX Spark which is really unified memory.

Also the RTX 5090 has 21760 cuda cores vs 6144 in the DGX Spark. (3.5 x as many). And with the much higher bandwidth in the 5090 you have a better shot at keeping them fed. So for embarrassingly parallel workloads the 5090 crushes the Spark.

So if you need to fit big models into VRAM and don’t care about speed too much because you are for example, building something on your desktop that’ll run on data center hardware in production, the DGX Spark is your answer.

If you need speed and 32G of VRAM is plenty, and you don’t care about modeling network interconnections in production, then the RTX 5090 is what you want.

kouteiheika

5d ago

2 replies

> building something on your desktop that’ll run on data center hardware in production, the DGX Spark is your answer

It isn't, because it's a different architecture than the datacenter hardware. They're both called "Blackwell", but that's a lie[1] and you still need "real" datacenter Blackwell card for development work. (For example, you can't configure/tune vLLM on Spark, and then move it into a B200 and even expect it to work, etc.)

[1] -- https://github.com/NVIDIA/dgx-spark-playbooks/issues/22

benreesman

5d ago

1 reply

sm_120 (aka 1CTA) supports tensor cores and TMEM just fine: example 83 shows block-scaled NVFP4 (I've gotten 1850 ish dense TFLOPs at 600W, the 300W part caps out more like 1150). sage3 (which is no way in hell from China, myelin knows it by heart) cracks a petaflop in bidirectional noncausal.

The nvfuser code doesn't even call it sm_100 vs. sm_120: NVIDIA's internal nomenclature seems to be 2CTA/1CTA, it's a bin. So there are less MMA tilings in the released ISA as of 13.1 / r85 44.

The mnemonic tcgen05.mma doesn't mean anything, it's lowered onto real SASS. FWIW the people I know doing their own drivers say the whole ISA is there, but it doesn't matter.

The family of mnemonics that hits the "Jensen Keynote" path is roughly here: https://docs.nvidia.com/cuda/parallel-thread-execution/#warp....

10x path is hot today on Thor, Spark, 5090, 6000, and data center.

Getting it to trigger reliably on real tilings?

Well that's the game just now. :)

kouteiheika

5d ago

1 reply

Wait, so are you telling me all of the hardware/ISA is actually fully accessible and functional, and it's just an artificial PTX -> SASS compiler limitation?

Because the official NVidia stance is definitely that TMEM, etc. is not supported and doesn't work.

...I don't suppose you have a link to a repo with code that can trigger any of this officially forbidden functionality?

benreesman

5d ago

1 reply

I'm telling your it works now. It's just not called `tcgen05`.

Put this in nsight compute: https://github.com/NVIDIA/cutlass/blob/main/examples/79_blac...

(I said 83, it's 79).

If you want to know what NVIDIA really thinks, watch this repo: https://github.com/nVIDIA/fuser. The Polyhedral Wizards at play. All the big not-quite-Fields players are splashing around there. I'm doing lean4 proofs of a bunch of their stuff. https://v0-straylight-papers-touchups.vercel.app

It works now. It's just not the PTX mnemonic that you want to see.

kouteiheika

5d ago

Very interesting! Thanks! I'll definitely keep a close eye on that repo.

Anyhow, be that as it may, I was talking about the PTX mnemonics and such because I'd like to use this functionality from my own, custom kernels, and not necessarily only indirectly by triggering whatever lies at the bottom of NVidia's abstraction stack.

So what's your endgame with your proofs? You wrote "the breaking point was implementing an NVFP4 matmul" - so do you actually intend to implement an NVFP4 matmul? (: If so I'd be very much interested; personally I'm definitely still in the "cargo-cults from CUTLASS examples" camp, but would love something more principled.

my123

4d ago

Note that sm_110 (Jetson Thor) has the tcgen05 ISA exposed (with TMEM and all) instead of the sm_120 model.

chao-

5d ago

It's also worth nothing that the 128GB of "VRAM" in the GB10 is even less straightforward than just being aware that the memory is shared with the CPU cores. There's a lot of details in memory performance that differ across both the different core types, and the two core clusters:

https://chipsandcheese.com/p/inside-nvidia-gb10s-memory-subs...

cpgxiii

5d ago

1 reply

Absent disassembly and direct comparison between a DGX Spark and a Dell GB10, I don't think there's sufficient evidence to say what is meaningfully different between these devices (beyond the obvious of the power LED). Anything over 240W is beyond the USB-C EPR spec, and while Dell does have a question ably-compliant USB-C 280W supply, you'd have to compare actual power consumption to see if the Dell supply is actually providing more power. I suspect any other minor differences in experience/performance are more explainable as the consequences on increasing maturity of the DGX software stack than anything unique to the Dell version; particularly any comparisons to very early DGX Spark behavior need to keep in mind that the software and firmware have seen a number of updates.

geerlingguy

5d ago

Comparing notes with Wendell from Level1Techs, the ASUS and Dell GB10 boxes were both able to sustain better performance due to their better thermal management. That's a fairly significant improvement. The Spark's crusted gold facade seems more form over function.

nightski

5d ago

It's a product without a purpose.

supermatt

4d ago

Jeff, This is the second time you have been given a prosumer level cluster pretty much built for local LLM inference and on both occasions you have performed benchmarks without batching.

If you still have the hardware (this and the Mac cluster) can you PLEASE get some advice and run some actually useful benchmarks?

Batching on a single consumer GPU often results in 3-4x the throughout. We have literally no idea what that looks like on a cluster, because you arent performing useful benchmarks.

npalli

5d ago

Seems you are paying the Dell tax of 15%. The same setup is $4K from NVidia, Lenovo and $3K for 1TB at Asus.

https://www.dell.com/en-us/shop/desktop-computers/dell-pro-m...

graham33

5d ago

I have NixOS running on my DGX Spark: https://github.com/graham33/nixos-dgx-spark, would be interested to know if the USB image also boots on the Dell Pro Max GB10.

View full discussion on Hacker News

ID: 46457027Type: storyLast synced: 1/2/2026, 10:10:36 PM

Want the full context?

Jump to the original sources

Read the primary article or dive into the live Hacker News thread when you're ready.

Open link View on HN