I really want a 128gb+ machine but it's brutal to be at only 256 GB/s for $4k (especially with the drawbacks of both ARM and AMD).
I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+
GB/s comes out it'd be at $10k, way out of most consumers' hands.
A Mac Studio is a much better buy in terms of memory bandwidth, but impossible to buy in a 128 GB configuration. Honestly there aren’t great options right now and it’s probably better to wait for the market to be less insane.
I looked for one and it's impossible to find, let alone at a reasonable price + it does suffer from being harder to train/use less common models and workflows (e.g. arbitrary comfyui ones). Spark at least doesnt have that drawback, while AMD has both drawbacks.
Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.
No, equities naturally trend up with economic growth. RAM is up because of a supply shock, as new capacity comes on prices will drop, it’s a commodity.
Yes, in the very long term, but in the medium term where the listed capacities even matter we are not close to that. Really, the cost per gb of vram on flagships hasnt went down since 1080ti and thats not accouting for the recent increases which will likely last for years.
Asus rog 13 128gb is still sub $3k last I checked and ticks the boxes. If you get tired of AI it's a kick ass tablet except the weight and can run AAA games for now decently enough.
RAM is more like agricultural products (with short shelf-life) than commodities like fossil fuels, mineral ores, etc. You can manage an inventory or speculate on production, but you cannot really hold a "portfolio" of it in any sensible way.
So, you should get into RAM futures if you believe this is more than a transient arbitrage sort of situation. All extant RAM will become obsolete as the demand shifts to newer, fancier versions.
Yes but on what timescale? Replacing the > 5 year old sticks in my laptop would currently cost well over half what the machine ran me when it was brand new.
The Mac Studio M3 Ultra with 96GB of RAM with 1Tb SSD costs $6799, but has triple the memory bandwidth [1]. It's almost twice the price of Strix Halo and when the M5 Ultras come out it will be interesting to see how much they go for. I expect a lot of demand for them!
There was a post recently, that showed the DGX spark is twice as fast as the M3 Ultra at prompt processing, but half as fast at output tokens [2]. They used gps-oss-120 for that test with a small context.
I’ve got a 128gb m5 max mbp and two sparks. For my real-world use cases, a single spark running DS4 Flash will have fully responded by the time my Mac has even started generating tokens. I figured I’d have more generation heavy work when I also bought the Mac, but it has done very little work running LLMs since I got the first Spark.
I mostly run them clustered for DS4 and am quite happy with the performance, and the cost isn’t that much more for two than the MBP while giving me double the unified memory.
I’ll probably pick up a third to run multiple smaller models. I don’t understand why people would buy a halo over a spark at comparable prices, particularly because if you want to cluster, the cx7 be beats the shit out of them when it comes to latency and throughput
Apple rumor mill is suggesting that we may see M5 Mac Studio announced in September at that event. Apple just increased their pricing across the board, so I'm not holding my breath that they will be reasonably priced. They also cancelled development of their M6 in favor of the future M7 which I suspect will be a major, AI-focused upgrade compared to the M4->M5 upgrade
Yeah, folks should be aware that if you're filling up the memory on a Strix Halo for an inference workload, you're going to be getting uncomfortably slow token rates. Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token. It is not a realistic interactive coding model for agentic use.
I like my Strix Halo and keep it chewing on stuff, mostly non-interactive workloads (security audits of software mostly, training experiments, etc.), I get a lot of use out of it. If you want to experiment with AI, it is a good platform for that, though at $4k you can get an Nvidia-based Asus Ascend GX10, which is probably better. But, if you want a local model for interactive agentic use, you're going to be running either Qwen 3.6 or Gemma 4, which will fit comfortably on 2x64GB GPUs (even old GPUs will run them faster than the Strix Halo...I have dual Radeon Pro V620s which are faster, and they're six years old), or snugly on 32GB. A 48GB or 64GB Mac would run them well. Two Radeon AI Pro R9700 GPUs is probably the sweet spot, right now for GPUs. Not the cost of a good used car, like a 5090 or 4090, but plenty of memory and performance for local inference. Also, not finicky and weird and needing custom 3D printed fan shrouds like the old server GPUs on eBay.
At the moment, there just isn't a model that works better on a 128GB inference machine like this that don't also work fine on 64GB machines, which may be faster (very few 32GB GPUs will be slower, though I wouldn't recommend buying any GPU that isn't currently actively supported by the vendor drivers and CUDA or ROCm...so probably don't buy an MI50 or V100 or whatever).
I said, "Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token."
Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.
To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.
I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+ GB/s comes out it'd be at $10k, way out of most consumers' hands.
Edit: capitalized gb/s to GB/s.