Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

10 grand for 256GB memory. Likely double that for 512GB, but won't be available or finalized until October. Thunderbolt 5 is highest bandwidth external IO available at 120Gb/s. 1.2TB/s claimed max internal memory bandwidth.

Not exactly "future proof" for >1T parameter models but good for targeting specific lower-parameter models, or if you can rely on pipeline parallelism and run a cluster.



> Not exactly "future proof"

Computers are never "future proof".


They do however last a much longer time now.

A 8 years old graphics card can still play modern games. Ten years ago playing a modern game on hardware that old would've been unthinkable.

And with how the market is right now, we'll be stuck on the current "reference level" of hardware for a while longer.


I recently got a refurbished Strix Halo 128GB desktop for a steal ($2200) and I have a feeling it's going to be my last PC purchase for the next 3-5 years. Who knows how long this RAMpocalypse will last. My Lunar Lake 32GB laptop is already a couple years old, and it's not a speed demon but it ticks all the laptop boxes (portable, great battery).


About 20 years ago my dad bought me a $5k computer, it was future proof for about "5 years" before we had to upgrade its internal parts (more memory, new graphics card).

It was future proof but not really because it struggled a lot in its final years.


My 6 year old mediocre gaming PC is still a mediocre gaming PC.

I expect it to stay a mediocre gaming PC for the next 3 years maybe 5 years.

16GB RAM RTX 3060 TI (8 GB VRAM)


Nice. Curious, what do you play with that thing? Do you think you'll still play those in 3-5 years? And are there games you do want to play but can't?


I like to test new releases that are story heavy.

I also like 4K with 90 to 120fps which leads me most of the time to use almost minimal settings.

Look at steam hardware reviews. My setup is quite the default thing. Games optimize for it


He had that option,l. However buying apple hardware todqy you are stuck and need to replace all of it even if for example the CPU is plenty fast but you need more memory or GPU power.


i'd say it depends (as with everything)

i'm still using an old i7 3770k @ 4.8ghz with 16gb (ddr3) ram running linux for random tasks like executing tests. obviously, power consumption is higher.

my main machine is a MBP M1 Max which i use for everything. i also have my main linux desktop workstation that has a 5950x with 128gb (ddr4) ram.

i'll probably get 2-5 years out of my MBP, and my AMD workstation will probably be good for another 5-10 years.

i'm not a gamer, but i have a 3080. i'm sure my 5950x will still be good for gaming in 10 years if paired with a modern GPU.


I’d say something like a PS4 is future proof. $350 got like 10 years of modern software.


Nobody calls a PS4 a computer. We all know it is internally with the hardware, but the whole package makes it not a computer.


PS4 is as close as it gets to the colloquial computer; it's a AMD APU running BSD with homebrew linux builds available.

And if not, would you consider a raspberry pi a computer? or your iphone a computer?


> Computers are never "future proof".

Upgradeable components however could go a loooong stretch towards that goal. It can't be that hard to follow a common form factor for at least the housing across two or three generations to allow a reuse of everything but the main PCB.


Think it would have same memory bandwidth if the RAM was upgradeable?

Would be nice if someone knowledgeable about electrical engineering and manufacturing processes could lay out some valid reasons for manufacturers to integrate RAM onto the motherboard.

https://news.ycombinator.com/item?id=49041256#49082206


It isn't integrated into the motherboard, it's integrated onto the same package as the CPU/GPU which allows for better signal integrity and higher speeds. You can get somewhat close to the same speeds while modular with tech like LPCAMM2, but there are some pretty difficult challenges to overcome to close the gap completely. As an example, the Framework Laptop 13 Pro CPUs support up to 9600MT/s (same as M5), and Micron sells LPCAMM2 modules that can run at 8533MT/s, but the 13 Pro only officially supports 7467MT/s.


Each connector will introduce "insertion loss" to the chain, eating into your signal integrity budget.

The SO DIMM Ram slot, which is designed, idk, 30 years ago? is not really capable of handling the frequency we are targeting (close to 10GT/s)


If it was upgradable, then yes, spending more on top of it every year would make it future proof, but that's not the point. It's that spending 10 grand doesn't get you a future proof computer today.


The main PCB is pretty much everything that has value. The rest is a heatsink, case and PSU.


The way the Apple M-series does ram that might be difficult to pull off.


> The way the Apple M-series does ram that might be difficult to pull off.

Well it might be an idea to keep the layout of the mainboard and connectors the same.

That way, instead of having to upgrade the whole machine, all it would need is a new mainboard. Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.


> Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.

It's not the same thing though. On the M-series, CPU and GPU share a unified memory architecture and ram is much more tightly coupled to get it to go faster. A closer example would be the Framework desktop, actually, where memory is also soldered in for the same reason.


That actually would be interesting - yes, the computer itself is basically a PCB, but it’s also wrapped in a couple pounds of aluminum, a power supply, cooling fans, and a few other things that don’t need to be consumables. Upgrade a mini or a studio by swapping the new board into the old case - yeah, you’re not saving much money, but you also don’t need to throw out the entire rest of the case, and you can ship the main board in the space of a couple CD cases.

It’s a very non-Apple thing to do, but it’d be pretty awesome if they did.


>Computers are never "future proof". Till now.

My thinkpad t14g1 with 16gb ram somehow still relevant today.


My main PC at home is 15 years old at this point, running linux. It's fine.

I don't play games, but it still chugs along perfectly fine and for what I use it for I rarely think that it's too slow.


Same thing, I still have at home a desktop I've built when I was 17 (now I'm 30!) and still runs fine for day to day tasks, of course with Linux on it. It has 8Gb of RAM (DDR3), 512Gb SSD (not original of course because back then they did cost an astronomic price) and an i7 CPU that I've overclocked, an NVIDIA GPU with 2Gb of memory that still plays older games fine.

To me there is no reason to update a PC if you don't use Windows or MacOS that forces you to purchase a new hardware to do the same things. Just look at Windows 11, full of useless AI features, weights a ton, and in the end it's probably faster a PC from 25 years ago with Windows XP.


> 10 grand for 256GB memory.

A NVIDIA RTX 6000, 96 GB at 1.7 TB/s, is 13 grand.

This 256 GB at 1.2 TB/s Mac is extremely competitive, it will be sold out everywhere.


The relevant comparison isn't one mac studio to one RTX 6000, it's a 24 channel DDR5 system, which also has ~1.2TB/s of memory bandwidth (or more when Xeon 6 compatible 8800mt/s memory becomes widely available), vastly higher prefill due to more CPU horsepower, orders of magnitude faster networking, can hook into GPU accelerators, can be upgraded etc. A baseline 384GB system from eg Puget is ~30K vs ~12K for the 256GB Mac Studio and you do get value for the money.


Prefill is gated by GPU compute, not CPU compute. A Epyc/Xeon CPU will never beat an M5 GPU in prefill. If you want something that can match M5 Ultra in prefill, you'd need to add beefy Nvidia GPU. However, the problem is that the GPU memory is separate from the 384GB system memory. And here lies that advantage of Apple Silicon. It's unified memory and accessible by the CPU and GPU.


So 3x more, plus the cost of a GPU (another 10k?). How is that value for money to get slightly better performance?


It can be more than "slightly", particularly if the model you're interested in (or will be interested in in 6 months) doesn't fit on the mac studio. You also need to account for eg storing 10TB of random checkpoints, load time when experimenting, and so on. When you start actually needing throughput these are all capability gaps in practical use, not just x% benchmark differences.

If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one.


A 3x price difference means you can get 3 256GB mac studios which you can connect through thunderbolt and with RDMA a total of 768GB ram with compute/memory bandwidth scaling basically linearly (with a small overhead cost).


Doesn't make any sense since you can chain 3 M5 Ultra studios for the same price to get 768 which is 2x more than your system.

Lastly, I want to clarify that prefill on an x86 CPU is drastically slower than on an M5 Ultra GPU.


You can find model variants to scale up to whatever capacity you have. Queen has models that just barely fit in 256gb, I ran them… okay… on my Studio.


Excellent post. Heck yes. And with MRDIMMs coming, we're going to get another >50% boost in throughput per channel real soon, with a massive uptick in max capacity (4x).

It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers.


How's the compute side now, I wonder? Because while the Ultras have impressive memory bandwidth for inference, processing prompts still takes a dog's age on my M3 Ultra. I heard the M5 makes some strides forward in this area, though, and the M7 in particular promises to go a lot further.


M5 is excellent, they’ve finally gotten their own tensor cores.

Good for inference; however if you like to train, data format support and effective performance is limited (M5 Pro). Some hardware features are not exposed or extremely slow.

You’ll be fine for inference, but pales in comparison to what a RTX 6000 Pro can do for compute/matmuls/training.


4x faster prompt processing than M3 Ultra.


I've been using Macs as my primary machine for over three decades – but in the past year a Microcenter-Prebuilt (with a 5070Ti) has helped make Ubuntu my preferable computing environment (I have three Silicon Macs with similar access to 16GB vRAM – the 5070Ti "smokes" them in compute).

But my next machine for LLMs will probably be the Mac Mini (with 64GB vRAM access) for low-$2300s – which is what I paid for my MC-Prebuilt (and am still very happy with – Ubuntu is great).

After decades of installing and forgetting various shades of linux-distro, Ubuntu has kept me "using" the computer as a tool, instead of "just tinkering with it").


Except the RTX 6000 will run circles around the Mac studio in just about every way. Memory bandwidth is literally the only spec where Apple is competitive, and while high memory bandwidth is necessary for LLMs to perform well, many people strangely don't understand that memory bandwidth alone is not sufficient.


Mac studio wins in memory capacity, price, perf/watt and value.

RTX 6000 wins in performance, if your model can fit into the VRAM.

There are very obvious and clear advantages to a Mac Studio. It's an entire system for one and you're getting a world class CPU as well.


> Mac studio wins in memory capacity, price, perf/watt and value.

[citation needed]. I have personally specced out and built an nvidia GPU-based machine which after some optimization, handily beat the Mac Studio in terms of tokens/watt for LLM inference with most models. This was in the M2 Ultra era, and I haven't run the numbers for the later generations, but nvidia's cards have gotten faster just as Apple's CPUs/GPUs have, so I would guess that it's still possible to do.

> RTX 6000 wins in performance, if your model can fit into the VRAM.

"if your model can fit into the VRAM" can be true for the Mac as well.

> There are very obvious and clear advantages to a Mac Studio.

There are certain advantages for sure, depending on your use case. They may _seem_ to be obvious, but as evidenced above, I believe that many people overestimate the Mac's superiority on the metrics you cite when comparing a Mac vs. a dedicated GPU for LLM inference.


> "if your model can fit into the VRAM" can be true for the Mac as well.

It is much more likely for your model to fit in large unified memory of a Mac than the smaller more limited memory of a GPU. Even going with two 5090s, you now have to shard your model and that is a PITA.

But it turns out that MoE is the solution both for running models on macs of limited computer power means (not as fast as GPUs), and on multiple GPUs that require sharding the model.


> It is much more likely for your model to fit in large unified memory of a Mac than the smaller more limited memory of a GPU.

I bristle at general statements like this when it obviously depends on the specific Mac and GPU in question. But yes, comparing a maxed out M5 Ultra with an RTX 6000, the Mac has much more memory.

> Even going with two 5090s, you now have to shard your model and that is a PITA.

Every modern tool does this for you automatically. It is absolutely not a pain in the least (e.g. llama.cpp ships with pipeline parallelism enabled by default).


Sharding a dense model using Tensor Parallelism (TP) across dual RTX 5090s has a significantly worse performance penalty over PCIe than sharding an MoE model.

If you are using multiple GPUs, MoE is basically going to be your only workable choice unless you can leverage pipeline parallelism (only half your GPUs can work on a prompt at a time, so you need to process prompts back to back in a pipeline setup, and they better be doing similar things because your vram is limited).


Have you ever actually set up a multi-GPU system for inference? Based on my experience you are drastically overstating the problem. Both tensor and pipeline parallelism (without NVLink) produce a machine which is faster than any Mac on the planet, which is what we’re discussing here. Yes, each has pros and cons, and neither scales perfectly linearly. But it works great regardless.


No, and at $4000+ per 5090, I'm unlikely to have any experience anytime soon.


You can parallelize inference with much cheaper GPUs as well!

I’ve got a 4060 ti 16gb, and I’m thinking about getting another. I previously specced out a cluster using multiple 3090s. At the time, the 3090s were going for $700 on eBay. They’re more than that now, but there’s no need to spend $4k per GPU at all.


  [citation needed].
No need. You can infer the logic with this line I wrote:

  RTX 6000 wins in performance, if your model can fit into the VRAM.
I'm not sure what the controversy is here.


You made claims about performance per watt and other metrics which were completely unsubstantiated and aren't backed up by the line you quoted. That's what I was asking for citations about.


What other metrics?


Do I need to quote your own statement to you, verbatim, again?

> Mac studio wins in memory capacity, price, perf/watt and value.


Mac Studio winning in memory capacity is just a fact. That's why I'm confused.


Yes, I agree. What about the other metrics, including performance per watt which I’ve now mentioned four times and you’ve ignored the previous three times?


https://www.youtube.com/watch?v=nwIZ5VI3Eus

This is a good video to watch.

  Yes, I agree.
Glad you agree. I was just confused why you were questioning it. The big biggest advantage for Apple Silicon is that you can get much more VRAM per $ over Nvidia cards. No controversy.


You seem to be acting deliberately obtuse.

> What about [...] performance per watt which I’ve now mentioned ~~four~~ five times and you’ve ignored the previous ~~three~~ four times?


In the video.


No, it's not in the video.

The video was not focused on performance per watt at all, and the best attempt that the video makes at measuring performance per watt actually shows the opposite: that the 5090 system is more efficient than the Mac.

When he's running the 27B model, at about 5:02, he shows that the Mac Studio is pulling 251.5W, and the PC is pulling 315.3W.

Then he shows the results:

Mac: 27.62 tok/s PC: 40.92 tok/s

Doing the math, we arrive at 0.11 tokens/W for the Mac, and 0.13 tokens/W for the PC. The PC is about 20% more efficient.

The only other time he even shows the power usage at all is near the beginning, running a 4B model which is trivially small for both systems.

So once again, I ask, what are the sources for your performance/watt claim?


Energy efficiency is in another league with the Mac Studio for local workloads.

I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly.


I’ve run the numbers on this, and an optimized nVidia build can meet or beat Apple platforms in terms of tokens per watt, which is the most important efficiency metric if what you care about is using the least amount of energy to generate a given response.

Yes, the Mac might get lukewarm, but it will take 2-3+ times longer to do the same task.


Yes I agree that at full speed with parallel workloads, the tokens per watt are better using Nvidia GPU servers.

Also while the pre fill performance sucks, the Mac isn’t that slow and can use much better model compared to a similarly priced Nvidia workstation so it’s not really taking much longer in practice.

It’s taking longer than in the cloud for sure. At least a local computer uses the local energy grid that is pretty clean and not fossil energy.


Okay, so you said “Energy efficiency is in another league with the Mac Studio”; I assume you meant that it was more efficient than an nVidia GPU setup. But when pressed, you agreed that’s not actually the case. So it’s not actually that much different in terms of energy efficiency.

> Also while the pre fill performance sucks, the Mac isn’t that slow and can use much better model compared to a similarly priced Nvidia workstation so it’s not really taking much longer in practice.

This is exactly backwards; assuming you mean that the Mac has more RAM so you can use a larger model, the mac is going to be _even slower_ since the fastest Mac isn’t as fast as the average nVidia setup.

My point is that comparing apples to apples (no pun intended), an nVidia setup is both faster and (possibly with some tweaking) more efficient than a Mac.


It better because you’ll need a few of them to run some larger models (I’ll be just as vague citing which models).


I was not speaking about specific models so I didn’t feel the need to cite any. Not sure why the backhanded insult was necessary. You need multiple Mac Studios to run the largest models as well, so neither is a one-size-fits-all device.

If you can show me a model for which a Mac is faster than the RTX 6000 then I’ll be happy to update or retract my statement.


That's in the same ballpark as two 128GB AI machines like the Asus GX10 or DGX Spark or Strix Halo. And, it seems very likely to perform better than either of those for inference. And, 256GB brings some pretty good models into play.

But, that doesn't make it a good deal. It just means the Apple tax doesn't apply when stacked up against AI machines and with memory prices being so out of whack. I'm still planning to wait until the RAMpocalypse ends before I buy any more hardware.


Is there anything in the works or planned that suggests the RAMpocalypse will end anytime soon? e.g. new fabs being built, permitted, planned, etc...


Define "soon"

RAM production is completely sold out for 2027[1] which means the prices are locked in until after then.

It takes about 2 years from the time ground if broken for a new fab to be built and producing RAM.

There were some new fabs announced between February and April this year by both the Korean and Chinese manufactures, so that new capacity might start having an impact in 2028 in the most optimistic scenario.

Samsung says supply will remain tight in 2028[2], and Micron says "tight beyond 2027"

The best hope is that new (Chinese) players overbuild fab capacity and supply outstrips demand. That isn't likely, but perhaps in the late 2028-2029 timeframe could happen.

[1] https://www.techpowerup.com/351344/memory-makers-seal-2027-d...

[2] https://www.tweaktown.com/news/112966/memory-shortages-will-...

[3] https://s25.q4cdn.com/621799436/files/doc_events/2026/06/Q3-...


No one making statements has an incentive to tell the truth though (in fact, they're all incentivised to keep proclaiming the RAMpocalypse will never end).

Producers (Samsung, SK Hynix etc.) will not say "prices expected to drop" or "demand expected to drop" even if it was true because then consumers would start delaying purchased.

The big buyers (OpenAI, hyperscalers etc.) have no incentive to say "supply expected to start opening up" because that would imply their growth trajectory is flattening; also a huge part of their moat now is just deployed RAM.


There was an article last week that a Chinese RAM manufacturer was planning to add new fabs to be able to ramp up. More or less simultaneously there was also an article on Apple considering switching to using china sourced RAM for China destined devices.


Lots, if your time horizon is ~5 years for not buying new hardware, you're good.

Otherwise you'll have to wait to see if the AI circular financing club collapses- if you still have a job, there should be deals to be had...


There are signs that the money faucet is being turned down. Various investments that were announced have been quietly canceled or reduced in scale. I don't think it'll happen soon, but it seems unlikely to be more than a couple years. If I were a betting man, 12-18 months seems right. The AI companies that can make enough money will survive, the ones running on investor cash and debt, won't.

Efficiency is improving, both in hardware and in software and in intelligence density (smaller models can effectively do more of the AI work that needs doing), so I think the pure data center plays will falter. If there isn't some other business attached, they're never going to recoup their investment. Anthropic and OpenAI are buying all the compute they can find right now, but efficiency gains, especially those coming out of Chinese labs where they must be more efficient to compete, will make it less and less of a problem.

I mean, think about the hardware we use for AI. It's basically an accident. GPUs were not designed for AI (though they are becoming more focused on AI). The specialized AI hardware industry is just ramping up.

So, we're still early in the curve for how efficient both the hardware and software can be at performing these tasks, and given the effectiveness of recent very small models (e.g. DeepSeek V4 Flash 0731 and Qwen 3.8 27B), I just don't see a long future for giant data centers built around billions of dollars worth of last years graphics cards. As with the crypto mining operations, at some point, it becomes more expensive to run the hardware than it makes in revenue. And, as with the crypto mining operations, when the money dries up, the hardware hits eBay and prices drop.


There are some news about AI/LLM progress rate flattening out: People uses the cheaper model more than better model. Some AI startup in Chinese lose half of value.


The memory bandwidth on the Spark and Strix Halo is a fraction of the M5 Ultra. "Perform better" is probably a huge understatement.


Yes, if I was going to throw away ten grand on a computer that can run models that are much worse than what I can rent from a variety of providers for a few bucks a month, I'd buy the Apple.


Where are you getting DeepSeek for a few bucks a month? I spend around $3/day on it, sometimes $4.


I didn't mention DeepSeek, though with the caching effectiveness in Reasonix, you can get DeepSeek for pretty danged cheap directly from deepseek.com, even at the higher prices. I was referring to the many subscriptions available that will sell you quite a lot of usage for $20/month. Given how slow local inference is, you're probably not going to ever get your money back out of it if you just use it for inference (or training/fine-tuning, since you can rent GPUs for a buck or two an hour and ten grand is a lot of hours).

When memory prices come back down, I'll be down to the Apple Store. But, it doesn't make sense to buy hardware right now.


I don't think the $20/month packages come with API access at all.

I bought a refurbished M3 Max a couple of years ago just to try things out, I get 90 toks/s which is good enough for a lot of tasks. But ya, when DRAM prices come back down I'll definitely look for a better local rig.


The Qwen and Z.ai coding plans include API access, last I checked. OpenCode Go also has API access to a bunch of models for $10/month.


There is no "future proof" for >1T param models, there is no present or future where you can run a model that size on consumer hardware.


I don't know that "consumer hardware" is a useful distinction anymore, it's just "what's your budget and what's your speed requirement".


Consumer hardware means:

- 120v input plug

- not rack-mounted

- has a video out port


To me, "consumer item" is when i don't have to contact sale people for price.


Ah, a MacBook Professional is Consumer /s


In computing "Pro" is not the opposite of "Consumer".

Putting aside the fact it is a marketing label, "Pro" usually means "designed for work" while "Consumer" (in this context) means "doesn't need a special environment".

In computing the distinction is primarily noise, power and cooling requirements.

If a computer is designed to use home power and is quiet enough to use without annoying people and doesn't require specialist cooling then it is a consumer device, even if it is used for work.


Apple's use of "Professional" is just a fancy alias for high end or premium, it in no way indicates anything about being "for professionals" (most extreme example: "professional" iPhone models)

Hence why they had to make up the "Studio" brand for the workstation market, because they'd already fully removed any meaning from "Professional"


Yes?


Pro?


Want to bet that consumers with enough money buy it?


6 Mac Studios for 100k USD. Consider it a rule of thumb now - 100k to run 1T params, scales linearly.


If you have 100k you can just buy a few b200s.


One does not simply "buy a few b200". They come in 8 packs minimally, at >5x that budget.


Out of curiosity, what do you do to distribute the weights and inference across the multiple machines?


Thunderbolt 5, and software


at a 100k$ you're not buying apple products to be limited by the 120gb/s thunderbolt port.


3 of those thunderbolt 5 ports, so you can do a fully connected 4 machine cluster topology.


It's unclear to me how bandwidth scales with multiple connections. Many-to-many does not seem ideal. Daisy chaining would be fine for straight pipeline work. There doesn't seem to be an equivalent of a ethernet switch for thunderbolt 5 though.


Apple put on the announce page that 4 Studios in this config can run inference at 3x the speed as 1 Studio.


I read each has an independent controller for full bandwidth.


It will always be wasteful to have a GPU sleeping next to you 99% of the time. Something like OpenRouter is the solution imho, for me at least. I realize that that some people care about privacy more though, and I respect that.


10 grand for 256GB new Ultra sounds too cheap in today’s crazy DRAM market, it feels too good to be true.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: