Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

A 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s

The 3080ti is 912.4 GB/s



But it also has 8GB of RAM.


And the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...


You got me curious so I looked up the previous chips[0].

Memory bandwidth M1: 68 GB/s M2: 100 GB/s (47% increase) M3: 100 GB/s (0% increase) M4: 120 GB/s (20% increase) M5: 153 GB/s (27.5% increase)

So, M6: 170 GB/s (11% increase) doesn’t seem impossible, though I would have expected more.

[0]: https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4...


The different models of chips and memory config have very different memory speeds as well.

Eg the M4 Max 128GB has a bandwidth speed of 500GB/s+. And that's true for other models as well.

But as you note, the base speed has also increased over the versions.


From a bandwidth perspective, the ultra is like 8 M5s fused together (@ 150GB/s), that's how it gets to the 1200.

Historically the Pro doubles the base, the Max doubles the Pro, and the Ultra doubles the Max.

If an M6 ultra were released today it would be 1.36TB/s.


does that mean they're measuring bandwidth differently than how others (like nvidia) does it? memory bandwidth is the gating factor of running models locally, so if it's actually 8x 150GB/s, it may help something like prefill, but would it actually speed up decode comparatively?


No, they use the same definition of memory bandwidth as others, but Apple Silicon has a lot of memory channels. In previous generations, prefill has been compute-limited and decode is fast.

https://blog.exolabs.net/nvidia-dgx-spark/ outlines a combination of a DGX Spark and an M3 Ultra that took advantage of fast prefill on the Nvidia hardware and fast decode on Apple Silicon.


Not a hardware engineer, but it's mainly because of RAM wires/channels (not implying that this is "simple" form an engineering perspective).

Using the published bandwidths, the math is 170 * 1 and 153 * 8.


But 170GB/s is almost nothing? None of the RTX 50 series GPUs has that low bandwidth, you have to go back two generations of nvidia GPUs to get closer to that, and then it's the cheapest of the series, RTX 3050, which has ~170GB/s.

Even the GTX 1080, launched ten years ago, has double the bandwidth!

This must be some different way of measuring the bandwidth right? Since they explicitly say this for AI, but the numbers they share don't show that at all. Or I gravely misunderstand something here.


> Since they explicitly say this for AI

That's marketing spin indeed (or lies, if you prefer).


GFX vram is still faster.


Enjoy Gemma 4 E2B at blistering speeds, I guess?


My point was more: this was the 2nd lowest end card from a generation 6 years ago, and it had way higher bandwidth than today's alleged flagship.


Small amounts of fast ram vs huge amounts of slower ram.

It costs more than the strix to just buy regular ddr5 ram sticks today.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: