And the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...
does that mean they're measuring bandwidth differently than how others (like nvidia) does it? memory bandwidth is the gating factor of running models locally, so if it's actually 8x 150GB/s, it may help something like prefill, but would it actually speed up decode comparatively?
No, they use the same definition of memory bandwidth as others, but Apple Silicon has a lot of memory channels. In previous generations, prefill has been compute-limited and decode is fast.
https://blog.exolabs.net/nvidia-dgx-spark/ outlines a combination of a DGX Spark and an M3 Ultra that took advantage of fast prefill on the Nvidia hardware and fast decode on Apple Silicon.
But 170GB/s is almost nothing? None of the RTX 50 series GPUs has that low bandwidth, you have to go back two generations of nvidia GPUs to get closer to that, and then it's the cheapest of the series, RTX 3050, which has ~170GB/s.
Even the GTX 1080, launched ten years ago, has double the bandwidth!
This must be some different way of measuring the bandwidth right? Since they explicitly say this for AI, but the numbers they share don't show that at all. Or I gravely misunderstand something here.
The 3080ti is 912.4 GB/s