I’ve run the numbers on this, and an optimized nVidia build can meet or beat Apple platforms in terms of tokens per watt, which is the most important efficiency metric if what you care about is using the least amount of energy to generate a given response.
Yes, the Mac might get lukewarm, but it will take 2-3+ times longer to do the same task.
Yes I agree that at full speed with parallel workloads, the tokens per watt are better using Nvidia GPU servers.
Also while the pre fill performance sucks, the Mac isn’t that slow and can use much better model compared to a similarly priced Nvidia workstation so it’s not really taking much longer in practice.
It’s taking longer than in the cloud for sure. At least a local computer uses the local energy grid that is pretty clean and not fossil energy.
Okay, so you said “Energy efficiency is in another league with the Mac Studio”; I assume you meant that it was more efficient than an nVidia GPU setup. But when pressed, you agreed that’s not actually the case. So it’s not actually that much different in terms of energy efficiency.
> Also while the pre fill performance sucks, the Mac isn’t that slow and can use much better model compared to a similarly priced Nvidia workstation so it’s not really taking much longer in practice.
This is exactly backwards; assuming you mean that the Mac has more RAM so you can use a larger model, the mac is going to be _even slower_ since the fastest Mac isn’t as fast as the average nVidia setup.
My point is that comparing apples to apples (no pun intended), an nVidia setup is both faster and (possibly with some tweaking) more efficient than a Mac.
I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly.