Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Have you ever actually set up a multi-GPU system for inference? Based on my experience you are drastically overstating the problem. Both tensor and pipeline parallelism (without NVLink) produce a machine which is faster than any Mac on the planet, which is what we’re discussing here. Yes, each has pros and cons, and neither scales perfectly linearly. But it works great regardless.


No, and at $4000+ per 5090, I'm unlikely to have any experience anytime soon.


You can parallelize inference with much cheaper GPUs as well!

I’ve got a 4060 ti 16gb, and I’m thinking about getting another. I previously specced out a cluster using multiple 3090s. At the time, the 3090s were going for $700 on eBay. They’re more than that now, but there’s no need to spend $4k per GPU at all.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: