Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...


50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?


I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.

qwen3.5:122b-a10b is significantly faster at around 60-65.


No magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...


With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further


It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them


I tried 8-bit, perhaps I should try 6-bit.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: