Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
sscaryterry
27 days ago
|
parent
|
context
|
favorite
| on:
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
smcleod
27 days ago
[–]
50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
Casteil
27 days ago
|
parent
|
next
[–]
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.
qwen3.5:122b-a10b is significantly faster at around 60-65.
smcleod
27 days ago
|
root
|
parent
|
next
[–]
No magic, just oMLX with MTP. You can look through the speed the community is getting here:
https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...
syntaxing
27 days ago
|
root
|
parent
|
prev
|
next
[–]
With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
Casteil
27 days ago
|
root
|
parent
|
next
[–]
It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
sscaryterry
27 days ago
|
parent
|
prev
[–]
I tried 8-bit, perhaps I should try 6-bit.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: