Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

No way you can get that much each without extremely quants

For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp

Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense

 help



You can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.

Can confirm. I've a single R9700 and have maxed out at 45 tokens/second on llama.cpp with Q4 Qwen 3.6 27B (with MTP)

Q4 isn't an extreme quant, and I average 75 toks/s on code, 45 tok/s on prose with MTP.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: