Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and over 400 tok/s on concurrent requests which is plenty fast for a local model of this strength.
 help



Ok, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request.

Even without ninfer I would get over 80 on LM studio with default settings, so it should be noticeably more on 6000. You might want to try different a different inference engine or settings.

can't second ninfer enough. amazing tech

Is there an equivalent but for 4090s?

The repository has many forks, suggesting that folks are trying to (vibe) code support for different GPUs. Might be worth a shot.

dang only for certain nvidia GPUs, had my hopes up



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: