Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is great. I have a weird system layout (192gb system ram, 8gb vram). the mixture of experts models have been nice when i can run the dense reasoning layers on the gpu (which somehow fit?!) and then the expert on the cpu.

its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: