Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Claude might as well be .5 tok/s. I end up waiting several minutes and what it tells me could usually be summarized in under 100 words.

So I could potentially live with this if it was concise.



This is not how thinking works. Claude uses tens of thousands of thinking tokens to get to 100 words. Kimi is no different.


Additionally, GP can use the word ‘concise’ (or similar) in their prompt if they want more concise output from a model.


That doesn’t mean “use fewer thinking tokens”. It might use more as it mulls over how to make its response concise.


effort level is a separate parameter, and is probably what they want


I have one setup that gets about 1.5 tok/s of a very large on prem LLM, on a system that lives under my desk. It's used for overnight project review runs and code review that it is fed at the end of each work day. When I look at it the next morning it has done quite a lot of useful work. Dealing with a big slow LLM as an effective tool is really about planning the workflow to feed it.


In my experience Kimi k3 is even more verbose.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: