Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.
 help



My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.

Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"


Agree. Switching models with a keypress is for developer coding.

Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.


For every production use case I build a eval suite which I use for prompt tuning and model evaluation and config. How else do you establish your model and prompt combination works? How do you upgrade to a newer model or decide on a fallback?

This same suite can simply be run very handoff to switch to a new prod model.


Jev can run your evals too :D



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: