Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: