Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Not just RLHF but also RLVR, and isn't that the litter lesson though?

My sense of the Sutton Dwarkesh interview was that he was calling out that he didn't mean just longer datasets, but rather learning through exploration and that's exactly RL.



They just need more contact with reality. That's what RL is right? Contact with narrow subsets of reality.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: