Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

From the METR report:

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.



No human could have read the reasoning traces by themselves:

> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.


So maybe not build a non deterministic black box that can't be reasoned about ?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: