Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> OpenAI looked at user data, stole world class researchers' work

This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else.

It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.

 help



No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell

"I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."

Let's see what statement OpenAI will come up with for their side of the story

EDIT: precised my thought on user data vs session data


I was downvoted initially, look what OpenAI shared https://openai.com/index/navier-stokes-solution/ ...

"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve."

"When a further trained version of our internal model became available over the course of the effort"

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "


If it's siloed the same way the HF bots were, that doesn't exactly bode well. I'd be amazed if there weren't some big companies sending a fleet of lawyers at OpenAI's ZRPs after this news

If you have work happening in a part of your latent space that's got a much lower representation in your dataset then it's pretty plausible to include it. It doesn't actually matter who the user is if there's not a lot of people in the world working on problem X and you have a dataset of work on problem X.

They likely train on logs.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: