A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
I work in catastrophe risk modeling and it's a multi billion dollar industry.
We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models.
If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
> If OpenAI is indeed using customer data to train their models to win a $1m prize
Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS:
Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing.
Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format.
There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results...
> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
If you still had that question, you can answer it now.
But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.
The fantastic grey area that was engineered over the past decade is "profiling", so my guess is the answer will be "we didn't use your customer data for training, but we cannot rule out that it has been used to create profiles of your customers to train our model"
Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.
You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).
Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).
In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.
So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.
The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.
And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.
So I'd say a whistleblower is pretty out of luck even becoming one.
You can just spin up deep research agents that ingest many sources at once to produce reports that don't replicate any one source too much. Since agents compare against sources they provide across-source analysis - what is the distribution of positions on this topic, is it debated or settled. Not truth, just summarizing, but I think this would be very useful for training.
Besides reporting on search sources you can also run the same queries on multiple LLMs closed book mode, and judge their distribution as well. It helps a lot if models are more aware of their knowledge holes. Scale it up for billions of topics if you have the pockets, the DR data is copyright free.
Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them.
But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
And require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.
Pretty much everything or at least a lot of what you use as an OpenAI (or Anthropic or whatever) customer was once someone elses product that just got appropriated by OpenAI.
> We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
I feel this is exactly what will happen as they cause all sites to go closed source to protect their intellectual property and the AI companies offer only biased information. They are replace the business on internet model by bankrupting everyone with their own tools.
This is predatory pricing under most antitrust laws (imho, not a lawyer)
and it is very easy to do when you dont need to pay for the raw material.
This is the business case already. And has been the case with tech companies for a long time. Your phones built in photo manager replaced a lot of what Photoshop does.
> If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
I mean how could you expect them to not given they've trained the existing models on effectively the sum total of all human knowledge available on the internet without regard to copyright/ownership of that material.
It's a little trite but this absolutely runs into the "Frog and the Scorpion", it is simply in their nature.
I mean this is a basic question, is your data used to improve models, did you opt in or not. There's nothing crazy here and it's not identifiable. You'll just conveniently find the next model iteration knows how to do it.
Any enterprise worth their salt already considers this stuff.
"which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him"
"Our aim was to see whether our system was also capable of this impressive feat"
"OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler"
For some reason I have a hard time believing people when they use language like this.
phrases along the lines of "I don't want you to take harm while trying to accuse us" is quite an "impressive feat".
Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this.
However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor:
1) the agents spin for days and produce too much output to review
2) using LLMs to process that output skips many important details
Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training.
Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this.
OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement).
And the difficulty is harder than just the extreme scale of text searching. but also explodes with organizational difficulty since there are so many people tweaking/shifting data independently upstream of the actual training run, and no they will not all add the telemetry you wish they did.
In the ideal, should it be this hard? Well, no, but that's org wrangling for you.
It feels convenient to not spend time on engineering around tooling that could be used to answer a question like “did you violate copyright by training on X?”
Don't attribute to malice what is better explained by coordination headwinds in extremely large companies.
The engineering around tooling wasn't remotely the issue. It's getting all the (thousands?) data researchers mostly iterating on fine tuning datasets that would get bristly if they couldn't work outside version control in a python notebook iteratively tweaking their dataset that processed and reprocessed a few datasets until a threshold was reached.
The only _guaranteed_ chains of custody are down at the compute job and file read level. Which in a massively distributed computing job is... [redacted] nodes reading [redacted] fanouts of "datasets" that is just an abstraction over [redacted] individual files.
There's no malice here. Just way way way more complex than you'd first think.
It definitely is solvable though. Data versioning is a thing and it can work quite transparently to the mutations done on the data.
To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal.
Not tracking data changesets like code changesets is certainly a choice, not really a constraint anymore. A similar choice I feel is implied by "extreme scale of text searching".
> no they will not all add the telemetry you wish they did
...is just a failure of the corporate policy surrounding data handling. Is git-for-data already considered telemetry?
Of course the truth is provenance of data is something best institutionally forgotten as quickly as possible. The only thing that matters is it's there, that the data has no history, and that's why it can be used in whatever way deemed necessary.
All of your points are valid, and believe me I was trying to make them. The problem is one of culture. Most of the people doing this kind of work didn't like version control, and their work was really just running notebooks (like iPython or Google Colab) until a number was good enough and they'd submit the file for inclusion into training runs.
You can call it a policy failure, but these people were in very high talent demand and so top down dictates would risk "X people leaving lab Y for lab Z" headlines and morale hits.
I am not saying this is good. I am telling you that on the ground it is so much messier than it should be.
The appearance of heroic efforts to get authoritative lists of what datasets went into which major model versions prevented actual data laundering (up to intent and mistakes). But don't attribute malice to that which is far far easier to explain with coordination headwinds: https://komoroske.com/slime-mold/
Malice is not required, which is precisely why I added, "in effect". If the effect is the same as data laundering, that is reason enough to encourage the practice, regardless of the original motives (and I'm sure there are plenty of legitimate ones).
I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced somehow
I don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.'
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
> Can you explain the difficulty in engineering a search apparatus over a corpus of text data?
My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
Especially because the data that gets fed into training is first anonymized, so they’d need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristan’s permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.
It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?
What surprises me is they're not more boldly/plainly lying about it.
How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
Unless they know exactly the researcher’s account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably don’t know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers.
I’m not saying they didn’t do anything unethical. I’m just saying even if they were ethical, there’s plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
> One can in hindsight see that our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
I don't know how much more clearly they can write:
> When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models.
One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality.
The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
ChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated:
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more
The "Learn more" link takes you to the link you've shared.
As I understand it, there is substantial question as to whether that actually stops them training on your data, it just perhaps changes what derivative processes are applied and used.
All I can say with certainty that the legal team at our typical Multi-Billion dollar Silicon Valley company had zero faith in any licensing arrangements with Anthropic or OpenAI, regardless of they $$$ involved, and that even getting to the point where Amazon Bedrock on Dedicated GPUs (we're already a big AWS customer - so definite cost advantages to dealing with them) - took 4-6 months of legal review before we could allow our engineers to start using Claude and OpenAI coding agents. Still can't use Fable because of their Data Retention requirements.
The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
Absolutely 100% not true. I have colleagues in 4 "FAANG adjacent" companies plus the one I work at - zero of them have any faith in Enterprise Agreements from either OpenAI or Anthropic.
There's a reason why people spend more $$$ with Data Bricks, Palantir, AWS Bedrock etc.. and don't even consider using Anthropic or OpenAI APIs directly - it's because those guarantees provide very little in the way of data-discovery, audit requirements, or liquidated damages should it ever be discovered there was data leakage.
At least with these other companies, while the LD is likewise not great (typically limited to the amount of money you paid them) - you at least have some data-governance guarantees around running on dedicated hardware - no multi-tenancy, no third-party access outside of the AWS operators who keep the HW running - but are very much not in the business of looking at your data.
I think this is mostly a function of what's at risk - when company valuations get into the 10s of billions of dollars, the risk of IP leaking into what could be seen as competitive companies (OpenAI/Anthropic would be happy to take over the world - I don't sense that AWS or Azure, are as ruthless in stepping on their customers business, unless of course they are a SAAS provider) is just too significant a liability to take - particularly when you can de-risk.
This is not true. They claim not to train by default for business and enterprise agreements, but for plus and pro plans they enable it by default and you can allegedly turn it off (I don’t trust them very much though, I’m sure there is something in the T&C saying they can modify that deal any time)
It sounds like OpenAI is trying to appease the author when they don’t have to by allowing him to rewrite their proof. They probably don’t believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
He's not "asking for a coauthor from Anthropic"; he already has a coauthor, who he's already been collaborating with, who happens to also be employed by Anthropic (but whose research in this area is not done as part of their employment at Anthropic).
Given that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
Here's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
It's too late. Sub models are deployed at every major organization in the United States and all it will take is turning off the option to improve the model for them to train directly on your own personal workflow, which CEOs will greedily eat up instantly if they can reduce labor costs. If they can brute force N-S, automating your finance or SWE job will be trivial. GG to most jobs connected to a computer in the next 5 years.
BTW, this was always the plan from day 1. You will pour all your training and experience into training the model and receive a pink slip as compensation.
If the work done is just "we made other people's work searchable without their consent" it's not quite the same as what they're implying in the marketing of "our model solved this problem".
Sam+Seb are struggling with their ideological allegiance. This amounts to a confession that there are no reseaechers, only research managers, left at OpenAI. Maybe they even know that they are losing credibility from their main investor(s). They desperately need a domain expert to salvage credibility.
They have no credibility with academia left, obviously, but their main competitor still does. No Millennium prize incoming, I'd wager. For openAI. Let's see mAth get political for once!!
One might be more certain that levent is now going to corner all the institutional support. Go go go!
For those of us who are into local models and preach it, we are called paranoid. I have often said this, if you are doing any real novel work, or putting your profitable business data/workflow into these models, you're a fool.
Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit.
Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.
Yup! I wanted to side with the mathematician on this one but I read the statement only to discover that they were also pushing an llm on someone else’s idea producing mountains of slop.
Have LLMs actually improved anything? Is mathematics better off than if these slop proofs didn’t exist? Who or what is actually benefiting here.
> Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems.
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.