Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.
 help



Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.

You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).

Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).

In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.

So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.

The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.

And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.

So I'd say a whistleblower is pretty out of luck even becoming one.


You can just spin up deep research agents that ingest many sources at once to produce reports that don't replicate any one source too much. Since agents compare against sources they provide across-source analysis - what is the distribution of positions on this topic, is it debated or settled. Not truth, just summarizing, but I think this would be very useful for training.

Besides reporting on search sources you can also run the same queries on multiple LLMs closed book mode, and judge their distribution as well. It helps a lot if models are more aware of their knowledge holes. Scale it up for billions of topics if you have the pockets, the DR data is copyright free.


Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.

Billions lost, while waiting for their trillion ipo. Im sure they would manage...

When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them.

But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.


And require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: