Hacker Newsnew | past | comments | ask | show | jobs | submit | SR2Z's commentslogin

> on the other hand the complete dismissal of copyright by AI labs

Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

The only thing they get in trouble for is pirating the works to get their hands on them.


> Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.

*USA only.

the UK has fair dealing, which is more restrictive

https://www.gov.uk/guidance/exceptions-to-copyright#fair-dea...

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...


"Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing. In particular, last I checked OpenAI and Microsoft are still badly threatened by the NYT lawsuit: https://law.justia.com/cases/federal/district-courts/new-yor... https://www.cnet.com/tech/services-and-software/publishers-o...

This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.


> "Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing.

yep

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.


Not to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

To be clear that was one of the few resolved cases where the judge agreed training was fair use. But the piracy was enough of a distraction that I don't consider that a particularly useful precedent. I am much more interested in the NYT case, which quite clearly shows GPT was trained on NYT articles and can spit them out verbatim (and has since been validated by academic research; all the commercial models are capable of mass plagiarism).

> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.

It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.


You're both reading tea leaves.

Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.

  In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.

Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.


> With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.

Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?


Perhaps Anthropic is indirectly paying to have their competitors sued...

My narritive is that the terms of the settlement would be full and final.

It was a class action, with payment going to authors and publishers, and the legal team will get paid too.

My guess is that funding is a major issue for the legal team. Authors presumably can't pay for lawyers unless a percentage of winnings, although publishers may have invested.

But the legal team will ask the beneficiaries to use some of the warchest to fund different campaigns against every other AI company. I would assume the legal team wants to win again. They've now got a good story to sell to rights holders, who presumably like money and don't like risks.

I haven't even got to my armchair yet this morning.


Man that's depressing to read someone defending this

This is a pretty common clause when a company open sources something that was previously locked away. The AOM AV1 codec has a similar rule for its patent pool to discourage trolling.

"You are violating the DRY principle with code like <blahblah>. Look through the codebase in depth, and identify possible places to consolidate common logic."

Yeah, there's no substitute for taste, but this is not that big of a deal. I infinitely prefer repeated code to crappy, leaky abstractions. Let the model generate some slop, then tighten it up either by hand or with more prompting.


Wasn't the famine a potato-specific blight?

Yes but the death toll was in large part because of Britain forcing the Irish to export everything that wasn’t a potato among a myriad of other oppressive practices.

And while I was taught repeatedly in school that all the Irish folks came to the US due to the potato famine, I don't ever recall anyone mentioning the forced exporation of other food to India, etc. Didn't learn it until I was like 50 years old. lol.

Have you ever met an actual Gen Z? They have no problem with swear words. Many of them love Key and Peele, whose humor is like 90% racist jokes.

If wokeness actually did capture a whole generation then why even bother complaining?


No but it's a useful shorthand to describe a type of bad writing.

I also think that people should focus on substance and not if AI was used, but AI writes like shit and I find myself retching a bit when I have to read long AI-written documents. Do they say something useful? Maybe, but when my eyes are glazing over because it's just so exhausting trying to parse what's written, I can't tell.

I certainly think less of people when they have such poor taste that they think writing like that is acceptable.


This is just an AI comment masquerading as not trying to prove a point.

Since you didn't address the substance, I guess we'll never know :)

I’m as pro AI as any.

Slop is slop. When the “it works but isn’t great” phrases end up slipping into a strong conceptual core, it compromises the perception of the ideas.

Perhaps our AI will cater to us by rewriting the content we read, and each of us intermediate all communication with systems that make that slop bearable.

Or perhaps, we learn that we kinda still need to give a shit when writing to land on the perception we’re trying to create within our readers


To be fair internet (or rather shovelware websites like Medium) were already flooded by crap articles based on a set of templates. Of course the issue is that LLMs are actually better than the robot-humans who used to write them so now it takes more time until you figure whether it’s worth reading or not..

Yeah, I'm not sure the level of trust extended to a company like Amazon or Google will also be extended to one run by Elon Musk, who is notorious for not respecting terms like this.

I don't know where you get this idea. A human being who used AI to generate something may actually claim copyright over the product.

You might be getting confused with a court case that ruled that AI could not have sole copyright, but that case just says that only a human being can hold copyright.


>I don't know where you get this idea.

I got the idea by reading the law,

https://www.congress.gov/crs-product/LSB10922

For example, the copyright is valid for X years after the author's death. The copyright transfers to the author's "widow or widower" or "surviving children or grandchildren". Does AI die? Who is the AI's widow? Who is the AI's children and grandchildren?

Oh, yeah, you're just wrong. The law explicity states that copyright belongs to HUMAN authors. Not AI. And courts have repeatedly ruled AI generated work cannot be copyrighted. You're the one who seems a little confused.


> Oh, yeah, you're just wrong. The law explicity states that copyright belongs to HUMAN authors. Not AI. And courts have repeatedly ruled AI generated work cannot be copyrighted. You're the one who seems a little confused.

No, I'm not confused. Go actually read your own link before you get condescending with me:

> Assuming that copyrightable works require a human author, works created by humans with the assistance of generative AI might be entitled to copyright protection depending on the nature of human involvement in the creative process.

The only settled legal finding is that copyright must be assigned to human beings.

The guidance of the Copyright Registration Office is not legally binding, which is where I think you may be confused. You do not need to register copyrights because they exist intrinsically when you create works.

If push comes to shove, you prove this in civil court to sue someone. They can argue a fair use defense. For example, training LLMs on copyrighted text has been repeatedly ruled fair use, as a transformative work.

The law is quite clear that how a work was created does not matter for copyright. All that matters is whether or not a thing qualifies as an original work of authorship. The boundaries of what AI-generated work can be copyrighted will be tested in court in the coming years.


Why are you calling this corporate surveillance? These cameras are installed on the government's orders for the government's use.

It's way more nefarious this way!


Because this isn't clearly against the law, nor should it be. If websites want to ban based on IP address lots of innocent users get caught in the cross-fire.

I'm not sure what the solution would look like - maybe Cloudflare's payment required for requests beyond a certain limit? But I think that the world needs user freedoms now more than ever.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: