Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.

Who are the native people?

 help



> So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?


To acquiesce to manufactured "social truths", to resign complacently to them being "modus operandi" means to place yourself outside of the society you live in.

You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.

Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.


Who knew the fix all along was just scolding people for being the wrong kind of frustrated

Who said it was besides you?

To take "being frustrated" as an end to personal responsibility and consequent action is defeatism.

When your previous attempts didn't work, change your approach. When you don't have one, you must find or invent one. When you don't know how, research that.

"Doing nothing because being frustrated" and scolding people who point out your circular reasoning does nothing besides guaranteeing defeat and demise.


What concrete action besides your comments here did you take? I am looking for inspiration.

I heard this in Worf's voice, with a long pause before "inspiration."

heavy metal guitar chords commence

[flagged]


I've been on my deathbed a few times over, and worse. Whatever life I left is a freebie for me.

Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read. Write your crawler, and crawl the web on your own if you want... and give it away!

You glaze over the fact, they stole the data to begin with and never recuperate their sources. That's not "reasonable".

They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.

You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.

So no, you're the unreasonable one here. So are they.


Search engines that obey robots.txt are stealing?

If someone does not want on a search engine index and they say not to index, and then get indexed anyhow is stealing.

But a “please read my content and make available to your users” then claim doing so is stealing seems a bit out there.

What am I missing?


There is a difference between "read my content and make it available to your users" and "make money by reading my content and making it available to your users". Terms of use like this are trying to say "we are the only ones who get to make money from this", and therein lies the problem.

To sharpen the distinction slightly (about increasingly abusive search-engines):

1. "You may make my content findable to your users as a service to them, taking a finder's-fee from ads etc."

2. "You may plagiarize my content, serving that to your users and cutting me out of the user-traffic almost entirely."


> we are the only ones who get to make money from this

Which, somewhat ironically, is the stance of copyright maximalists decrying “theft” be aggregators and AI while paying not a cent to the people whose shoulders they stand on.

It seems the complaint isn’t so much about trying to claim exclusive perpetual rights to an incremental creative layer built on millennia of history, but about the wrong people making the claim.


Well said. I'm sure a subset of the content that Cloudflare has indexed also has terms like Cloudflares, that they happily ignored.

Yes, search engine should be non-profits. But they must send users to profit from.

Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license. Irks me every time Dario at Anthropic calls for regulation to prevent distillation.

You do realize that US IP law does work that way, right?

You can have a copyright of a collection of facts even if you don't have a copyright of each individual fact in the collection.


You've mixed up your context here. Web pages are not generally "facts" under US copyright law, and redistributing them from a database doesn't turn them into facts, nor does it engage the normal "collection of facts" (aka database) rules about US copyright. And even if the webpages were "facts", the only part of "collections of facts" that gets copyright protection is the creative part. Which can't be the plain content of the facts.

Cite your source.

Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.

I never claimed the individual facts/data were given copyright to the collector. Let's stick to what I said if you are going to quibble with my comment.

[1] https://www.bitlaw.com/copyright/database.html


Here's the relevant portion of the original comment:

> Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license.

Here's what you said:

> You do realize that US IP law does work that way, right? You can have a copyright of a collection of facts even if you don't have a copyright of each individual fact in the collection.

Here's where you've gone wrong in those original comments:

- The original poster is talking about copyright-able works (in the context of this thread, those are websites).

- We're eliding "websites" into "facts" somehow - and as I've said, facts aren't copyrightable (https://www.copyright.gov/title17/92chap1.html#102) - and collections of facts (which are only have copyright over creative portions)

- And your first sentence is misleading as the poster is talking about the content of websites being treated as copyrighted by the collector, so we're already off track with the implication that the "fact" (website) is copyrighted when it's part of a database by the database constructor.

- I then proceeded to talk about how collections of facts interact here (ie: 'even supposing if' websites were somehow facts, which they generally aren't under copyright law).

- That's where the "only the creative parts of the database are copyright-able" (https://supreme.justia.com/cases/federal/us/499/340/, https://law.justia.com/cases/federal/district-courts/FSupp2/...) comes in. And again, refer back to the original comment, which is about the websites themselves, so we're kind of off track talking about creative parts of databases, but it is somewhat relevant due to comments further up talking mentioning the "relevance scores, or rankings" as something that the agreement prevents using freely. Those "relevance scores, or rankings" might be protected by copyright, though I suspect it would be a close call.

- Making a database of all websites is not creative (see the Feist Pubs., Inc. v. Rural Tel. Svc. Co., Inc., 499 U.S. 340 (1991) linked above, where a database of all phone numbers was determined to not be creative in itself). If the results of a particular search (ie: which results correspond to a particular search) are copyrighted hasn't been fully litigated, but Google themselves didn't claim their search results were copyrighted in the ongoing Google LLC v. SerpApi, LLC (4:25-cv-10826) litigation (https://storage.courtlistener.com/recap/gov.uscourts.cand.46..., III. B. 1 has a discussion).

Now to the most recent comment:

> Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.

You haven't distinguished it from my previous comment with a statement on what is copyrightable in collections of facts. If you believe that non-creative parts of collections are copyrightable, then examine the references above in this comment, and the section in what you've linked titled "Feist: Originality and Creativity". If you don't believe that, then you may want to be more clear.


You are mixing up law and morality.

Except web page content isn’t facts.

It is just like if someone scoured the internet for images and sold them in a book as a collection. They couldn’t do that without rights to those images.

Now imagine that the book “author” says you can’t copy the images from the book because they claim to own those images, without any copyright of any kind.

That is what they are trying to do. But with a different kind of copyrighted work.


No, the Search API isn't claiming it is taking over a copyright of the original crawled webpage. Quit trying to strawman this thread.

The Search API is acting as a database of a collection of webpages where the collection layer is the copyrightable thing. This is analogous to Google being able to copyright its search engine without having copyrights to all of the underlying content within the search corpus.

If your quibble is about the crawling / scraping, I'm not going to argue. It's easy for mass scraping to fall into unethical / illegal territory and I think that is the original sin of the AI foundation models.

Databases are copyrightable whether their content are facts (one example of content) or other data. I used the word facts because I knew from memory that was protected in US IP law.

Citation: https://www.bitlaw.com/copyright/database.html


Your citation undermines your point: "U.S. Supreme Court ruled that a compilation work such as a database must contain a minimum level of creativity in order to be protectable under the Copyright Act." https://www.bitlaw.com/copyright/database.html#Feist

A full copy of the indexable web isn't a creative work at all. It isn't a collection that has been curated, it is just the whole thing.


There’s a huge body of statute and case law about the differences between aggregations / databases and their contents. See: https://www.justia.com/intellectual-property/copyright/lists...

Like it or not (I’m in “not”), this is hardly a new thing.


It's probably not illegal per se. But if you violate their ToS they can just stop providing you services.

The data is the moat, the compute is a cost center.

> Who are the native people?

Geocities and vBulletin users!


> Who are the native people?

All of us!


“You're looking at 'em...”

- Tony Soprano


It's not weird, it's normal business attitude. If something brings you money, do it. If it loses you money, don't do it. "Morals" and "ethics" are for suckers, good capitalists only consider the likely consequences of their actions - realistically likely, not what's likely in an ideal world. You cannot become a ten-millionaire without thinking like this.

They don't forbid saving data because it's immoral, they forbid it because they want your dependence on them and your continued monetary transfers.


What if someone is offered money to kill you?

If they're a good capitalist and they think they can keep the money without negative consequences, they should take the money.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: