Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm not a fan of Anubis for various reasons but the idea that bot traffic in only harmful with dynamic pages must die. CPU (yes, even to serve static pages) is not free, bandwidth is definitely not free. There's an idea that serving a static page to a bot has a marginal cost tending to zero, but it's never really zero and serving them by millions definitely has a cost.

Also, while some pages may look like static HTML pages, they may be generated on the fly by an expensive and/or slow backend, which adds to the cost. I happen to maintain servers for academics and some content management systems are slow and have an expensive CPU cost. While it's OK for the low number of humans interested in the subjects they deal with, it's definitely not fit for massive bot scrapping. And before you ask, no, it's not always practical to have cache upfront or to pre-generate all pages.



Anubis is not free either, it is a matter of how much it costs to run Anubis vs to let bots in.

When I see LKML using Anubis when the pages it serves are tens of kB, all presumably static, I wonder if they do it out of spite rather than to protect their servers.


It’s done out of necessity, you can read about why here: https://people.kernel.org/monsieuricon/creepy-crawlies

As a user/reader/viewer I absolutely hate Anubis and usually turn around when I see it pop up (at least on my phone where it takes ages to compute), but with stats like that, I get why a site operator would resort to using it.


I think the kernel.org post proves the parent point rather than contradicts it.

> At any one time, across 5 geo-distributed nodes, there are 14 CPU cores doing nothing but rendering git commits as html.

14 CPU cores total for running a website like kernel.org is laughable. This is not worth burning cycles in Anubis on client's devices, this is not worth the time of the engineer who worked on it. Provisioning more hardware would have been literally better for everyone.


I can’t say I really disagree, and as a visitor of the site that’s the solution I would prefer.


> But no, let's in fact choose the stupidest possible way of doing it — by rendering everything as HTML commit by commit and then parsing it.

This drives me crazy with so-called SOTA LLMs that have "achieved AGI".

Fable, Sol, Astra, will start by trying to reverse engineer a binary to figure out how something works when software is open source and one search query away.

You let them know it's open source, and they will start using github API instead of just cloning and grepping.


There are numerous services that will let you host static pages for free or nearly free. There are also numerous services that sit in front of your website that can block bots and reduce load on your origin server, many of which are also free, or very low cost relative to the service they provide.

The situation you are in is far less dire sounding when you consider that you have these options available to you.


Except that I don't have these options per employer policies.


So your employer is having a problem, and prevents you from using any of the available options to solve it? And you have asked them about all of them/told them about the problem? They don't like saving money?

Well, sounds like it's not your problem then.


> Also, while some pages may look like static HTML pages, they may be generated on the fly by an expensive and/or slow backend

Yes and that should be fixed before you subject real users to resource-wasting scripts.


Sadly, I do not control resource-wasting scripts running crawlers.

And to answer your intended demand: this is in some cases impossible or unreasonable. And things were working fine before LLM DDoS.


It can't always be fixed. Also, I don't use Anubis.


It's fixed enough for how much they want it to be fixed (I.e. how much they are being paid to fix it)


Please submit a fix to cgit then.


May I see it?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: