I’d link to some blog posts about this an example, but the site they’re from went down a while ago.
At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host’s resources beyond being an ⊛ to webmasters?
Edit: Let me elaborate. A lot of the answers i’m seeing are just restating the problem without explaining why LLM scrapers are apparently all either coded by idiots or assholes. I know that they’re hitting sites with unreasonable numbers of requests, wasting bandwidth, and making tools like Anubis too important. I know that’s what’s happening. I’m asking why. What they gain from not being even a little intelligent about this.
With all the money and effort (and maybe even brainpower) going into this, surely there’s some explanation beyond incompetence.
You are sitting behind a receptionist’s desk with three coworkers. You need to check people into your office for visits.
Because there are four of you, and people are generally good about queuing, your normal operations run smoothly.
Now fifty people walk into the building at once. They don’t really care about checking in, or even visiting the office. They want to know if you have free coffee, a public bathroom, your elevator inspection on file, the date of your last fire inspection, and dozens more inane questions that aren’t what you’re used to handling. They refuse to wait in a line, they’re talking over one another, and if you take more than ten seconds to answer, they walk out of the building but come back to bother you a few minutes later.
Does your office run well, and can you check in a legitimate visitor in a timely manner still?
I would have said the 50+ new people would queue up in a line, but without giving the receptionist a toilet break. The receptionist would be so busy with them that the he can’t afford to leave the counter and ends up peeing on the spot, causing the building to be shut down due to hazardous materials.
Non-expert here, so conjecture warning.
Loading the info from the page once is fine for training one LLM at one point in time. But there are a bunch of different companies and people doing that, and they probably keep doing it again because they want newer information.
And training is only one aspect of AI scrapers. You also have agents constantly doing websearches and summarizing, synthesizing the info at user request, all day every day.
Everybody in this thread is talking about what they’re doing, but not a single reply has been able to explain why the bad behavior somehow benefits the companies doing it.
I don’t believe the only reason is incompetence; there’s got to be somehing else to it.
The unfortunate answer is because it’s cheaper to not give a shit. Sending a request and waiting for a timeout costs next to nothing, and scales linearly in terms of compute cost. The overwhelmed server on the other end slows exponentially with each concurrent request. The crawlers are set to maximize the efficiency of local resources, which include both wall-clock time and developer time. Why send one request at a time when your server can handle tens of thousands?
Try x; wait 60 seconds, if fail: put on a list to try again later.
Costs nothing to write and nothing to run. And if you own the hardware, and are paying for power already, may as well extract maximum dollar per watt.
It’s 100% a mix of various incompetences. I develop anti fraud systems and deal with cralwers all the time. The web is actually really complex, probably the most complex technology in history of human kind and I’m not even kidding.
Because a -r is likely what they are doing. Recursively checking everything every 3 microseconds
So they’re just checking as fast as possible to get every update as fast as possible?
This still sounds like idiot design, scraping hard enough to take down sites. It’s like cutting open the goose that lays lead eggs in the hope that you can get more lead and convince people that lead is better than gold.
Welcome to reason 2 to hate the current bubble.
As soon as a page loads, the AI goes to a new page. Not just one new page, every link on the page.
It’s functionally a ddos attack. Because it quickly spirals exponentially. The limit isn’t how fast the AI can scrape, it’s how much bandwidth the website server has.
But isn’t this how ordinary search crawlers and tools like wget also work? I’ve never heard of them causing these same problems.
Same difference as being hit by a golf cart rolling down a sleight slope and 100 Semi’s going 100mph
The problem isn’t what they’re doing, it’s the speed and depth. There’s no concern for efficiency because they’re not paying for hardware and utilities.
Everything AI is focused on doing asuch as possible as fast as possible, with the hope optimization will happen organically to the point it becomes profitable.
But it won’t.
Correct. Google indexing your web page is rate limited for this reason, and you can even include a directive in your robots.txt to specify your own rate limit if you’d like the intervals to be longer (or shorter). The AI scrapers completely ignore your robots.txt. Except, I am certain, for abusing it as if it were a site map. Anything you list there is simply a target you’ve revealed to them. (“Hey, robots.txt says we shouldn’t crawl /foo/bar.html. That means there’s a page there! Let’s hammer it with 900 page load requests per second!”)
This is not exactly true because web browsers have caching.
Honestly a lot of web is just extremely poorly optimized which is not a problem when you have have 1,000 page requests per day but suddenly you get 10,000 and things start breaking.
Optimizing pages is very easy these days and a good website can serve hundreds of thousands of requests per day same as first 1,000 because caching is extremely powerful.
However, same goes for crawlers which can be really poorly written and extremely unoptimized. This is especially the case with broad crawlers like AI crawlers that crawl all pages as they use a real web browser causing more expense. They use real browsers because websites these days use client side rendering to defer costs to the client AND anti-bot system that require a browser to bypass.
Tl;dr: lots of very bad code
Edit: leave it to luddites to down vote almost 30 years of web dev experience lol
because you are clearly minimizing the impact of AI crawlers. sites don’t go from 1000 page requests /day to 10,000. I’ve seen sites going from 50~100 request per minute to 10.000 request per minute and they are mostly bots. no matter how efficient your site is, if the server cannot manage the connections it does not matter whether the content is cached or not.
And caching does not solve everything because many cloud providers charge for it, so even if you could potentially deliver them, it suddenly becomes too expensive to do so.
Nosensical. I wrote an entire paragraph on ai crawlers being bad.


