I’d link to some blog posts about this an example, but the site they’re from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host’s resources beyond being an ⊛ to webmasters?

Edit: Let me elaborate. A lot of the answers i’m seeing are just restating the problem without explaining why LLM scrapers are apparently all either coded by idiots or assholes. I know that they’re hitting sites with unreasonable numbers of requests, wasting bandwidth, and making tools like Anubis too important. I know that’s what’s happening. I’m asking why. What they gain from not being even a little intelligent about this.

With all the money and effort (and maybe even brainpower) going into this, surely there’s some explanation beyond incompetence.

  • dgdft@lemmy.world
    link
    fedilink
    English
    arrow-up
    27
    ·
    25 days ago

    You are sitting behind a receptionist’s desk with three coworkers. You need to check people into your office for visits.

    Because there are four of you, and people are generally good about queuing, your normal operations run smoothly.

    Now fifty people walk into the building at once. They don’t really care about checking in, or even visiting the office. They want to know if you have free coffee, a public bathroom, your elevator inspection on file, the date of your last fire inspection, and dozens more inane questions that aren’t what you’re used to handling. They refuse to wait in a line, they’re talking over one another, and if you take more than ten seconds to answer, they walk out of the building but come back to bother you a few minutes later.

    Does your office run well, and can you check in a legitimate visitor in a timely manner still?

    • Blue Label@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      24 days ago

      I would have said the 50+ new people would queue up in a line, but without giving the receptionist a toilet break. The receptionist would be so busy with them that the he can’t afford to leave the counter and ends up peeing on the spot, causing the building to be shut down due to hazardous materials.

  • emb@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    edit-2
    25 days ago

    Non-expert here, so conjecture warning.

    Loading the info from the page once is fine for training one LLM at one point in time. But there are a bunch of different companies and people doing that, and they probably keep doing it again because they want newer information.

    And training is only one aspect of AI scrapers. You also have agents constantly doing websearches and summarizing, synthesizing the info at user request, all day every day.

  • grue@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    25 days ago

    Everybody in this thread is talking about what they’re doing, but not a single reply has been able to explain why the bad behavior somehow benefits the companies doing it.

    I don’t believe the only reason is incompetence; there’s got to be somehing else to it.

    • Dran@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      25 days ago

      The unfortunate answer is because it’s cheaper to not give a shit. Sending a request and waiting for a timeout costs next to nothing, and scales linearly in terms of compute cost. The overwhelmed server on the other end slows exponentially with each concurrent request. The crawlers are set to maximize the efficiency of local resources, which include both wall-clock time and developer time. Why send one request at a time when your server can handle tens of thousands?

      Try x; wait 60 seconds, if fail: put on a list to try again later.

      Costs nothing to write and nothing to run. And if you own the hardware, and are paying for power already, may as well extract maximum dollar per watt.

    • Dr. Moose@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      24 days ago

      It’s 100% a mix of various incompetences. I develop anti fraud systems and deal with cralwers all the time. The web is actually really complex, probably the most complex technology in history of human kind and I’m not even kidding.

  • slazer2au@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    25 days ago

    Because a -r is likely what they are doing. Recursively checking everything every 3 microseconds

    • IndigoGolem@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      25 days ago

      So they’re just checking as fast as possible to get every update as fast as possible?

      This still sounds like idiot design, scraping hard enough to take down sites. It’s like cutting open the goose that lays lead eggs in the hope that you can get more lead and convince people that lead is better than gold.

  • givesomefucks@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    25 days ago

    As soon as a page loads, the AI goes to a new page. Not just one new page, every link on the page.

    It’s functionally a ddos attack. Because it quickly spirals exponentially. The limit isn’t how fast the AI can scrape, it’s how much bandwidth the website server has.

    • IndigoGolem@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      25 days ago

      But isn’t this how ordinary search crawlers and tools like wget also work? I’ve never heard of them causing these same problems.

      • givesomefucks@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        25 days ago

        Same difference as being hit by a golf cart rolling down a sleight slope and 100 Semi’s going 100mph

        The problem isn’t what they’re doing, it’s the speed and depth. There’s no concern for efficiency because they’re not paying for hardware and utilities.

        Everything AI is focused on doing asuch as possible as fast as possible, with the hope optimization will happen organically to the point it becomes profitable.

        But it won’t.

        • dual_sport_dork 🐧🗡️@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          25 days ago

          Correct. Google indexing your web page is rate limited for this reason, and you can even include a directive in your robots.txt to specify your own rate limit if you’d like the intervals to be longer (or shorter). The AI scrapers completely ignore your robots.txt. Except, I am certain, for abusing it as if it were a site map. Anything you list there is simply a target you’ve revealed to them. (“Hey, robots.txt says we shouldn’t crawl /foo/bar.html. That means there’s a page there! Let’s hammer it with 900 page load requests per second!”)

  • Dr. Moose@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    7
    ·
    edit-2
    24 days ago

    Honestly a lot of web is just extremely poorly optimized which is not a problem when you have have 1,000 page requests per day but suddenly you get 10,000 and things start breaking.

    Optimizing pages is very easy these days and a good website can serve hundreds of thousands of requests per day same as first 1,000 because caching is extremely powerful.

    However, same goes for crawlers which can be really poorly written and extremely unoptimized. This is especially the case with broad crawlers like AI crawlers that crawl all pages as they use a real web browser causing more expense. They use real browsers because websites these days use client side rendering to defer costs to the client AND anti-bot system that require a browser to bypass.

    Tl;dr: lots of very bad code

    Edit: leave it to luddites to down vote almost 30 years of web dev experience lol

    • sanzky@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      1
      ·
      24 days ago

      because you are clearly minimizing the impact of AI crawlers. sites don’t go from 1000 page requests /day to 10,000. I’ve seen sites going from 50~100 request per minute to 10.000 request per minute and they are mostly bots. no matter how efficient your site is, if the server cannot manage the connections it does not matter whether the content is cached or not.

      And caching does not solve everything because many cloud providers charge for it, so even if you could potentially deliver them, it suddenly becomes too expensive to do so.