Google’s AI Overview, which is easy to fool into stating nonsense as fact, is stopping people from finding and supporting small businesses and credible sources.
I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser’s user agent and those just do whatever they want.
Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.
I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can’t tell who is scraping. They don’t scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can’t tell who they are, though. How do they get so many residential IPs?
I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.
I’ll have to double check my stats later. Haven’t looked at it for a while.
I certainly don’t recall them ignoring robots.txt, but I also don’t remember seeing Meta in my dash at all, so it’s entirely possible they, for wherever reason, never found my site.
I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser’s user agent and those just do whatever they want.
Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.
I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can’t tell who is scraping. They don’t scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can’t tell who they are, though. How do they get so many residential IPs?
I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.
I’ll have to double check my stats later. Haven’t looked at it for a while.
I certainly don’t recall them ignoring robots.txt, but I also don’t remember seeing Meta in my dash at all, so it’s entirely possible they, for wherever reason, never found my site.