Cloudflare is taking a stand against AI website scrapers

  • 📰 engadget
  • ⏱ Reading Time:
  • 63 sec. here
  • 4 min. at publisher
  • 📊 Quality Score:
  • News: 32%
  • Publisher: 63%

Cloudflare News

Perplexity AI

Anna has been a freelance writer for more than a decade. In that time, she's covered everything from electronics to esports, from marketing to magic. Her tech and entertainment reporting has appeared on Ars Technica, Mashable, Digital Trends, and more. She especially loves playing, making, and geeking out over video games.

Cloudflare has released a new free tool that prevents AI companies' bots from scraping its clients' websites for content to train large language models. The cloud service provider is making this tool available to its entire customer base, including those on free plans."This feature will automatically be updated over time as we see new fingerprints of offending bots we identify as widely scraping the web for model training," the company said.

Cloudflare also identified the most active bots from the past year. The Bytedance-owned Bytespider bot attempted to access 40 percent of websites under Cloudflare's purview, andtried on 35 percent. They were half of the top four AI bot crawlers by number of requests on Cloudflare's network, along with Amazonbot and ClaudeBot.

It's proving very difficult to fully and consistently block AI bots from accessing content. The arms race to build models faster has led to instances of companies skirting or outright breaking the existing rules around blocking scrapers.of scraping websites without the required permissions. But having a backend company at the scale of Cloudflare getting serious about trying to put the kibosh on this behavior could lead to some results.

"We fear that some AI companies intent on circumventing rules to access content will persistently adapt to evade bot detection," the company said."We will continue to keep watch and add more bot blocks to our AI Scrapers and Crawlers rule and evolve our machine learning models to help keep the Internet a place where content creators can thrive and keep full control over which models their content is used to train or run inference on.

 

Thank you for your comment. Your comment will be published after being reviewed.
Please try again later.
We have summarized this news so that you can read it quickly. If you are interested in the news, you can read the full text here. Read more:

 /  🏆 276. in TECHNOLOGY

Technology Technology Latest News, Technology Technology Headlines