OpenAI and Anthropic ignore robots.txt to scrape web content for training data
OpenAI and Anthropic have been found to be ignoring or circumventing the robots.txt rule that prevents automated scraping of websites, according to a person with knowledge of TollBit's analytics. The AI companies are bypassing blocks to their web crawlers GPTBot and ClaudeBot to retrieve all content from publishers' websites for model training. The practice undermines the long-standing web standard and raises concerns about copyright infringement.
- Organisation involved
- OpenAI, Anthropic
- Product, system or model
- ChatGPT, Claude
This incident was imported from businessinsider.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.