← the record
WF-Q9UW6D

Alleged Inclusion of 12,000 Live API Keys in LLM Training Data Reportedly Poses Security Risks

A dataset used to train large language models allegedly contained 12,000 live API keys and authentication credentials. Some of these were reportedly still active and allowed unauthorized access. Truffle Security found these secrets in a December 2024 Common Crawl archive, which spans 250 billion web pages. The affected credentials could have been exploited for unauthorized data access, service disruptions, financial fraud, and a variety of other malicious uses.

Date it happened
2025-02-28
Organisation involved
Microsoft, OpenAI, Common Crawl, Microsoft Azure OpenAI Service
Product, system or model
Common Crawl dataset (December 2024 archive), Microsoft Copilot, Google Gemini, Anthropic Claude, ChatGPT, xAI Grok, DeepSeek, LLMs trained on compromised data
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AI Incident Database and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.