← the record
AIAAIC-1892

DeepSeek accused of using OpenAI models to train AI system

Chinese AI startup DeepSeek has been accused by OpenAI of improperly using its proprietary models to train a competing AI system, potentially violating OpenAI's terms of service and prompting concerns about plagiarism, copyright, transparency, and accountability. What happened OpenAI alleges that DeepSeek used a technique called "distillation" to extract knowledge from OpenAI's models t hrough its API , which involves using outputs from larger AI models to train smaller ones . OpenAI says it has evidence to support its case. In August 2024, OpenAI and Microsoft investigated and blocked accounts for suspected terms of service violations, which they now believe were associated with DeepSeek. Distillation is known to be common in the AI industry. Why it happened The controversy appears to have stemmed Deepeek's desire to move quickly and at relatively low cost to develop and release its language models. The Chinese company is also seen to have used creative methods to circumvent US chip restrictions. What it means It is unclear whether DeepSeek will be held accountable for its alleged theft of OpenAI data. More broadly, the controversy highlights the challenges of protecting proprietary AI models and the data used to train them, and raises questions about the sustainability of high-cost, closed, general purpose AI models. Critics pointed out that OpenAI has also benefited from using others' data, and accused the US company of hypocrisy. System 🤖 DeepSeek R1 Developer: DeepSeek Artificial Intelligence Co Country: USA Sector: Technology Purpose: Train model Technology: Generative AI; Large lan guage model Issue: Accountability; Cheating/plagiarism; Copyright; Transparency

Date it happened
2025-01-01
Organisation involved
DeepSeek Artificial Intelligence Co
Product, system or model
DeepSeek R1
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.