Chinese large language model thinks it is ChatGPT
A new large language model developed by the Chinese start-up DeepSeek gained attention for its impressive performance, as well as its peculiar behaviour of identifying itself as ChatGPT, suggesting it had been trained on OpenAI's product. What happened DeepSeek V3 has been reported to outperform established models like OpenAI's GPT-4 and Meta's Llama 3 on a number of benchmarks. The model features 671 billion parameters and was trained at a notably low cost of approximately USD 5.58 million over two months. However, it has also been observed that DeepSeek V3 sometimes claims to be a version of ChatGPT . Why it happened The phenomenon of DeepSeek V3 identifying itself as ChatGPT may stem from the training data used for its development, with commentators speculating that the system may have been trained on datasets that include outputs from ChatGPT. Deepseek has not explained the behaviour of its system. What it means Deepseek V3 may perform strongly, but the start-up has been noticeably reluctant to discuss the sources of its training data, prompting questions about its ethics and integrity. With AI models increasingly mimicing one another, understanding their unique characteristics and origins will become important differentators.
- Date it happened
- 2024-01-01
- Product, system or model
- DeepSeek V3
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.