← the record
AIAAIC-2104

AI creates error-plagued Wikipedia articles in obscure languages

Machine translation tools powered by AI are flooding Wikipedia with inaccurate articles in minority and obscure languages, compromising the reliability of content on the platform and undermining linguistic and cultural integrity. What happened Beginning around 2012, the Swedish-developed bot Lsjbot began mass-producing articles for Wikipedia editions in smaller languages such as Cebuano and Waray-Waray (Philippines) and later contributed to Swedish Wikipedia. These bots automatically generated short encyclopedia entries, often on geographic or biological topics, using data scraped from public databases. More recently, AI translation tools and large language models have produced similarly flawed content in low-resource languages such as Greenlandic, Yoruba, and Swahili, where linguistic complexity and lack of training data cause frequent factual and grammatical errors. The result has been millions of stilted, error-prone articles, some of which contain mistranslations or factual distortions, creating misleading records and crowding out human-authored, culturally informed material. Why it happened This occurred because Wikipedia’s open editing model and enthusiasm for rapid content growth created an incentive to automate article creation, especially for languages with few human editors. Bots and AI systems were deployed without robust community oversight or linguistic quality control. Transparency about the extent and nature of bot-generated content has been limited, and accountability mechanisms within Wikipedia’s decentralised governance have been weak. Additionally, AI companies and Wikipedia maintainers have not adequately disclosed model limitations or biases, allowing poorly trained systems to produce large volumes of low-quality text in languages they do not “understand.” What it means For the communities directly affected - particularly speakers of small, indigenous, or underrepresented languages - the proliferation of flawed AI-generated content risks misrepresenting their language, history, and culture, eroding trust in local-language Wikipedia editions. It can also discourage genuine human participation, as editors face an overwhelming volume of machine text to correct. For society at large, the problem highlights the fragility of online knowledge ecosystems when automated systems outpace human oversight, raising questions about AI accountability, data ethics, and linguistic equity in digital information spaces. System 🤖 Content Translate Lsjbot Developer: Sverker Johansson ; Wikipedia Country: Canada; Greenland; Kenya; New Zealand; Nigeria ; Phi li ppines; Spain; Sweden; Tanzania; Uganda; Zanzibar Sector: Media/entertainment/sports/arts Purpose: Tran slate articles Technology: Bot/int elligent agent; Machine learning Issue: Accountability; Accuracy/reliability; Mis/disinfor mation; Representation

Date it happened
2012-01-01
Organisation involved
Wikipedia
Product, system or model
Content Translate; Lsjbot
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.