Google's Perspective AI tricked by typos and leetspeak
Researchers at Aalto University and the University of Padua found that Google's Perspective AI hate speech detection system can be easily tricked by simple typos, adding spaces, or using leetspeak. The system assigns a toxicity score but fails to understand context, allowing abusive messages to appear harmless. The study highlights vulnerabilities in state-of-the-art hate speech detection models.
- Organisation involved
- Product, system or model
- Perspective
This incident was imported from mashable.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.