This news article is an AI-generated translation from Dutch to English.
Language is the weakest link in AI security. New research led by ML6's Cristóbal Sendin tested four widely used guardrail systems - Cisco AI Defense, AWS, Google Cloud, and Microsoft Azure, by firing hundreds of adversarial prompts at them in different languages. The findings expose a clear blind spot: guardrails perform noticeably worse outside English. A malicious prompt that gets blocked 97 times out of 100 in English slips through more often in Dutch or German, and even more frequently in less common languages.
The study also found that providers make deliberate trade-offs between strictness and usability. AWS, for instance, was notably more permissive, a choice the company confirmed reflects its philosophy of prioritizing a smooth developer experience over aggressive blocking.
Multi-turn attacks proved even harder to catch than single prompts. Cisco's Jan Heijdra notes that some models fail to stop up to 90% of multi-turn attacks, where attackers gradually steer a conversation toward harmful territory rather than stating their intent outright, an approach that's especially effective against open-weight models.
The research underscores a growing need for guardrails that account for both language diversity and conversational context, not just isolated prompts.
(Translation, original source: Dutch)
Read more in the original news article.