TY - GEN
T1 - Preliminary Results of LLM Vulnerability Testing in Less Common Languages
AU - Kumar, Yulia
AU - Mardi, James
AU - Yang, Guohao
AU - Kruger, Dov
AU - Li, Juan
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - This study probes vulnerabilities of leading foundation models using a mixed manual/semi-automated evaluation via a custom agentic app. Two OpenAI-API agents run in tandem: a task-oriented responder, whose goal is to maximize query completion, and a safety sentinel focused on detecting jailbreaking and filtering out disallowed content. The authors deployed a 20-step “bank-heist” jailbreak script previously published in English and ran the suite in French, Japanese, Hebrew, Arabic, and Haitian Creole to assess the cross-lingual robustness of guardrails. Metrics include response latency and an experimental Adversarial Response Scoring System (ARSS). For analysis, all prompts and responses were visualized in a shared vector space to trace safe requests, guardrail rejections, and jailbreaking causes. Experiments show high variation between languages and models, and that no model was really jailbreak-proof. This work shows the need for more robust guardrails and shutdown to prevent attackers from exploiting AI models.
AB - This study probes vulnerabilities of leading foundation models using a mixed manual/semi-automated evaluation via a custom agentic app. Two OpenAI-API agents run in tandem: a task-oriented responder, whose goal is to maximize query completion, and a safety sentinel focused on detecting jailbreaking and filtering out disallowed content. The authors deployed a 20-step “bank-heist” jailbreak script previously published in English and ran the suite in French, Japanese, Hebrew, Arabic, and Haitian Creole to assess the cross-lingual robustness of guardrails. Metrics include response latency and an experimental Adversarial Response Scoring System (ARSS). For analysis, all prompts and responses were visualized in a shared vector space to trace safe requests, guardrail rejections, and jailbreaking causes. Experiments show high variation between languages and models, and that no model was really jailbreak-proof. This work shows the need for more robust guardrails and shutdown to prevent attackers from exploiting AI models.
KW - agentic evaluation
KW - AI safety
KW - cross-lingual LLM testing
KW - guardrails
KW - semantic embeddings
KW - semi-autonomous jailbreaks
UR - https://www.scopus.com/pages/publications/105040926449
U2 - 10.1007/978-3-032-22205-3_17
DO - 10.1007/978-3-032-22205-3_17
M3 - Conference contribution
AN - SCOPUS:105040926449
SN - 9783032222046
T3 - Communications in Computer and Information Science
SP - 235
EP - 245
BT - Applied Cognitive Computing and Artificial Intelligence - 27th International Conference, ICAI 2025, and 9th International Conference, ACC 2025, Held as Part of the World Congress in Computer Science, Computer Engineering, and Applied Computing, CSCE 2025, Revised Selected Papers
A2 - Ferens, Ken
A2 - Deligiannidis, Leonidas
A2 - Arabnia, Hamid R.
A2 - de la Fuente, David
A2 - Olivas, José A.
PB - Springer Science and Business Media Deutschland GmbH
T2 - 27th International Conference on Artificial Intelligence, ICAI 2025 and the 9th International Conference on Applied Cognitive Computing, ACC 2025. Held as part of the federated 2025 World Congress in Computer Science, Computer Engineering, and Applied Computing, CSCE 2025
Y2 - 21 July 2025 through 24 July 2025
ER -