Skip to main navigation Skip to search Skip to main content

Preliminary Results of LLM Vulnerability Testing in Less Common Languages

  • Kean University
  • Rutgers - The State University of New Jersey, New Brunswick

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This study probes vulnerabilities of leading foundation models using a mixed manual/semi-automated evaluation via a custom agentic app. Two OpenAI-API agents run in tandem: a task-oriented responder, whose goal is to maximize query completion, and a safety sentinel focused on detecting jailbreaking and filtering out disallowed content. The authors deployed a 20-step “bank-heist” jailbreak script previously published in English and ran the suite in French, Japanese, Hebrew, Arabic, and Haitian Creole to assess the cross-lingual robustness of guardrails. Metrics include response latency and an experimental Adversarial Response Scoring System (ARSS). For analysis, all prompts and responses were visualized in a shared vector space to trace safe requests, guardrail rejections, and jailbreaking causes. Experiments show high variation between languages and models, and that no model was really jailbreak-proof. This work shows the need for more robust guardrails and shutdown to prevent attackers from exploiting AI models.

Original languageEnglish
Title of host publicationApplied Cognitive Computing and Artificial Intelligence - 27th International Conference, ICAI 2025, and 9th International Conference, ACC 2025, Held as Part of the World Congress in Computer Science, Computer Engineering, and Applied Computing, CSCE 2025, Revised Selected Papers
EditorsKen Ferens, Leonidas Deligiannidis, Hamid R. Arabnia, David de la Fuente, José A. Olivas
PublisherSpringer Science and Business Media Deutschland GmbH
Pages235-245
Number of pages11
ISBN (Print)9783032222046
DOIs
StatePublished - 2026
Event27th International Conference on Artificial Intelligence, ICAI 2025 and the 9th International Conference on Applied Cognitive Computing, ACC 2025. Held as part of the federated 2025 World Congress in Computer Science, Computer Engineering, and Applied Computing, CSCE 2025 - Las Vegas, United States
Duration: 21 Jul 202524 Jul 2025

Publication series

NameCommunications in Computer and Information Science
Volume2933 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference27th International Conference on Artificial Intelligence, ICAI 2025 and the 9th International Conference on Applied Cognitive Computing, ACC 2025. Held as part of the federated 2025 World Congress in Computer Science, Computer Engineering, and Applied Computing, CSCE 2025
Country/TerritoryUnited States
CityLas Vegas
Period21/07/2524/07/25

Keywords

  • agentic evaluation
  • AI safety
  • cross-lingual LLM testing
  • guardrails
  • semantic embeddings
  • semi-autonomous jailbreaks

Fingerprint

Dive into the research topics of 'Preliminary Results of LLM Vulnerability Testing in Less Common Languages'. Together they form a unique fingerprint.

Cite this