Skip to main navigation Skip to search Skip to main content

Benchmarking Automatic Speech Recognition for Aphasia: A Clinical Evaluation Framework

  • Kean University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

AI tools for speech therapy represent more than innovation; they signify necessity. With 89% of clinicians facing overwhelming caseloads and therapy wait times averaging 3-6 months, the demand for scalable rehabilitation support has never been higher. Although AI-driven communication platforms increasingly provide real-time feedback and sustained engagement, current Automatic Speech Recognition (ASR) systems still face significant limitations in accurately processing disordered speech such as aphasia. Aphasia is an acquired language disorder, the speech of individuals with aphasia is characterized by unintelligible words, jargon, or non-words. To add to this the speaker with aphasia may not recognize their errors and mostly have difficulty in comprehension. These limitations hinder equitable access to AI-driven rehabilitation tools. To bridge this gap, the first contribution of this study is evaluating four state-of-the-art ASR models, such as Whisper, NeMo-Conformer, Wav2Vec 2.0, and SpeechBrain, through the lens of Speech-Language Pathologist (SLP). The second contribution of the paper is utilizing a comprehensive benchmarking framework to assess how effectively these models capture clinically relevant aspects of aphasic speech, including lexical, syntactic, and fluency-related features. For evaluating the transcribed text, a combination of quantitative, qualitative (human-expert based), and linguistically grounded evaluation metrics is used, such as verb error rate, noun error rate, mean dependency length etc. of transcribed text.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
EditorsJuan Liu, Jingshan Huang, Xiaowo Wang, Fa Zhang, Xiufen Zou, Tian Tian, Xiaohua Hu, Bin Hu, Yi Xiong
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7599-7607
Number of pages9
ISBN (Electronic)9798331515577
DOIs
StatePublished - 2025
Event2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 - Wuhan, China
Duration: 15 Dec 202518 Dec 2025

Publication series

NameProceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025

Conference

Conference2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
Country/TerritoryChina
CityWuhan
Period15/12/2518/12/25

Keywords

  • Aphasia ASR model
  • Automatic Speech Recognition (ASR)
  • Facebook AI wave2vec 2.0
  • NVIDIA nemo
  • OpenAI whisper
  • SpeechBrain

Fingerprint

Dive into the research topics of 'Benchmarking Automatic Speech Recognition for Aphasia: A Clinical Evaluation Framework'. Together they form a unique fingerprint.

Cite this