TY - GEN
T1 - Benchmarking Automatic Speech Recognition for Aphasia
T2 - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
AU - Kollapally, Navya Martin
AU - Akes, Christa
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - AI tools for speech therapy represent more than innovation; they signify necessity. With 89% of clinicians facing overwhelming caseloads and therapy wait times averaging 3-6 months, the demand for scalable rehabilitation support has never been higher. Although AI-driven communication platforms increasingly provide real-time feedback and sustained engagement, current Automatic Speech Recognition (ASR) systems still face significant limitations in accurately processing disordered speech such as aphasia. Aphasia is an acquired language disorder, the speech of individuals with aphasia is characterized by unintelligible words, jargon, or non-words. To add to this the speaker with aphasia may not recognize their errors and mostly have difficulty in comprehension. These limitations hinder equitable access to AI-driven rehabilitation tools. To bridge this gap, the first contribution of this study is evaluating four state-of-the-art ASR models, such as Whisper, NeMo-Conformer, Wav2Vec 2.0, and SpeechBrain, through the lens of Speech-Language Pathologist (SLP). The second contribution of the paper is utilizing a comprehensive benchmarking framework to assess how effectively these models capture clinically relevant aspects of aphasic speech, including lexical, syntactic, and fluency-related features. For evaluating the transcribed text, a combination of quantitative, qualitative (human-expert based), and linguistically grounded evaluation metrics is used, such as verb error rate, noun error rate, mean dependency length etc. of transcribed text.
AB - AI tools for speech therapy represent more than innovation; they signify necessity. With 89% of clinicians facing overwhelming caseloads and therapy wait times averaging 3-6 months, the demand for scalable rehabilitation support has never been higher. Although AI-driven communication platforms increasingly provide real-time feedback and sustained engagement, current Automatic Speech Recognition (ASR) systems still face significant limitations in accurately processing disordered speech such as aphasia. Aphasia is an acquired language disorder, the speech of individuals with aphasia is characterized by unintelligible words, jargon, or non-words. To add to this the speaker with aphasia may not recognize their errors and mostly have difficulty in comprehension. These limitations hinder equitable access to AI-driven rehabilitation tools. To bridge this gap, the first contribution of this study is evaluating four state-of-the-art ASR models, such as Whisper, NeMo-Conformer, Wav2Vec 2.0, and SpeechBrain, through the lens of Speech-Language Pathologist (SLP). The second contribution of the paper is utilizing a comprehensive benchmarking framework to assess how effectively these models capture clinically relevant aspects of aphasic speech, including lexical, syntactic, and fluency-related features. For evaluating the transcribed text, a combination of quantitative, qualitative (human-expert based), and linguistically grounded evaluation metrics is used, such as verb error rate, noun error rate, mean dependency length etc. of transcribed text.
KW - Aphasia ASR model
KW - Automatic Speech Recognition (ASR)
KW - Facebook AI wave2vec 2.0
KW - NVIDIA nemo
KW - OpenAI whisper
KW - SpeechBrain
UR - https://www.scopus.com/pages/publications/105033561636
U2 - 10.1109/BIBM66473.2025.11356945
DO - 10.1109/BIBM66473.2025.11356945
M3 - Conference contribution
AN - SCOPUS:105033561636
T3 - Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
SP - 7599
EP - 7607
BT - Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
A2 - Liu, Juan
A2 - Huang, Jingshan
A2 - Wang, Xiaowo
A2 - Zhang, Fa
A2 - Zou, Xiufen
A2 - Tian, Tian
A2 - Hu, Xiaohua
A2 - Hu, Bin
A2 - Xiong, Yi
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 15 December 2025 through 18 December 2025
ER -