Automatic Speech Recognition Performance in Psychiatric Speech: Linguistic, Clinical, and Architectural Factors
Automatic speech recognition (ASR) is essential for automated speech analysis pipelines in psychosis-spectrum disorders, yet performance on clinical speech remains understudied. We evaluated three modern ASR models (Whisper large-v2, Canary-1B, Parakeet-TDT-1.1B) on speech from 145 participants (92 psychosis-spectrum, 53 healthy controls) across seven tasks spanning structured reading to spontaneous speech, yielding 1,015 recordings. Performance assessed included word error rate (WER), semantic fidelity, disfluency preservation rate, and hallucination scores. The DISCOURSE corpus is publicly available through TalkBank for future ASR models. Whisper (WER: 12.1%) and Canary (12.2%) achieved comparable accuracy; Parakeet showed higher error rates (WER: 21.8%, p < 0.001). Task type was the strongest determinant of accuracy: structured reading (8.5%) outperformed spontaneous free speech (14.3%). Patients showed modestly elevated WER compared to controls, but significant group differences emerged only in open-ended narrative tasks. Sex influenced accuracy for Whisper (p = 0.039), with males showing higher error rates. Age was significantly associated with WER for Canary only (ρ = 0.245, p = 0.003). PANSS scores did not significantly predict WER for Whisper. Second formant bandwidth variability predicted WER (ρ = -0.324, p_FDR = 0.006). All models preserved few disfluency markers (Whisper 22.6% retained, Parakeet 5.3%), with patients showing significantly higher preservation rates than controls. This was acoustically mediated by loudness variability and was specific to disfluency preservation: loudness variability showed no mediation of general WER (all p > 0.31). Findings are limited to English-speaking populations; given that immigration status is a risk factor for psychosis, multilingual validation is essential. Whisper large-v2 is recommended for clinical speech research in psychosis-spectrum disorders based on its accuracy, robustness to symptom severity, and superior disfluency preservation.
Recommended citation: El-Mufti, N., Mackinley, M., Dzialoszynski, P., Palaniyappan, L., & Voppel, A. (under review). Automatic Speech Recognition Performance in Psychiatric Speech: Linguistic, Clinical, and Architectural Factors.
Download Paper

OpenReview