DOI RECORD
VOICE-AE: automated CTCAE scoring from ambient clinical audio using speech recognition and large language models
Abstract
Abstract Treatment-related adverse event documentation is labor-intensive and subject to inter-observer variability. We developed and evaluated VOICE-AE, a multi-stage pipeline integrating automatic speech recognition, speaker diarization and large language models (LLMs) for automated extraction and grading of Common Terminology Criteria for Adverse Events (CTCAE) toxicities from clinician-patient audio recordings. Encounters of prostate cancer patients undergoing radiotherapy were recorded and processed through VOICE-AE. Six blinded human raters independently graded 10 genitourinary/gastrointestinal toxicities from the same transcripts; hierarchical clustering identified an expert-rater subgroup that established the consensus reference standard. A total of 31 patients (39 encounters) were enrolled, yielding 390 encounter–toxicity assessments, of which 300 (76.9%) carried a reference grade of 0. Six-rater Krippendorff alpha was 0.677 for ordinal grading. Against reference standard, VOICE-AE achieved a mean weighted kappa of 0.717 and Krippendorff alpha of 0.669. Exact CTCAE grade match was 82.6% and within-one-grade accuracy 99.2%; because grade 0 predominated, a trivial classifier always predicting grade 0 would achieve 76.9% and 93.6% respectively. Performance was weaker at the grades that drive clinical action (grade 1 F1 0.58; grade 2 F1 0.57), and binary detection sensitivity was 77.8% with specificity 89.0% and precision 68.0%. In a like-for-like comparison restricted to the three expert-cluster raters, mean weighted kappa between VOICE-AE and those experts (0.666) fell just below agreement among the experts themselves (0.681). VOICE-AE performed within the range of inter-expert agreement, supporting the feasibility of automated, post-encounter AE documentation. Accuracy at grade ≥2 and detection of severe toxicity remain unproven.
Go to Main Website