RESEARCH. CONNECTED.

Explore the scholarly record.

Discover prefix ownership, publishers, journals and DOI metadata. Understand the data behind every publication.

VOICE-AE: automated CTCAE scoring from ambient clinical audio using speech recognition and large language models

Federico Mastroleo, Mariana Borras-Osorio iD, Jonathan E. Moonen, Jake A. Jordan, Keldon K. Lin, Qian Liu, Mi Zhou, Brad J. Stish, Daniel J. Ma, Robert W. Mutter iD, Jann N. Sarkaria, Nadia N. Laack, Satomi Shiraishi iD, Andrew Y. K. Foong, David M. Routman, Mark R. Waddle iD

DOI10.1038/s41746-026-03373-z
PublisherSpringer Science and Business Media LLC
Journal / Sourcenpj Digital Medicine
Published2026-10-10
Metadata Deposited2026-10-10 (updated: 2026-10-10)
Subject—
Languageen
ISSN2398-6352
Typejournal-article
Volume / Issue / Pages— / — / —
Citations0
References deposited0
Access / license metadataOpen license identified License 1 ↗A reuse license does not by itself establish whether the full text is freely readable.

Abstract

Abstract Treatment-related adverse event documentation is labor-intensive and subject to inter-observer variability. We developed and evaluated VOICE-AE, a multi-stage pipeline integrating automatic speech recognition, speaker diarization and large language models (LLMs) for automated extraction and grading of Common Terminology Criteria for Adverse Events (CTCAE) toxicities from clinician-patient audio recordings. Encounters of prostate cancer patients undergoing radiotherapy were recorded and processed through VOICE-AE. Six blinded human raters independently graded 10 genitourinary/gastrointestinal toxicities from the same transcripts; hierarchical clustering identified an expert-rater subgroup that established the consensus reference standard. A total of 31 patients (39 encounters) were enrolled, yielding 390 encounter–toxicity assessments, of which 300 (76.9%) carried a reference grade of 0. Six-rater Krippendorff alpha was 0.677 for ordinal grading. Against reference standard, VOICE-AE achieved a mean weighted kappa of 0.717 and Krippendorff alpha of 0.669. Exact CTCAE grade match was 82.6% and within-one-grade accuracy 99.2%; because grade 0 predominated, a trivial classifier always predicting grade 0 would achieve 76.9% and 93.6% respectively. Performance was weaker at the grades that drive clinical action (grade 1 F1 0.58; grade 2 F1 0.57), and binary detection sensitivity was 77.8% with specificity 89.0% and precision 68.0%. In a like-for-like comparison restricted to the three expert-cluster raters, mean weighted kappa between VOICE-AE and those experts (0.666) fell just below agreement among the experts themselves (0.681). VOICE-AE performed within the range of inter-expert agreement, supporting the feasibility of automated, post-encounter AE documentation. Accuracy at grade ≥2 and detection of severe toxicity remain unproven.