PubMed2026
BACKGROUND: Radiation pneumonitis (RP) remains a significant treatment-related toxicity in patients with unresectable, locally advanced non-small cell lung cancer (NSCLC) undergoing chemoradiotherapy (CRT). Most existing predictive models rely on static baseline demographic or dosimetry variables and lack real-time clinical applicability. We developed a novel predictive framework that integrates longitudinal symptom data extracted from clinical notes using natural language processing (NLP) with clinical and dosimetry features to improve early RP prediction. METHODS: We retrospectively identified 227 patients with locally advanced NSCLC treated with definitive CRT at a high-volume cancer center in the United States. We included all patients older than 18 years who were diagnosed between Jan 1, 2006, and Dec 31, 2022 with histologically or cytologically confirmed unresectable Stage 2 or 3 NSCLC and treated with conformal radiotherapy to a minimum dose of ≥45 Gy with or without chemotherapy. Of these, 31 RP events were identified through manual adjudication using radiologic criteria and chart review. NLP was used to extract the temporal relationship of 16 pre-specified symptoms with treatment from over 100,000 clinical notes spanning pre- and during-treatment intervals. We trained and validated machine learning models on combinations of baseline clinical data, radiation dosimetry, and NLP-derived symptom features. Model performance was evaluated using a nested cross-validation framework, with an outer cross-validation loop reserved for performance assessment and an inner cross-validation loop used for model training and integration, and summarized using area under the receiver operating characteristic curve (AUC) and partial AUC (pAUC) at high specificity thresholds. Clinical utility was evaluated using decision curve analysis (DCA). FINDINGS: The best-performing model incorporated longitudinal NLP features and achieved a median AUC of 0.759 (90% confidence interval 0.753-0.766), significantly outperforming baseline models using only dosimetry (AUC 0.613) or clinical variables (AUC 0.635). NLP-based features such as cough trajectory, shortness of breath, and wheezing were among the most important predictors. Inclusion of NLP-derived symptom data improved early identification of high-risk patients, particularly in the clinically relevant high-specificity range (pAUC 0.021 vs. 0.010 for dosimetry alone). DCA showed that the calibrated MLP model provided greater net benefit than default strategies of treating all or no patients across clinically relevant threshold possibilities. INTERPRETATION: In this early work, NLP-based extraction of longitudinal symptoms from routine clinical documentation meaningfully enhances RP prediction in patients undergoing CRT for NSCLC. This approach leverages existing electronic health record infrastructure to deliver real-time, scalable, and interpretable risk estimates, offering a pathway toward potential early intervention and personalized toxicity management. The model and DCA requires external and prospective validation before clinical deployment; as such, future work should focus on this validation and integration into clinical decision support systems. FUNDING: AstraZeneca.