PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning↗

Machine learning for an expert system to predict preterm birth risk.

OBJECTIVE: Develop a prototype expert system for preterm birth risk assessment of pregnant women. Normal gestation involves a term of 40 weeks, but because 8-12% of the newborns in the United States are delivered prior to 37 weeks' gestation, problems associated with prematurity continue to plague individuals, families, and the health care system. DESIGN: A knowledge-base development methodology used machine learning, statistical analysis, and validation techniques to analyze three large datasets (18,890 subjects and 214 variables). The dependent (i.e., decision) variable studied was weeks of gestation at delivery, with dichotomous coding of preterm delivery (prior to 37 weeks) and full-term delivery (37+ weeks). RESULTS: Machine learning with a program named Learning from Examples using Rough Sets (LERS) induced 520 usable rules that were entered into a prototype expert system. The prototype expert system was 53-88% accurate in predicting preterm delivery for 9,419 patients. CONCLUSION: The prototype expert system was more accurate than traditional manual techniques in predicting preterm birth.

Adult↗

miRNA Target Prediction: An Overview of the Past and Current Tools.

MicroRNAs (miRNAs) are among the most studied molecules in recent years, and since their discovery, many miRNAs have been identified across various species. As members of the non-coding RNA family, miRNAs are key players in post-transcriptional gene regulation. These molecules can inhibit translation or promote degradation of messenger RNA (mRNA) by binding to the 3' untranslated region (UTR) of mRNA, thereby influencing almost all biological processes. To identify a miRNA's biological role, it is essential to predict the target sites to which it binds, a goal made possible through bioinformatics tools. This chapter discusses the bioinformatics tools commonly used for this purpose. Also, it analyzes the main factors considered in target prediction, such as seed match, free energy, conservation, site accessibility, multiple binding site contribution, and machine learning and deep learning approaches. Understanding the principles underlying these predictive methodologies is crucial for advancing one's biological research on miRNAs.

MicroRNAs↗

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans↗

How advances in machine learning drive early detection and risk prediction of early-onset colorectal cancer.

Early-onset colorectal cancer (EOCRC), defined as colorectal cancer diagnosed before age 50, is rising across high- and middle-income settings whilst organised screening stays anchored to older age thresholds. Blood-based liquid biopsy, combined with machine learning, is the most plausible route to early detection in this group because it does not depend on bowel preparation, endoscopy capacity, or adherence to stool-based testing. The gap is structural: incidence climbs fastest in the population below the age at which any guideline-endorsed modality is offered. The analytical challenge is that early-stage tumour-derived signals in plasma are low in abundance and distributed across heterogeneous molecular layers: circulating tumour DNA mutations, aberrant methylation, cfDNA fragmentomics, and small non-coding RNA. Machine learning converts these into a single calibrated probability. This review examines where artificial intelligence (AI)-driven liquid biopsy genuinely adds diagnostic value in EOCRC, distinguishes components in which learned models are decorative from those in which they are mechanistically necessary, and identifies the validation deficit separating research cohorts from deployable clinical tools. It summarises the first-generation tools used clinically for early detection and post-treatment monitoring, then considers analytes from exosome-bound microRNAs to long-read whole-genome sequencing of circulating plasma DNA, which reads cytosine modification natively, resolves methylation and fragmentation on single molecules, and characterises structural events short reads cannot anchor. Any analyte can feed a learned model, but more diverse input yields better discrimination. The central argument is that approved, guideline-included blood tests were validated in populations aged 45 and above, and their performance in younger patients cannot be assumed.

cfDNA fragmentomics↗

An in-silico method for prediction of polyadenylation signals in human sequences.

This paper presents a machine learning method to predict polyadenylation signals (PASes) in human DNA and mRNA sequences by analysing features around them. This method consists of three sequential steps of feature manipulation: generation, selection and integration of features. In the first step, new features are generated using k-gram nucleotide acid or amino acid patterns. In the second step, a number of important features are selected by an entropy-based algorithm. In the third step, support vector machines are employed to recognize true PASes from a large number of candidates. Our study shows that true PASes in DNA and mRNA sequences can be characterized by different features, and also shows that both upstream and downstream sequence elements are important for recognizing PASes from DNA sequences. We tested our method on several public data sets as well as our own extracted data sets. In most cases, we achieved better validation results than those reported previously on the same data sets. The important motifs observed are highly consistent with those reported in literature.

Base Sequence↗

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)↗

In silico prediction method for plant Nucleotide-binding leucine-rich repeat- and pathogen effector interactions.

Plant Nucleotide-binding leucine-rich repeat (NLR) proteins play a crucial role in effector recognition and activation of Effector triggered immunity following pathogen infection. Genome sequencing advancements have led to the identification of a myriad of NLRs in numerous agriculturally important plant species. However, deciphering which NLRs recognize specific pathogen effectors remains challenging. Predicting NLR-effector interactions in silico will provide a more targeted approach for experimental validation, critical for elucidating function, and advancing our understanding of NLR-triggered immunity. In this study, NLR-effector protein complex structures were predicted using AlphaFold2-Multimer for all experimentally validated NLR-effector interactions reported in literature. Binding affinities- and energies were predicted using 97 machine learning models from Area-Affinity. We show that AlphaFold2-Multimer predicted structures have acceptable accuracy and can be used to investigate NLR-effector interactions in silico. Binding affinities for 58 NLR-effector complexes ranged between -8.5 and -10.6 log(K), and binding energies between -11.8 and -14.4&#x2009;kcal/mol-1, depending on the Area-Affinity model used. For 2427 "forced" NLR-effector complexes, these estimates showed larger variability, enabling identification of novel NLR-effector interactions with 99% accuracy using an Ensemble machine learning model. The narrow range of binding energies- and affinities for "true" interactions suggest a specific change in Gibbs free energy, and thus conformational change, is required for NLR activation. This is the first study to provide a method for predicting NLR-effector interactions, applicable to all pathosystems. Finally, the NLR-Effector Interaction Classification (NEIC) resource can streamline research efforts by identifying NLRs important for plant-pathogen resistance, advancing our understanding of plant immunity.

Plant Proteins↗

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization↗

Prematurity and Genetic Liability for Autism Spectrum Disorder.

BACKGROUND: Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by diverse presentations and a strong genetic component. Environmental factors, such as prematurity, have also been linked to increased liability for ASD, though the interaction between genetic predisposition and prematurity remains unclear. This study aims to investigate the impact of genetic liability and preterm birth on ASD conditions. METHODS: We analyzed phenotype and genetic data from two large ASD cohorts, the Simons Foundation Powering Autism Research for Knowledge (SPARK) and Simons Simplex Collection (SSC), encompassing 78,559 individuals for phenotype analysis, 12,519 individuals with genome sequencing data, and 8,104 individuals with exome sequencing data. Statistical significance of differences in clinical measures was evaluated between individuals with different ASD and preterm status. We assessed the rare variants burden using generalized estimating equations (GEE) models and polygenic load using ASD-associated polygenic risk score (PRS). Furthermore, we developed a machine learning model to predict ASD in preterm children using phenotype and genetic features available at birth. RESULTS: Individuals with both preterm birth and ASD exhibit more severe phenotypic outcomes despite similar levels of genetic liability for ASD across the term and preterm groups. Notably, preterm ASD individuals showed an elevated rate of de novo variants identified in exome sequencing (GEE model, p=0.005) in comparison to the non-ASD preterm group. Additionally, a GEE model showed that a higher ASD PRS, preterm birth, and male sex were positively associated with a higher predicted probability for ASD, reaching a probability close to 90% in SPARK. Lastly, we developed a machine learning model using phenotype and genetic features available at birth with limited predictive power (AUROC = 0.65). CONCLUSIONS: Preterm birth may exacerbate the multimorbidity present in ASD, which was not due to the ASD genetic factors. However, increased genetic factors may elevate the likelihood of a preterm child being diagnosed with ASD. Additionally, a polygenic load of ASD-associated variants had an additive role with preterm birth in the predicted probability for ASD, especially for boys. We propose that incorporating genetic assessment into neonatal care could benefit early ASD identification and intervention for preterm infants.

Autism Spectrum Disorder↗

Predicting 5-Year Mortality in Non-Small-Cell Lung Cancer Using the Korean Central Cancer Registry: Model Development and Validation Study.

BACKGROUND: Non-small-cell lung cancer (NSCLC) is one of the most common cancers and a leading cause of cancer-related mortality, making prognostic prediction clinically essential. Machine learning models are increasingly used to assess prognosis; however, developing systems that combine high discrimination with clear, clinically interpretable reasoning remains challenging. OBJECTIVE: This study aimed to develop deep learning models that predict 5-year mortality in NSCLC using data from the Korea Central Cancer Registry and quantify feature importance through permutation testing. METHODS: We identified 3144 patients diagnosed between 2014 and 2017 who had complete clinical data, pulmonary function test results, histological information, genomic data, and staging details. After preprocessing, the cohort was divided into stratified training, validation, and test sets in a 70%-15%-15% ratio. Five models were tuned using Hyperband across 10 predefined feature groups. The primary evaluation metric was the area under the receiver operating characteristic curve (AUC); additional metrics included accuracy, F1-score, precision, and recall. Groupwise permutation importance was calculated for each model, and the concordance of importance rankings was assessed using the Friedman test. RESULTS: All 5 models yielded comparable discrimination values on the test set (AUC=0.875-0.879). Model A was selected as the primary model and achieved an AUC of 0.879, an accuracy of 0.806, an F1-score of 0.824, and a Brier score of 0.142. Permuting the stage resulted in the largest decrease in AUC (0.217), followed by the pulmonary function test (0.016). Gene mutation had a modest overall impact but became more influential within the adenocarcinoma subset. The Friedman test showed no statistically significant differences in importance rankings across the models (P=.93). CONCLUSIONS: A grouped-input deep learning framework achieved discrimination comparable to a conventional Cox proportional hazards model using the same routine clinical variables for 5-year mortality prediction in NSCLC. Group-level permutation importance provided stable and reproducible insights into the clinical factors influencing risk, which may guide future model refinement and clinical decision-making.

Humans↗

Complex hybrid models combining deterministic and machine learning components for numerical climate modeling and weather prediction.

A new practical application of neural network (NN) techniques to environmental numerical modeling has been developed. Namely, a new type of numerical model, a complex hybrid environmental model based on a synergetic combination of deterministic and machine learning model components, has been introduced. Conceptual and practical possibilities of developing hybrid models are discussed in this paper for applications to climate modeling and weather prediction. The approach presented here uses NN as a statistical or machine learning technique to develop highly accurate and fast emulations for time consuming model physics components (model physics parameterizations). The NN emulations of the most time consuming model physics components, short and long wave radiation parameterizations or full model radiation, presented in this paper are combined with the remaining deterministic components (like model dynamics) of the original complex environmental model--a general circulation model or global climate model (GCM)--to constitute a hybrid GCM (HGCM). The parallel GCM and HGCM simulations produce very similar results but HGCM is significantly faster. The speed-up of model calculations opens the opportunity for model improvement. Examples of developed HGCMs illustrate the feasibility and efficiency of the new approach for modeling complex multidimensional interdisciplinary systems.

Artificial Intelligence↗

Decision-tree approach to the immunophenotype-based prognosis of the B-cell chronic lymphocytic leukemia.

Use of a nonlinear prediction method, such as machine learning, is a valuable choice in predicting progression rate of disease when applied to the highly variable and correlated biological data such as those in patients with chronic lymphocytic leukemia (CLL). In this work, decision-tree approach to cell phenotype-based prognosis of CLL was adopted. The panel of 33 (32 different phenotypic features and serum concentration of sCD23) parameters was simultaneously presented to the C4.5 decision tree which extracted the most informative of them and subsequently performed classification of CLL patients against the modified Rai staging system. It has been shown that substantial correlation between the percentage of expression of the CD23 molecule on CD19+ B-cells, the level of sCD23, the percentage of CD45RA+, and the absolute number of CD4CD45RA+RO+ T-cells and the clinical stages, exists. The prediction vector, composed of their concatenated values, was able to correctly associate 83% of the cases in the low-risk group (Rai stage 0), 100% of the cases in the intermediate-risk group (Rai stage I and II), and 89% of the cases in the high-risk group (Rai stage III and IV) of CLL patients. Predictivity of this vector was 100%, 95%, and 89%, respectively. In conclusion, from the described analysis, it may be inferred that two processes play important roles in the progression rate of CLL: 1.deregulated function of the CD23 gene in B-cells accompanied by the appearance of its cleaved product sCD23 in the sera; and 2. functionally impaired and imbalanced CD4 T-cell subpopulations found in the peripheral blood of CLL patients.

Aged↗

Unveiling m7G modification patterns and causal drivers governing intracranial aneurysm rupture risk through multi-omics validation and m7G-MeRIP-seq profiling.

Intracranial aneurysm (IA) rupture causes severe brain hemorrhage with high mortality, yet its molecular drivers remain unclear and better risk prediction is urgently needed. Using transcriptomics, single-cell analysis, and genetic data, we investigated the role of N7-methylguanosine (m7G) RNA modification in IA. We identified distinct m7G modification patterns, validated their methylation features in patient samples, and incorporated these patterns into a machine learning-based rupture prediction model. The presence and characteristics of m7G patterns significantly improved model performance, achieving high predictive accuracy across three independent cohorts (AUC 0.91-0.95). Genetic analyses further identified three causal m7G-related genes (NSUN2, IFIT5, SNUPN), and laboratory experiments confirmed their altered expression and methylation in ruptured aneurysms. Overall, our findings demonstrate that m7G modifications play a key role in IA rupture. The validated prediction model offers strong clinical potential for rupture risk assessment, and the identified genes represent promising therapeutic targets.

Humans↗

A probabilistic learning approach to whole-genome operon prediction.

We present a computational approach to predicting operons in the genomes of prokaryotic organisms. Our approach uses machine learning methods to induce predictive models for this task from a rich variety of data types including sequence data, gene expression data, and functional annotations associated with genes. We use multiple learned models that individually predict promoters, terminators and operons themselves. A key part of our approach is a dynamic programming method that uses our predictions to map every known and putative gene in a given genome into its most probable operon. We evaluate our approach using data from the E. coli K-12 genome.

Gene Expression Profiling↗

Prognostic significance of DNA damage response-related markers in esophageal squamous cell carcinoma using machine learning approaches.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) lacks reliable prognostic biomarkers. Homologous recombination deficiency (HRD) has been implicated in genomic instability across multiple cancers, but its prognostic significance in ESCC remains unexplored. This study aimed to evaluate HRD score as a prognostic biomarker and develop a machine learning-based predictive model for ESCC. METHODS: Transcriptomic and clinical data from 78 ESCC patients were obtained from The Cancer Genome Atlas (TCGA) and randomly split into training (70%) and test (30%) cohorts. Prognostic models were constructed using 112 machine learning algorithm combinations based on DNA damage response (DDR)-related genes. Gene set enrichment analysis (GSEA), somatic mutation profiling, and immune cell infiltration estimation via CIBERSORT were performed to characterize HRD-associated molecular features. RESULTS: High HRD scores were significantly associated with poorer overall survival (P<0.05). Among 112 algorithm combinations, the survival support vector machine (Survival-SVM) model demonstrated optimal performance [training concordance index (C-index): 0.741; test C-index: 0.708], identifying six hub genes: PARP1, MBD4, TELO2, NSMCE3, SMUG1, and BABAM1. A nomogram incorporating risk score (RS) and clinical variables achieved strong predictive accuracy for 1- to 3-year survival [area under the curve (AUC) >0.7]. High-HRD tumors exhibited distinct mutational patterns (TP53 and TTN) and enriched glutathione metabolism and cytochrome P450 pathways. Immune infiltration analysis revealed significant differences in plasma cell and neutrophil infiltration between risk groups (P<0.05), suggesting HRD-associated immune microenvironment remodeling. CONCLUSIONS: We developed a novel HRD-based prognostic model incorporating six DDR-related genes that demonstrates robust predictive performance in ESCC. HRD score is identified as an independent prognostic factor associated with genomic instability, immune microenvironment alterations, and clinical outcomes. These findings provide a theoretical basis for personalized treatment strategies, including potential applications of PARP inhibitors and immunotherapy in ESCC.

Esophageal squamous cell carcinoma (ESCC)↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗

Predicting dire outcomes of patients with community acquired pneumonia.

Community-acquired pneumonia (CAP) is an important clinical condition with regard to patient mortality, patient morbidity, and healthcare resource utilization. The assessment of the likely clinical course of a CAP patient can significantly influence decision making about whether to treat the patient as an inpatient or as an outpatient. That decision can in turn influence resource utilization, as well as patient well being. Predicting dire outcomes, such as mortality or severe clinical complications, is a particularly important component in assessing the clinical course of patients. We used a training set of 1601 CAP patient cases to construct 11 statistical and machine-learning models that predict dire outcomes. We evaluated the resulting models on 686 additional CAP-patient cases. The primary goal was not to compare these learning algorithms as a study end point; rather, it was to develop the best model possible to predict dire outcomes. A special version of an artificial neural network (NN) model predicted dire outcomes the best. Using the 686 test cases, we estimated the expected healthcare quality and cost impact of applying the NN model in practice. The particular, quantitative results of this analysis are based on a number of assumptions that we make explicit; they will require further study and validation. Nonetheless, the general implication of the analysis seems robust, namely, that even small improvements in predictive performance for prevalent and costly diseases, such as CAP, are likely to result in significant improvements in the quality and efficiency of healthcare delivery. Therefore, seeking models with the highest possible level of predictive performance is important. Consequently, seeking ever better machine-learning and statistical modeling methods is of great practical significance.

Community-Acquired Infections↗