PubMed HealthSearch

SEARCH · PubMed Health

Results for “Predictive Learning Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

From conditioning to category learning: an adaptive network model.

We used adaptive network theory to extend the Rescorla-Wagner (1972) least mean squares (LMS) model of associative learning to phenomena of human learning and judgment. In three experiments subjects learned to categorize hypothetical patients with particular symptom patterns as having certain diseases. When one disease is far more likely than another, the model predicts that subjects will substantially overestimate the diagnosticity of the more valid symptom for the rare disease. The results of Experiments 1 and 2 provide clear support for this prediction in contradistinction to predictions from probability matching, exemplar retrieval, or simple prototype learning models. Experiment 3 contrasted the adaptive network model with one predicting pattern-probability matching when patients always had four symptoms (chosen from four opponent pairs) rather than the presence or absence of each of four symptoms, as in Experiment 1. The results again support the Rescorla-Wagner LMS learning rule as embedded within an adaptive network model.

Association Learning

Integration of Gene Expression and Digital Histology to Predict Treatment-Specific Responses in Breast Cancer.

Deep learning models applied to digital histology can predict gene expression signatures (GES) and offer a low-cost, rapidly available alternative to molecular testing at the time of diagnosis. We optimized transformer-based models to infer GES results and applied this approach to pre-treatment H&E-stained biopsies from 1,940 breast cancer patients treated with neoadjuvant chemotherapy in clinical trial and real-world cohorts. The most predictive histology-derived GES for pathologic complete response (pCR) in the I-SPY2 trial was validated in four external cohorts: CALGB 40601, CALGB 40603, a trial of durvalumab plus CT, and standard-of-care CT-treated patients from the University of Chicago. Among HER2-negative patients, a transformer-based model trained using a signature composed of estrogen-regulated genes, proliferation, apoptosis, and interferon response genes predicted pCR with an AUC of 0.794, outperforming models based on clinical features alone (AUC 0.704, p = 0.001), pathologist TIL assessment, and a model trained directly to predict response from I-SPY2 cases. Tertiles of this signature stratify patients into clinically relevant groups with increasing likelihood of complete response, with pCR rates ≥50% in the top tertile regardless of treatment or hormone receptor status. Additional transformer-based signature models predicted response to specific therapies (but not chemotherapy alone), including a HER2 signaling signature in IO-treated patients, and a claudin-low signature in bevacizumab treated patients. In HER2- cohorts with available gene expression data and histology, models trained on expression data performed similarly to digital histology predictions, but the combination of gene expression and histology outperformed histology alone. These findings suggest that histology-based GES provides additive information to RNA sequencing data and can inform precision treatment selection across breast cancer subtypes.

Journal Article

Importance of attributions as a predictor of how people cope with failure.

This study examined the extent to which causal attributions were predictive of depressed mood in college students who experienced a negative event. In a replication and extension of a study by Metalsky, Abramson, Seligman, Semmel, and Peterson (1982), we evaluated students' attributional style and their attributions for an examination performance in the college classroom. Additionally, an indirect probe was used to assess unsolicited attributions. Subjects were asked about their plans to prepare for the next examination in order to test for the motivational deficits predicted by the reformulated learned helplessness (RLH) model. Unlike Metalsky et al., attributional style did not predict depressed mood following a disappointing examination performance. Attributions for the particular examination performance were predictive of depressed mood for students who were disappointed in their examination performance. Few subjects, 31%, gave attributions in response to the indirect probe, and there was no support for the prediction that unexpected negative events would lead to subjects' making more attributions. Internal, stable, and global attributions for poor examination performance resulted in students making more plans to study for the next examination, a finding contrary to what is predicted by the RLH model.

Achievement

Quantitative molecular analysis predicts 5-hydroxytryptamine3 receptor binding affinity.

A quantitative molecular model was derived to predict drug affinities for 5-hydroxytryptamine3 (5-HT3) receptors. The model was based on the molecular characteristics of a "learning set" of 40 pharmacological agents that had been analyzed previously in radioligand binding studies. Molecules were analyzed for various structural features, i.e., the presence of a benzenoid ring and nitrogen atom, substitutions on the benzenoid ring, the location of the substitutions on the nitrogen, and the molecular characteristics of the most direct pathway from the benzenoid ring to the nitrogen. Weighting factors, based on published 5-HT3 receptor affinity data, were then assigned to each of 10 molecular characteristics. The derived computational model predicts accurately the affinities of the learning set for the 5-HT3 receptor (r = 0.98; p less than 0.001). The computational model was then used to predict the receptor affinities of a "test set" of 40 pharmacological agents. The predicted values for these agents also correlate significantly (r = 0.83; p less than 0.001) with drug affinities for the 5-HT3 receptor, as determined by radioligand binding assays. This first line screening approach allows for the accurate prediction of drug affinities based on molecular characteristics with minimal dependence upon animal tissues or radioactivity.

Animals

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature

Attention and learning processes in the identification and categorization of integral stimuli.

The relationship between subjects' identification and categorization learning of integral-dimension stimuli was studied within the framework of an exemplar-based generalization model. The model was used to predict subjects' learning in six different categorization conditions on the basis of data obtained in a single identification learning condition. A crucial assumption in the model is that because of selective attention to component dimensions, similarity relations may change in systematic ways across different experimental contexts. The theoretical analysis provided evidence that, at least under unspeeded conditions, selective attention may play a critical role in determining the identification-categorization relationship for integral stimuli. Evidence was also provided that similarity among exemplars decreased as a function of identification learning. Various alternative classification models, including prototype, multiple-prototype, average distance, and "value-on-dimensions" models, were unable to account for the results.

Attention

Combining exemplar-based category representations and connectionist learning rules.

Adaptive network and exemplar-similarity models were compared on their ability to predict category learning and transfer data. An exemplar-based network (Kruschke, 1990a, 1990b, 1992) that combines key aspects of both modeling approaches was also tested. The exemplar-based network incorporates an exemplar-based category representation in which exemplars become associated to categories through the same error-driven, interactive learning rules that are assumed in standard adaptive networks. Experiment 1, which partially replicated and extended the probabilistic classification learning paradigm of Gluck and Bower (1988a), demonstrated the importance of an error-driven learning rule. Experiment 2, which extended the classification learning paradigm of Medin and Schaffer (1978) that discriminated between exemplar and prototype models, demonstrated the importance of an exemplar-based category representation. Only the exemplar-based network accounted for all the major qualitative phenomena; it also achieved good quantitative predictions of the learning and transfer data in both experiments.

Concept Formation

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4⁺ memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC > 0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics

Proteomic signature of dementia risk in type 2 diabetes.

INTRODUCTION: Type 2 diabetes (T2D) significantly increases dementia risk, yet the molecular mechanisms underlying this association remain unclear. OBJECTIVES: This study aimed to identify protein signatures that distinguish dementia risk in T2D patients, develop a proteomic prediction model, and elucidate biological pathways connecting T2D and dementia. METHODS: We analyzed 2,920 plasma proteins from 52,958 participants (including 3,292 with T2D) in the UK Biobank Pharma Proteomics Project with a median follow-up of 14.6 years. Cox regression models with interaction terms identified T2D-specific protein associations with dementia risk. Machine learning models were developed to predict dementia in T2D patients. Pathway analysis and weighted gene co-expression network analysis identified biological mechanisms linking T2D and dementia. RESULTS: We identified 471 proteins with significant interaction effects between T2D and dementia risk. In non-T2D individuals, elevated levels of neuronal pentraxin receptor (NPTXR, HR = 0.74, 95 %CI:0.66-0.83) and carbonic anhydrase 14 (CA14, HR = 0.67, 95 %CI:0.60-0.75) were exclusively associated with decreased dementia risk. Conversely, in T2D patients, elevated rho guanine nucleotide exchange factor 12 (ARHGEF12, HR = 1.45, 95 %CI:1.10-1.91) was specifically associated with increased dementia risk. A 51-protein model accurately predicted 15-year dementia risk in T2D patients (AUC = 0.835, C-index = 0.829), outperforming conventional clinical risk scores and maintaining high accuracy for Alzheimer's disease and vascular dementia. Pathway analysis revealed enrichment of IL6-JAK-STAT3 signaling in T2D-related dementia, while dysregulation of fatty acid metabolism was specific to T2D-associated Alzheimer's disease. CONCLUSIONS: This large-scale proteomic analysis identifies specific molecular signatures that differentiate dementia risk in diabetic and non-diabetic populations, with potential applications for early risk stratification and targeted interventions. The identified pathways provide novel insights into the pathophysiological processes connecting T2D and dementia and suggest potential therapeutic targets.

Humans

Cat lung hemodynamics: comparison of experimental results and model predictions.

Commonly, attempts have been made to learn about the structure and function of the pulmonary vascular bed from measurements of arterial and venous pressures and blood flow rate under steady-state conditions (e.g., from pressure vs. flow data) or dynamic conditions (e.g., from vascular occlusion data). Zhuang et al. (J. Appl. Physiol. 55: 1341-1348, 1983) have presented a detailed model of steady-state cat lung hemodynamics based on direct measurements of anatomical and elasticity data. This model provides an opportunity to better understand the information content of the hemodynamic data. Therefore, in the present study we carried out a series of steady-state and dynamic experiments on isolated cat lungs. We then compared the results with those predicted by the model. We found that the model provided a good fit to the steady-state data. However, to fit the dynamic data, some modifications were necessary to account for the viscous behavior of the vessel walls and to move the first moment of the distribution of vascular resistance toward the arterial end of the vascular bed relative to that of the distribution of vascular compliance. Due to the sensitivity of the vascular resistance to small changes in vessel diameters and branching ratio, the modifications in morphometry represent small changes in morphometric data and are probably within the range of uncertainty in such data. The modifications had little effect on the steady-state model simulations but substantially improved the dynamic model simulations, suggesting that the dynamic data are quite sensitive to small changes in the relative distributions of vessel diameters and elasticity.

Animals

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

Learned helplessness, depression, and anxiety.

The learned helplessness model of depression predicts that depressives should tend to perceive reinforcement as response-independent in skill tasks. Depressed-anxious, nondepressed-anxious, and nondepressed-nonanxious college students estimated their chances for success in a skill or a chance task. (Virtually no depressed-nonanxious subjects could be obtained.) Depressed-anxious subjects showed less expectancy change in skill than nondepressed-anxious subjects, while these two groups exhibited similar expectancy change in chance. Nondepressed-anxious and nondepressed-nonanxious subjects did not differ in either skill or chance. The results for a discrimination learning problem were mixed. The groups did not differ in latency to shut off an aversive noise. So, depressed subjects perceptually distort the outcomes of skilled responding as being response-independent, and they may, under certain conditions, show deficits at learning the consequences of responses. These deficits may reflect learned helplessness and are specific to depression.

Anxiety

APNet, an explainable sparse deep learning model to discover differentially active drivers of severe COVID-19.

MOTIVATION: Computational analyses of bulk and single-cell omics provide translational insights into complex diseases, such as COVID-19, by revealing molecules, cellular phenotypes, and signalling patterns that contribute to unfavourable clinical outcomes. Current in silico approaches dovetail differential abundance, biostatistics, and machine learning, but often overlook nonlinear proteomic dynamics, like post-translational modifications, and provide limited biological interpretability beyond feature ranking. RESULTS: We introduce APNet, a novel computational pipeline that combines differential activity analysis based on SJARACNe co-expression networks with PASNet, a biologically informed sparse deep learning model, to perform explainable predictions for COVID-19 severity. The APNet driver-pathway network ingests SJARACNe co-regulation and classification weights to aid result interpretation and hypothesis generation. APNet outperforms alternative models in patient classification across three COVID-19 proteomic datasets, identifying predictive drivers and pathways, including some confirmed in single-cell omics and highlighting under-explored biomarker circuitries in COVID-19. AVAILABILITY AND IMPLEMENTATION: APNet's R, Python scripts, and Cytoscape methodologies are available at https://github.com/BiodataAnalysisGroup/APNet.

COVID-19

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease

NMR metabolomics and glycomics for cancer detection in patients with non-specific symptoms: a prospective observational cohort study.

BACKGROUND: Early cancer diagnosis in patients with non-specific symptoms is limited by the lack of discriminatory tests. Within the Oxfordshire Suspected CANcer (SCAN) pathway, exploratory biomarker work showed that serum 1H NMR-based metabolomics can identify cancer with high accuracy. SCAN2 evaluated whether integrating metabolomics with glycomics provides complementary molecular information and improves discrimination in a clinically complex, real-world population. METHODS: Serum from 369 SCAN patients (59 cancers) was analysed using AXINON® System-derived NMR metabolomics and HPLC-MS glycomics. Machine-learning models were trained to predict cancer status, with performance assessed by receiver operating characteristic (ROC) analysis of pooled cross-validated predictions. To place cancer risk in a broader clinical context, a second classifier modelling alternative non-cancer diagnosis was incorporated, and mean predicted probabilities from both models were jointly projected into a two-dimensional space, maintaining strict separation of training and test data. FINDINGS: In the full cohort, integration of glycomics with metabolomics achieved an AUC of 0.814 (95% CI 0.808-0.820). In a refined sub-cohort excluding major comorbidities and selected cancer types (32 cancers, 277 non-cancers), performance improved to an AUC of 0.884 (95% CI 0.879-0.890). Discriminatory features included cancer-associated biantennary fucosylated glycans alongside amino acid metabolites (glutamate, histidine) and lipoprotein-related measures. A classifier distinguishing metastatic from non-metastatic disease (n = 29 vs. 30) achieved an AUC of 0.80. Joint probability analysis in the full cohort preserved cancer-associated signatures across comorbidity burden, with projection-based classification achieving an accuracy of 89.2% (95% CI 85.7-92.6). INTERPRETATION: These findings validate the SCAN1 metabolomic signature in a more clinically complex cohort and indicate that integrating glycomics with metabolomics provides complementary biological information for cancer discrimination. Joint probability analysis provides an interpretable framework for cancer risk stratification within multimorbid diagnostic pathways, supporting the clinical potential of scalable multi-omics blood testing. FUNDING: EPSRC, EU Horizon 2020, Wellcome/MLSTF, Novo Nordisk Foundation.

Humans

Learned helplessness, depression, and the attribution of failure.

Depressed and nondepressed college students received experience with solvable, unsolvable, or no discrimination problems. When later tested on a series of patterned anagrams, depressed groups performed worse than nondepressed groups, and unsolvable groups performed worse than solvable and control groups. As predicted by the learned helplessness model of depression, nondepressed subjects given unsolvable problems showed anagram deficits parallel to those found in naturally occurring depression. When depressed subjects attributed their failure to the difficulty of the problems rather than to their own incompetence, performance improved strikingly. So, failure in itself is apparently not sufficient to produce helplessness deficits in man, but failure that leads to a decreased belief in personal competence is sufficient.

Achievement

Cross-Device Adaptation of Mirai for Mammography-Based Breast Cancer Risk Prediction.

Fine-tuning can adapt pretrained medical imaging models to new clinical datasets, but device-specific domain shifts may limit generalizability. We evaluated Mirai, a mammography-based deep learning model for breast cancer risk prediction, in a large screening cohort containing Hologic and General Electric (GE) full-field digital mammography systems, including GE Premium View (GE PV) and Tissue Equalization (GE TE) post-processing software. Native Mirai showed lower performance on TE images than on Hologic or PV images. Fine-tuning on TE images improved TE performance, particularly for short-term risk prediction, but substantially reduced performance on Hologic images, consistent with catastrophic forgetting. To mitigate this effect, we developed a device-invariant model using interleaved multi-device sampling and conditional adversarial training. This approach largely restored Hologic performance while maintaining improved TE performance, providing better robustness across heterogeneous imaging platforms. Comparison of cumulative and annual risk AUCs over a five-year time horizon further showed that performance gains were driven mainly by short- and intermediate-term predictions. These findings highlight both the value and dangers of device-specific fine-tuning and support balanced domain-adaptation strategies for deploying mammography-based risk models across diverse clinical imaging environments.

Journal Article

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning