PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease↗

Protein beta-turn prediction using nearest-neighbor method.

MOTIVATION: With the emerging success of protein secondary structure prediction through the applications of various statistical and machine learning techniques, similar techniques have been applied to protein beta-turn prediction. In this study, we perform protein beta-turn prediction using a k-nearest neighbor method, which is combined with a filter that uses predicted protein secondary structure information. Traditional beta-turn prediction from k-nearest neighbor method is modified to account for the unbalanced ratio of the natural occurrence of beta-turns and non-beta-turns. RESULTS: Our prediction scheme is tested on a set of 426 non-homologous protein sequences. The prediction scheme consists of two stages: k-nearest neighbor method stage and filtering stage. Variations of the k-nearest neighbor method were used to take property of beta-turns into consideration. Our filtering method uses beta-turn/non-beta-turn estimates from the k-nearest neighbor method stage and predicted protein secondary structure information from PSI-PRED in order to get new beta-turn/non-beta-turn estimate. Our result is compared with the previously best known beta-turn prediction method on the dataset of 426 non-homologous protein sequences and is shown to give slightly superior performance at significantly lower computational complexity. AVAILABILITY: Contact the author for information on the source code of the programs used.

Algorithms↗

Predictive models for protein crystallization.

Crystallization of proteins is a nontrivial task, and despite the substantial efforts in robotic automation, crystallization screening is still largely based on trial-and-error sampling of a limited subset of suitable reagents and experimental parameters. Funding of high throughput crystallography pilot projects through the NIH Protein Structure Initiative provides the opportunity to collect crystallization data in a comprehensive and statistically valid form. Data mining and machine learning algorithms thus have the potential to deliver predictive models for protein crystallization. However, the underlying complex physical reality of crystallization, combined with a generally ill-defined and sparsely populated sampling space, and inconsistent scoring and annotation make the development of predictive models non-trivial. We discuss the conceptual problems, and review strengths and limitations of current approaches towards crystallization prediction, emphasizing the importance of comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high throughput crystallography initiatives, exchange of information and standardization should be encouraged, aiming to effectively integrate data mining and machine learning efforts into a comprehensive predictive framework for protein crystallization. Similar experimental design and knowledge discovery strategies should be applied to valid analysis and prediction of protein expression, solubilization, and purification, as well as crystal handling and cryo-protection.

Bayes Theorem↗

PSoL: a positive sample only learning algorithm for finding non-coding RNA genes.

MOTIVATION: Small non-coding RNA (ncRNA) genes play important regulatory roles in a variety of cellular processes. However, detection of ncRNA genes is a great challenge to both experimental and computational approaches. In this study, we describe a new approach called positive sample only learning (PSoL) to predict ncRNA genes in the Escherichia coli genome. Although PSoL is a machine learning method for classification, it requires no negative training data, which, in general, is hard to define properly and affects the performance of machine learning dramatically. In addition, using the support vector machine (SVM) as the core learning algorithm, PSoL can integrate many different kinds of information to improve the accuracy of prediction. Besides the application of PSoL for predicting ncRNAs, PSoL is applicable to many other bioinformatics problems as well. RESULTS: The PSoL method is assessed by 5-fold cross-validation experiments which show that PSoL can achieve about 80% accuracy in recovery of known ncRNAs. We compared PSoL predictions with five previously published results. The PSoL method has the highest percentage of predictions overlapping with those from other methods.

Algorithms↗

HINN: Hierarchical Input Neural Network identifies multi-omics biomarker for cognitive decline.

Understanding complex diseases requires models that can integrate diverse layers of biological data while yielding insights that are biologically interpretable. Although multi-omics integration with machine learning (ML) has advanced disease prediction and biomarker discovery, most existing approaches overlook the hierarchical and regulatory relationships that connect these molecular layers. Here, we present the Hierarchical Input Neural Network (HINN), a deep learning framework that incorporates known cross-omics relationships directly into its architecture, capturing the flow of information from genomics to epigenomics, transcriptomics, and downstream biological processes. By embedding these relationships, HINN improves both predictive performance and biological interpretability. We applied HINN to blood-derived multi-omics data from individuals with Alzheimer's disease or mild cognitive impairment to predict cognitive scores from standardized assessments. HINN outperformed both baseline and state-of-the-art models and pinpointed multi-omics biomarkers-including SNPs and promoter-region CpG sites in ATP6V1C1 and RCHY1 -that were significantly correlated with plasma p-Tau181 levels. These features map to biologically relevant processes with potential implications for cognitive decline. Our findings demonstrate how combining deep learning with biological knowledge can uncover interpretable, blood-based biomarkers for cognitive decline due to complex diseases such as Alzheimer's. All code and data are openly available at https://github.com/bozdaglab/HINN.

Alzheimer’s disease↗

NMR metabolomics and glycomics for cancer detection in patients with non-specific symptoms: a prospective observational cohort study.

BACKGROUND: Early cancer diagnosis in patients with non-specific symptoms is limited by the lack of discriminatory tests. Within the Oxfordshire Suspected CANcer (SCAN) pathway, exploratory biomarker work showed that serum 1H NMR-based metabolomics can identify cancer with high accuracy. SCAN2 evaluated whether integrating metabolomics with glycomics provides complementary molecular information and improves discrimination in a clinically complex, real-world population. METHODS: Serum from 369 SCAN patients (59 cancers) was analysed using AXINON® System-derived NMR metabolomics and HPLC-MS glycomics. Machine-learning models were trained to predict cancer status, with performance assessed by receiver operating characteristic (ROC) analysis of pooled cross-validated predictions. To place cancer risk in a broader clinical context, a second classifier modelling alternative non-cancer diagnosis was incorporated, and mean predicted probabilities from both models were jointly projected into a two-dimensional space, maintaining strict separation of training and test data. FINDINGS: In the full cohort, integration of glycomics with metabolomics achieved an AUC of 0.814 (95% CI 0.808-0.820). In a refined sub-cohort excluding major comorbidities and selected cancer types (32 cancers, 277 non-cancers), performance improved to an AUC of 0.884 (95% CI 0.879-0.890). Discriminatory features included cancer-associated biantennary fucosylated glycans alongside amino acid metabolites (glutamate, histidine) and lipoprotein-related measures. A classifier distinguishing metastatic from non-metastatic disease (n = 29 vs. 30) achieved an AUC of 0.80. Joint probability analysis in the full cohort preserved cancer-associated signatures across comorbidity burden, with projection-based classification achieving an accuracy of 89.2% (95% CI 85.7-92.6). INTERPRETATION: These findings validate the SCAN1 metabolomic signature in a more clinically complex cohort and indicate that integrating glycomics with metabolomics provides complementary biological information for cancer discrimination. Joint probability analysis provides an interpretable framework for cancer risk stratification within multimorbid diagnostic pathways, supporting the clinical potential of scalable multi-omics blood testing. FUNDING: EPSRC, EU Horizon 2020, Wellcome/MLSTF, Novo Nordisk Foundation.

Humans↗

Drug design by machine learning: support vector machines for pharmaceutical data analysis.

We show that the support vector machine (SVM) classification algorithm, a recent development from the machine learning community, proves its potential for structure-activity relationship analysis. In a benchmark test, the SVM is compared to several machine learning techniques currently used in the field. The classification task involves predicting the inhibition of dihydrofolate reductase by pyrimidines, using data obtained from the UCI machine learning repository. Three artificial neural networks, a radial basis function network, and a C5.0 decision tree are all outperformed by the SVM. The SVM is significantly better than all of these, bar a manually capacity-controlled neural network, which takes considerably longer to train.

Algorithms↗

Integrative multi-omics analyses suggest a candidate microbial metabolite-associated host gene network in ulcerative colitis.

Ulcerative colitis (UC) is associated with gut microbial dysbiosis, but the host molecular alterations potentially linked to microbially derived metabolites remain incompletely understood. We integrated Mendelian randomization (MR), microbial metabolite annotation, computational target prediction, colonic transcriptomics, network analysis, and machine learning. MiBioGen microbiome GWAS data were used as exposures and FinnGen Release 12 ULCERENTER as the outcome. Metabolites linked to MR-prioritized taxa were retrieved from GutMGene, and human targets were predicted using SwissTargetPrediction and SEA. UC-related genes were defined by integrating differential expression analysis and WGCNA and then intersected with predicted metabolite targets. MR prioritized one family and eight genera showing nominal genetically supported associations with UC, but none remained significant after Benjamini-Hochberg FDR correction. Three prioritized genera were linked to 15 microbe-metabolite records, corresponding to 13 unique metabolites; nine were retained for target prediction, yielding 277 unique predicted human targets. Transcriptomic analysis identified 1,530 DEGs and a 312-gene MEgrey60 module, with 273 overlapping genes, producing 1,569 unique UC-related genes. Their intersection with the 277 predicted targets yielded 47 candidate genes. Enrichment analyses highlighted mainly metabolic and lipid-related processes. Random Forest showed the highest mean AUC across the two independent external benchmarking cohorts, and SHAP prioritized EPHX1, HSD17B2, IGFBP5, and MMP10. IBDome analysis showed inflammation-associated expression differences in these genes. This study provides a genomics-informed, hypothesis-generating framework that prioritizes candidate microbe-metabolite-host relationships in UC for future experimental validation.

Humans↗

Prediction of transporter family from protein sequence by support vector machine approach.

Transporters play key roles in cellular transport and metabolic processes, and in facilitating drug delivery and excretion. These proteins are classified into families based on the transporter classification (TC) system. Determination of the TC family of transporters facilitates the study of their cellular and pharmacological functions. Methods for predicting TC family without sequence alignments or clustering are particularly useful for studying novel transporters whose function cannot be determined by sequence similarity. This work explores the use of a machine learning method, support vector machines (SVMs), for predicting the family of transporters from their sequence without the use of sequence similarity. A total of 10,636 transporters in 13 TC subclasses, 1914 transporters in eight TC families, and 168,341 nontransporter proteins are used to train and test the SVM prediction system. Testing results by using a separate set of 4351 transporters and 83,151 nontransporter proteins show that the overall accuracy for predicting members of these TC subclasses and families is 83.4% and 88.0%, respectively, and that of nonmembers is 99.3% and 96.6%, respectively. The accuracies for predicting members and nonmembers of individual TC subclasses are in the range of 70.7-96.1% and 97.6-99.9%, respectively, and those of individual TC families are in the range of 60.6-97.1% and 91.5-99.4%, respectively. A further test by using 26,139 transmembrane proteins outside each of the 13 TC subclasses shows that 90.4-99.6% of these are correctly predicted. Our study suggests that the SVM is potentially useful for facilitating functional study of transporters irrespective of sequence similarity.

Amino Acid Sequence↗

Feature-map vectors: a new class of informative descriptors for computational drug discovery.

In order to develop robust machine-learning or statistical models for predicting biological activity, descriptors that capture the essence of the protein-ligand interaction are required. In the absence of structural information from X-ray or NMR experiments, deriving informative descriptors can be difficult. We have developed feature-map vectors (FMVs), a new class of descriptors based on chemical features, to address this challenge. FMVs, which are derived from the conformational models of a few actives, are low dimensional, problem specific, and highly interpretable. By using shape-based alignments and scoring with chemical features, FMVs can combine information about a molecule's shape and the pharmacophores it can match. In five validation studies, bag classifiers built using FMVs have shown high enrichments for identifying actives for five diverse targets: CDK2, 5-HT(3), DHFR, thrombin, and ACE. The interpretability of these descriptors has been demonstrated for CDK2 and 5-HT(3), where the method automatically discovers the standard literature pharmacophore.

Algorithms↗

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95 % CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma↗

Architectural logic of the 3D genome: mechanisms of dysregulation and emerging cancer therapeutics.

The three-dimensional (3D) genome provides an essential layer of organization that shapes genome function in space and time. Chromatin compartments and topologically associating domains (TADs) arise from the interplay between intrinsic properties of chromatin and architectural factors, including cohesin and CTCF. Despite substantial progress in defining these structural features, whether 3D genome architecture plays a causal role in regulating processes such as transcription, DNA replication, and DNA repair, or instead reflects underlying regulatory activity, remains unresolved. Here, we use the distinction between chromatin-intrinsic features and architectural factors as a framework to evaluate evidence for causality in genome structure-function relationships. We extend this framework to cancer, where both intrinsic alterations (including noncoding mutations, structural variants, and changes in chromatin state) and architectural factor perturbations (such as mutations in architectural proteins and dysregulation of transcriptional machinery) disrupt genome organization and contribute to disease progression. These findings suggest that alterations in genome structure can, in some contexts, actively reshape oncogenic programs. A major limitation in applying 3D genome insights to cancer biology is the cost and complexity of omics assays. Recent advances in artificial intelligence (AI) and machine learning (ML) enable inference and prediction of 3D genome organization from sequence and epigenomic features, providing insight into the extent to which genome folding is encoded intrinsically versus dynamically regulated in architectural factors. This perspective provides a unified view of how genome structure is established, how it relates to function, and how its disruption contributes to tumorigenesis.

3D genome↗

Descriptor-based protein remote homology identification.

Here, we report a novel protein sequence descriptor-based remote homology identification method, able to infer fold relationships without the explicit knowledge of structure. In a first phase, we have individually benchmarked 13 different descriptor types in fold identification experiments in a highly diverse set of protein sequences. The relevant descriptors were related to the fold class membership by using simple similarity measures in the descriptor spaces, such as the cosine angle. Our results revealed that the three best-performing sets of descriptors were the sequence-alignment-based descriptor using PSI-BLAST e-values, the descriptors based on the alignment of secondary structural elements (SSEA), and the descriptors based on the occurrence of PROSITE functional motifs. In a second phase, the three top-performing descriptors were combined to obtain a final method with improved performance, which we named DescFold. Class membership was predicted by Support Vector Machine (SVM) learning. In comparison with the individual PSI-BLAST-based descriptor, the rate of remote homology identification increased from 33.7% to 46.3%. We found out that the composite set of descriptors was able to identify the true remote homolog for nearly every sixth sequence at the 95% confidence level, or some 10% more than a single PSI-BLAST search. We have benchmarked the DescFold method against several other state-of-the-art fold recognition algorithms for the 172 LiveBench-8 targets, and we concluded that it was able to add value to the existing techniques by providing a confident hit for at least 10% of the sequences not identifiable by the previously known methods.

Algorithms↗

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship↗

nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.

Nonsynonymous single nucleotide polymorphisms (nsSNPs) are prevalent in genomes and are closely associated with inherited diseases. To facilitate identifying disease-associated nsSNPs from a large number of neutral nsSNPs, it is important to develop computational tools to predict the nsSNP's phenotypic effect (disease-associated versus neutral). nsSNPAnalyzer, a web-based software developed for this purpose, extracts structural and evolutionary information from a query nsSNP and uses a machine learning method called Random Forest to predict the nsSNP's phenotypic effect. nsSNPAnalyzer server is available at http://snpanalyzer.utmem.edu/.

Algorithms↗

Prediction of protein structural classes by support vector machines.

In this paper, we apply a new machine learning method which is called support vector machine to approach the prediction of protein structural class. The support vector machine method is performed based on the database derived from SCOP which is based upon domains of known structure and the evolutionary relationships and the principles that govern their 3D structure. As a result, high rates of both self-consistency and jackknife test are obtained. This indicates that the structural class of a protein inconsiderably correlated with its amino and composition, and the support vector machine can be referred as a powerful computational tool for predicting the structural classes of proteins.

Artificial Intelligence↗

Integrated analysis of established and novel microbial and chemical methods for microbial source tracking.

Several microbes and chemicals have been considered as potential tracers to identify fecal sources in the environment. However, to date, no one approach has been shown to accurately identify the origins of fecal pollution in aquatic environments. In this multilaboratory study, different microbial and chemical indicators were analyzed in order to distinguish human fecal sources from nonhuman fecal sources using wastewaters and slurries from diverse geographical areas within Europe. Twenty-six parameters, which were later combined to form derived variables for statistical analyses, were obtained by performing methods that were achievable in all the participant laboratories: enumeration of fecal coliform bacteria, enterococci, clostridia, somatic coliphages, F-specific RNA phages, bacteriophages infecting Bacteroides fragilis RYC2056 and Bacteroides thetaiotaomicron GA17, and total and sorbitol-fermenting bifidobacteria; genotyping of F-specific RNA phages; biochemical phenotyping of fecal coliform bacteria and enterococci using miniaturized tests; specific detection of Bifidobacterium adolescentis and Bifidobacterium dentium; and measurement of four fecal sterols. A number of potentially useful source indicators were detected (bacteriophages infecting B. thetaiotaomicron, certain genotypes of F-specific bacteriophages, sorbitol-fermenting bifidobacteria, 24-ethylcoprostanol, and epycoprostanol), although no one source identifier alone provided 100% correct classification of the fecal source. Subsequently, 38 variables (both single and derived) were defined from the measured microbial and chemical parameters in order to find the best subset of variables to develop predictive models using the lowest possible number of measured parameters. To this end, several statistical or machine learning methods were evaluated and provided two successful predictive models based on just two variables, giving 100% correct classification: the ratio of the densities of somatic coliphages and phages infecting Bacteroides thetaiotaomicron to the density of somatic coliphages and the ratio of the densities of fecal coliform bacteria and phages infecting Bacteroides thetaiotaomicron to the density of fecal coliform bacteria. Other models with high rates of correct classification were developed, but in these cases, higher numbers of variables were required.

Animals↗