PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Support vector machines for prediction and analysis of beta and gamma-turns in proteins.

Tight turns have long been recognized as one of the three important features of proteins, together with alpha-helix and beta-sheet. Tight turns play an important role in globular proteins from both the structural and functional points of view. More than 90% tight turns are beta-turns and most of the rest are gamma-turns. Analysis and prediction of beta-turns and gamma-turns is very useful for design of new molecules such as drugs, pesticides, and antigens. In this paper we investigated two aspects of applying support vector machine (SVM), a promising machine learning method for bioinformatics, to prediction and analysis of beta-turns and gamma-turns. First, we developed two SVM-based methods, called BTSVM and GTSVM, which predict beta-turns and gamma-turns in a protein from its sequence. When compared with other methods, BTSVM has a superior performance and GTSVM is competitive. Second, we used SVMs with a linear kernel to estimate the support of amino acids for the formation of beta-turns and gamma-turns depending on their position in a protein. Our analysis results are more comprehensive and easier to use than the previous results in designing turns in proteins.

Algorithms↗

Prediction of oxidoreductase-catalyzed reactions based on atomic properties of metabolites.

MOTIVATION: Our knowledge of metabolism is far from complete, and the gaps in our knowledge are being revealed by metabolomic detection of small-molecules not previously known to exist in cells. An important challenge is to determine the reactions in which these compounds participate, which can lead to the identification of gene products responsible for novel metabolic pathways. To address this challenge, we investigate how machine learning can be used to predict potential substrates and products of oxidoreductase-catalyzed reactions. RESULTS: We examined 1956 oxidation/reduction reactions in the KEGG database. The vast majority of these reactions (1626) can be divided into 12 subclasses, each of which is marked by a particular type of functional group transformation. For a given transformation, the local structures of reaction centers in substrates and products can be characterized by patterns. These patterns are not unique to reactants but are widely distributed among KEGG metabolites. To distinguish reactants from non-reactants, we trained classifiers (linear-kernel Support Vector Machines) using negative and positive examples. The input to a classifier is a set of atomic features that can be determined from the 2D chemical structure of a compound. Depending on the subclass of reaction, the accuracy of prediction for positives (negatives) is 64 to 93% (44 to 92%) when asking if a compound is a substrate and 71 to 98% (50 to 92%) when asking if a compound is a product. Sensitivity analysis reveals that this performance is robust to variations of the training data. Our results suggest that metabolic connectivity can be predicted with reasonable accuracy from the presence or absence of local structural motifs in compounds and their readily calculated atomic features. AVAILABILITY: Classifiers reported here can be used freely for noncommercial purposes via a Java program available upon request.

Algorithms↗

Constructive induction and protein tertiary structure prediction.

To date, the only methods that have been used successfully to predict protein structures have been based on identifying homologous proteins whose structures are known. However, such methods are limited by the fact that some proteins have similar structure but no significant sequence homology. We consider two ways of applying machine learning to facilitate protein structure prediction. We argue that a straightforward approach will not be able to improve the accuracy of classification achieved by clustering by alignment scores alone. In contrast, we present a novel constructive induction approach that learns better representations of amino acid sequences in terms of physical and chemical properties. Our learning method combines knowledge and search to shift the representation of sequences so that semantic similarity is more easily recognized by syntactic matching. Our approach promises not only to find new structural relationships among protein sequences, but also expands our understanding of the roles knowledge can play in learning via experience in this challenging domain.

Artificial Intelligence↗

Prediction of torsade-causing potential of drugs by support vector machine approach.

In an effort to facilitate drug discovery, computational methods for facilitating the prediction of various adverse drug reactions (ADRs) have been developed. So far, attention has not been sufficiently paid to the development of methods for the prediction of serious ADRs that occur less frequently. Some of these ADRs, such as torsade de pointes (TdP), are important issues in the approval of drugs for certain diseases. Thus there is a need to develop tools for facilitating the prediction of these ADRs. This work explores the use of a statistical learning method, support vector machine (SVM), for TdP prediction. TdP involves multiple mechanisms and SVM is a method suitable for such a problem. Our SVM classification system used a set of linear solvation energy relationship (LSER) descriptors and was optimized by leave-one-out cross validation procedure. Its prediction accuracy was evaluated by using an independent set of agents and by comparison with results obtained from other commonly used classification methods using the same dataset and optimization procedure. The accuracies for the SVM prediction of TdP-causing agents and non-TdP-causing agents are 97.4 and 84.6% respectively; one is substantially improved against and the other is comparable to the results obtained by other classification methods useful for multiple-mechanism prediction problems. This indicates the potential of SVM in facilitating the prediction of TdP-causing risk of small molecules and perhaps other ADRs that involve multiple mechanisms.

Algorithms↗

PDP-Miner: an AI/ML tool to detect prophage tail proteins with depolymerase domains across thousands of bacterial genomes.

MOTIVATION: Antibiotic resistance is predicted to become the leading cause of human mortality by 2050. Despite this, no other major antibiotic class has been approved for medical use since 1987. Nevertheless, phage tail proteins offer a promising alternative, given their depolymerase activity toward outer membrane polysaccharides. Several pathogenic bacteria harbor prophages, thus making these prophages' molecular target already known. RESULTS: We therefore developed a wrapper for an existing machine learning-based phage depolymerase prediction tool (Depolymerase-Predictor), called PDP-Miner, which annotates phage tail proteins ab initio, detects depolymerase activity within this candidate protein subset, and then performs post-hoc validation by annotating protein domains thereby allowing the user to investigate for protein domains indicative of depolymerase activity. This tool allowed identification of 10 high confidence phage depolymerase gene candidates across all 1294 Pseudomonas genomes available on the International Pseudomonas Consortium Database while also accurately reporting depolymerases in known phage genomes, similarly to other software like PhageDPO or DepoScope. AVAILABILITY AND IMPLEMENTATION: Source code, test datasets and documentation are freely available for download at http:///www.github.com/jeffgauthier/pdpminer. This software is free and open source under the GNU General Public License v3.0.

Prophages↗

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent↗

The prediction of human oral absorption for diffusion rate-limited drugs based on heuristic method and support vector machine.

Support vector machine (SVM), as a novel machine learning technique, was used for the prediction of the human oral absorption for a large and diverse data set using the five descriptors calculated from the molecular structure alone. The molecular descriptors were selected by heuristic method (HM) implemented in CODESSA. At the same time, in order to show the influence of different molecular descriptors on absorption and to well understand the absorption mechanism, HM was used to build several multivariable linear models using different numbers of molecular descriptors. Both the linear and non-linear model can give satisfactory prediction results: the square of correlation coefficient R(2) was 0.78 and 0.86 for the training set, and 0.70 and 0.73 for the test set respectively. In addition, this paper provides a new and effective method for predicting the absorption of the drugs from their structures and gives some insight into structural features related to the absorption of the drugs.

Administration, Oral↗

Two-sample comparison based on prediction error, with applications to candidate gene association studies.

To take advantage of the increasingly available high-density SNP maps across the genome, various tests that compare multilocus genotypes or estimated haplotypes between cases and controls have been developed for candidate gene association studies. Here we view this two-sample testing problem from the perspective of supervised machine learning and propose a new association test. The approach adopts the flexible and easy-to-understand classification tree model as the learning machine, and uses the estimated prediction error of the resulting prediction rule as the test statistic. This procedure not only provides an association test but also generates a prediction rule that can be useful in understanding the mechanisms underlying complex disease. Under the set-up of a haplotype-based transmission/disequilibrium test (TDT) type of analysis, we find through simulation studies that the proposed procedure has the correct type I error rates and is robust to population stratification. The power of the proposed procedure is sensitive to the chosen prediction error estimator. Among commonly used prediction error estimators, the .632+ estimator results in a test that has the best overall performance. We also find that the test using the .632+ estimator is more powerful than the standard single-point TDT analysis, the Pearson's goodness-of-fit test based on estimated haplotype frequencies, and two haplotype-based global tests implemented in the genetic analysis package FBAT. To illustrate the application of the proposed method in population-based association studies, we use the procedure to study the association between non-Hodgkin lymphoma and the IL10 gene.

Adult↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

Improving the reliability of medical software by predicting the dangerous software modules.

Software reliability analysis is inevitable for modern medical systems, since a large amount of medical system functionality is now dependent on software, and software does contribute to system failures. Most software reliability models are based on software failure data collected from the project. This creates a problem for the designers since, during the early stage, software failure data are not available. However, a valuable knowledge can be learned from the analysis of previous projects and applied to the new ones. This paper presents the approach that predicts the potentially dangerous software modules under development based on the analysis of the already finished modules using the machine-learning techniques. On the basis of the prediction given by our method software designers are able to devote more testing effort to the dangerous parts of the system, which results in a more reliable medical software system.

Algorithms↗

Machine learning prognostic model and drug survival analysis for lung adenocarcinoma in the context of radiotherapy.

BACKGROUND: Patients with lung adenocarcinoma (LUAD) receiving radiotherapy represent an important but underexplored clinical subgroup. These patients often undergo concomitant pharmacologic treatments, yet the prognostic impact and underlying determinants of such combined regimens remain poorly understood. OBJECTIVE: This retrospective observational study aimed to develop and validate a radiotherapy-specific machine learning prognostic model for LUAD and to compare survival across concomitant pharmacologic regimens. METHODS: In this retrospective observational study, using genomic and clinical data from TCGA, a radiotherapy-specific prognostic model for LUAD was developed and validated through ten machine learning algorithms. Survival analyses were conducted across distinct concomitant pharmacologic strategies, followed by functional enrichment to elucidate molecular mechanisms underlying differential outcomes. RESULTS: Demonstrating robust prognostic abilities, the model efficiently sorted patients into high- and low-risk categories. Both treatment type and risk score independently predicted overall survival, with significant interaction effects. Low-risk patients receiving targeted or combination therapy-mainly erlotinib, gefitinib, or bevacizumab-exhibited substantially improved survival compared with those receiving conventional chemotherapy. Enrichment of "Exogenous peptide presentation," "MHC class II assembly," "Peptide-MHC II assembly," and "Symbiotic interaction" pathways indicated immune modulation and host-tumor crosstalk as key mediators of treatment efficacy. CONCLUSION: This study establishes a radiotherapy-specific prognostic model for lung adenocarcinoma, demonstrating distinct molecular and therapeutic heterogeneity and highlighting the superior survival benefit of targeted combination therapy in low-risk patients.

Humans↗

Predicting protein-ligand binding affinities using novel geometrical descriptors and machine-learning methods.

Inspired by the concept of knowledge-based scoring functions, a new quantitative structure-activity relationship (QSAR) approach is introduced for scoring protein-ligand interactions. This approach considers that the strength of ligand binding is correlated with the nature of specific ligand/binding site atom pairs in a distance-dependent manner. In this technique, atom pair occurrence and distance-dependent atom pair features are used to generate an interaction score. Scoring and pattern recognition results obtained using Kernel PLS (partial least squares) modeling and a genetic algorithm-based feature selection method are discussed.

Algorithms↗

Proteomic signature of dementia risk in type 2 diabetes.

INTRODUCTION: Type 2 diabetes (T2D) significantly increases dementia risk, yet the molecular mechanisms underlying this association remain unclear. OBJECTIVES: This study aimed to identify protein signatures that distinguish dementia risk in T2D patients, develop a proteomic prediction model, and elucidate biological pathways connecting T2D and dementia. METHODS: We analyzed 2,920 plasma proteins from 52,958 participants (including 3,292 with T2D) in the UK Biobank Pharma Proteomics Project with a median follow-up of 14.6&#xa0;years. Cox regression models with interaction terms identified T2D-specific protein associations with dementia risk. Machine learning models were developed to predict dementia in T2D patients. Pathway analysis and weighted gene co-expression network analysis identified biological mechanisms linking T2D and dementia. RESULTS: We identified 471 proteins with significant interaction effects between T2D and dementia risk. In non-T2D individuals, elevated levels of neuronal pentraxin receptor (NPTXR, HR&#xa0;=&#xa0;0.74, 95&#xa0;%CI:0.66-0.83) and carbonic anhydrase 14 (CA14, HR&#xa0;=&#xa0;0.67, 95&#xa0;%CI:0.60-0.75) were exclusively associated with decreased dementia risk. Conversely, in T2D patients, elevated rho guanine nucleotide exchange factor 12 (ARHGEF12, HR&#xa0;=&#xa0;1.45, 95&#xa0;%CI:1.10-1.91) was specifically associated with increased dementia risk. A 51-protein model accurately predicted 15-year dementia risk in T2D patients (AUC&#xa0;=&#xa0;0.835, C-index&#xa0;=&#xa0;0.829), outperforming conventional clinical risk scores and maintaining high accuracy for Alzheimer's disease and vascular dementia. Pathway analysis revealed enrichment of IL6-JAK-STAT3 signaling in T2D-related dementia, while dysregulation of fatty acid metabolism was specific to T2D-associated Alzheimer's disease. CONCLUSIONS: This large-scale proteomic analysis identifies specific molecular signatures that differentiate dementia risk in diabetic and non-diabetic populations, with potential applications for early risk stratification and targeted interventions. The identified pathways provide novel insights into the pathophysiological processes connecting T2D and dementia and suggest potential therapeutic targets.

Humans↗

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4&#x207a; memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC&#x2009;>&#x2009;0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics↗

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

Support vector machine for predicting alpha-turn types.

Tight turns play an important role in globular proteins from both the structural and functional points of view. Of tight turns, beta-turns and gamma-turns have been extensively studied, but alpha-turns were little investigated. Recently, a systematic search for alpha-turns classified alpha-turns into nine different types according to their backbone trajectory features. In this paper, Support Vector Machines (SVMs), a new machine learning method, is proposed for predicting the alpha-turn types in proteins. The high rates of correct prediction imply that that the formation of different alpha-turn types is evidently correlated with the sequence of a pentapeptide, and hence can be approximately predicted based on the sequence information of the pentapeptide alone, although the incorporation of its interaction with the other part of a protein, the so-called "long distance interaction", will further improve the prediction quality.

Algorithms↗

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗