PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Combining multi-species genomic data for microRNA identification using a Naive Bayes classifier.

MOTIVATION: Most computational methodologies for microRNA gene prediction utilize techniques based on sequence conservation and/or structural similarity. In this study we describe a new technique, which is applicable across several species, for predicting miRNA genes. This technique is based on machine learning, using the Naive Bayes classifier. It automatically generates a model from the training data, which consists of sequence and structure information of known miRNAs from a variety of species. RESULTS: Our study shows that the application of machine learning techniques, along with the integration of data from multiple species is a useful and general approach for miRNA gene prediction. Based on our experiments, we believe that this new technique is applicable to an extensive range of eukaryotes' genomes. Specific structure and sequence features are first used to identify miRNAs followed by a comparative analysis to decrease the number of false positives (FPs). The resulting algorithm exhibits higher specificity and similar sensitivity compared to currently used algorithms that rely on conserved genomic regions to decrease the rate of FPs.

Algorithms↗

Probabilistic finite-state machines--part I.

Probabilistic finite-state machines are used today in a variety of areas in pattern recognition, or in fields to which pattern recognition is linked: computational linguistics, machine learning, time series analysis, circuit testing, computational biology, speech recognition, and machine translation are some of them. In Part I of this paper, we survey these generative objects and study their definitions and properties. In Part II, we will study the relation of probabilistic finite-state automata with other well-known devices that generate strings as hidden Markov models and n-grams and provide theorems, algorithms, and properties that represent a current state of the art of these objects.

Algorithms↗

Classification of metabolites with kernel-partial least squares (K-PLS).

Numerous experimental and computational approaches have been developed to predict human drug metabolism. Since databases of human drug metabolism information are widely available, these can be used to train computational algorithms and generate predictive approaches. In turn, they may be used to assist in the identification of possible metabolites from a large number of molecules in drug discovery based on molecular structure alone. In the current study we have used a commercially available database (MetaDrug) and extracted a fraction of the human drug metabolism data. These data were used along with augmented atom descriptors in a predictive machine learning model, kernel-partial least squares (K-PLS). A total of 317 molecules, including parent drugs and their primary and secondary (sequential) metabolites, were used to build these models corresponding to individual metabolism rules, representing the formation of discrete metabolites, e.g., N-dealkylation. Each model was internally validated to assess the capability to classify other molecules that were left out. Using receiver operator curve statistics models for N-dealkylation, O-dealkylation, aromatic hydroxylation, aliphatic hydroxylation, O-glucuronidation, and O-sulfation gave area under the curve values from 0.75 to 0.84 and were able to predict between 61 and 79% active molecules upon leave-one-out testing. This preliminary study indicates that K-PLS and possibly other similar machine learning methods (such as support vector machines) can be applied to predicting human drug metabolite formation in a classification manner. Improvements can be achieved using considerably larger datasets that contain more positive examples for the less frequently occurring metabolite rules, as well as the external evaluation of novel molecules.

Biotransformation↗

On classification capability of neural networks: a case study with otoneurological data.

We investigated the capability of multilayer perceptron neural networks and Kohonen neural networks to recognize difficult otoneurological diseases from each other. We found that they are efficient methods, but the distribution of a learning set should be rather uniform. Also it is important that the number of learning cases is sufficient. If the two mentioned conditions are satisfied, these neural networks are similarly efficient as some other machine learning methods. The conditions are known in the theory of neural networks [1,2], but not often taken seriously in practice. Both networks functioned as well, excluding the case with several input variables, where the Kohonen neural networks surpassed the perceptron.

Algorithms↗

VR interaction techniques for medical imaging applications.

Methods of virtual reality (VR) offer new ways of human-computer interaction. Medicine is predestined to benefit from this new technology in many ways. Virtual environments can support physicians in their work, alleviate communication between specialists from different fields or be established in educational and training applications. For the field of visualization and analysis of three-dimensional anatomical images (e.g. CT or MRI scans), an application is introduced which expedites recognition of spatial coherencies and the exploration and manipulation of the 3D data. To avoid long periods of learning and accustoming and to facilitate work in such an environment, a powerful human-oriented interface is required allowing interactions similar to the real world and utilization of our natural experiences. This paper shows the use of eye tracking parameters for a level-of-detail algorithm and the integration of a glove-based hand gesture recognition into the virtual environment as an essential component of the human-machine interface. Furthermore, virtual bronchoscopy and virtual angioscopy are presented as examples for the use of the virtual environment.

Diagnosis, Computer-Assisted↗

Predicting the subcellular localization of human proteins using machine learning and exploratory data analysis.

Identifying the subcellular localization of proteins is particularly helpful in the functional annotation of gene products. In this study, we use Machine Learning and Exploratory Data Analysis (EDA) techniques to examine and characterize amino acid sequences of human proteins localized in nine cellular compartments. A dataset of 3,749 protein sequences representing human proteins was extracted from the SWISS-PROT database. Feature vectors were created to capture specific amino acid sequence characteristics. Relative to a Support Vector Machine, a Multi-layer Perceptron, and a Naive Bayes classifier, the C4.5 Decision Tree algorithm was the most consistent performer across all nine compartments in reliably predicting the subcellular localization of proteins based on their amino acid sequences (average Precision=0.88; average Sensitivity=0.86). Furthermore, EDA graphics characterized essential features of proteins in each compartment. As examples, proteins localized on the plasma membrane had higher proportions of hydrophobic amino acids; cytoplasmic proteins had higher proportions of neutral amino acids; and mitochondrial proteins had higher proportions of neutral amino acids and lower proportions of polar amino acids. These data showed that the C4.5 classifier and EDA tools can be effective for characterizing and predicting the subcellular localization of human proteins based on their amino acid sequences.

Algorithms↗

Prediction of caspase cleavage sites using Bayesian bio-basis function neural networks.

MOTIVATION: Apoptosis has drawn the attention of researchers because of its importance in treating some diseases through finding a proper way to block or slow down the apoptosis process. Having understood that caspase cleavage is the key to apoptosis, we find novel methods or algorithms are essential for studying the specificity of caspase cleavage activity and this helps the effective drug design. As bio-basis function neural networks have proven to outperform some conventional neural learning algorithms, there is a motivation, in this study, to investigate the application of bio-basis function neural networks for the prediction of caspase cleavage sites. RESULTS: Thirteen protein sequences with experimentally determined caspase cleavage sites were downloaded from NCBI. Bayesian bio-basis function neural networks are investigated and the comparisons with single-layer perceptrons, multilayer perceptrons, the original bio-basis function neural networks and support vector machines are given. The impact of the sliding window size used to generate sub-sequences for modelling on prediction accuracy is studied. The results show that the Bayesian bio-basis function neural network with two Gaussian distributions for model parameters (weights) performed the best and the highest prediction accuracy is 97.15 +/- 1.13%. AVAILABILITY: The package of Bayesian bio-basis function neural network can be obtained by request to the author.

Algorithms↗

Efficiently mining gene expression data via a novel parameterless clustering method.

Clustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets.

Algorithms↗

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans↗

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans↗

Quantitative classification and natural clustering of Caenorhabditis elegans behavioral phenotypes.

Genetic analysis of nervous system function relies on the rigorous description of behavioral phenotypes. However, standard methods for classifying the behavioral patterns of mutant Caenorhabditis elegans rely on human observation and are therefore subjective and imprecise. Here we describe the application of machine learning to quantitatively define and classify the behavioral patterns of C. elegans nervous system mutants. We have used an automated tracking and image processing system to obtain measurements of a wide range of morphological and behavioral features from recordings of representative mutant types. Using principal component analysis, we represented the behavioral patterns of eight mutant types as data clouds distributed in multidimensional feature space. Cluster analysis using the k-means algorithm made it possible to quantitatively assess the relative similarities between different behavioral phenotypes and to identify natural phenotypic clusters among the data. Since the patterns of phenotypic similarity identified in this study closely paralleled the functional similarities of the mutant gene products, the complex phenotypic signatures obtained from these image data appeared to represent an effective diagnostic of the mutants' underlying molecular defects.

Animals↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core Cα RMSDs of 0.62 Å for MHC-I and 0.33 Å for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63 817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides↗

Biomarker discovery, disease classification, and similarity query processing on high-throughput MS/MS data of inborn errors of metabolism.

In newborn errors of metabolism, biomarkers are urgently needed for disease screening, diagnosis, and monitoring of therapeutic interventions. This article describes a 2-step approach to discover metabolic markers, which involves (1) the identification of marker candidates and (2) the prioritization of them based on expert knowledge of disease metabolism. For step 1, the authors developed a new algorithm, the biomarker identifier (BMI), to identify markers from quantified diseased versus normal tandem mass spectrometry data sets. BMI produces a ranked list of marker candidates and discards irrelevant metabolites based on a quality measure, taking into account the discriminatory performance, discriminatory space, and variance of metabolites' concentrations at the state of disease. To determine the ability of identified markers to classify subjects, the authors compared the discriminatory performance of several machine-learning paradigms and described a retrieval technique that searches and classifies abnormal metabolic profiles from a screening database. Seven inborn errors of metabolism-- phenylketonuria (PKU), glutaric acidemia type I (GA-I), 3-methylcrotonylglycinemia deficiency (3-MCCD), methylmalonic acidemia (MMA), propionic acidemia (PA), medium-chain acylCoAdehydrogenase deficiency (MCADD), and 3-OH long-chain acyl CoA dehydrogenase deficiency (LCHADD)-were investigated. All primarily prioritized marker candidates could be confirmed by literature. Some novel secondary candidates were identified (i.e., C16:1 and C4DC for PKU, C4DC for GA-I, and C18:1 forMCADD), which require further validation to confirm their biochemical role during health and disease.

Acyl-CoA Dehydrogenase↗

Machine learning approaches to lung cancer prediction from mass spectra.

We addressed the problem of discriminating between 24 diseased and 17 healthy specimens on the basis of protein mass spectra. To prepare the data, we performed mass to charge ratio (m/z) normalization, baseline elimination, and conversion of absolute peak height measures to height ratios. After preprocessing, the major difficulty encountered was the extremely large number of variables (1676 m/z values) versus the number of examples (41). Dimensionality reduction was treated as an integral part of the classification process; variable selection was coupled with model construction in a single ten-fold cross-validation loop. We explored different experimental setups involving two peak height representations, two variable selection methods, and six induction algorithms, all on both the original 1676-mass data set and on a prescreened 124-mass data set. Highest predictive accuracies (1-2 off-sample misclassifications) were achieved by a multilayer perceptron and Naïve Bayes, with the latter displaying more consistent performance (hence greater reliability) over varying experimental conditions. We attempted to identify the most discriminant peaks (proteins) on the basis of scores assigned by the two variable selection methods and by neural network based sensitivity analysis. These three scoring schemes consistently ranked four peaks as the most relevant discriminators: 11683, 1403, 17350 and 66107.

Algorithms↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

FiberID--a technique to identify fibrous protein subclasses.

Fibrous proteins such as collagen, silk, and elastin play critical biological roles, yet they have been the subject of few projects that use computational techniques to predict either their class or their structure. In this article, we present FiberID, a simple yet effective method for identifying and distinguishing three fibrous protein subclasses from their primary sequences. Using a combination of amino acid composition and fast Fourier measurements, FiberID can classify fibrous proteins belonging to these subclasses with high accuracy by using two standard machine learning techniques (decision trees and Naïve Bayesian classifiers). After presenting our results, we present several fibrous sequences that are regularly misclassified by FiberID as sequences of potential interest for further study. Finally, we analyze the decision trees developed by FiberID for potential insights regarding the structure of these proteins.

Algorithms↗

An entity tagger for recognizing acquired genomic variations in cancer literature.

VTag is an application for identifying the type, genomic location and genomic state-change of acquired genomic aberrations described in text. The application uses a machine learning technique called conditional random fields. VTag was tested with 345 training and 200 evaluation documents pertaining to cancer genetics. Our experiments resulted in 0.8541 precision, 0.7870 recall and 0.8192 F-measure on the evaluation set.

Abstracting and Indexing↗