PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

CAKL: Commutative algebra k-mer learning of genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer learning (CAKL) as the first-ever nonlinear algebraic framework for analyzing genomic sequences. CAKL bridges between commutative algebra, algebraic topology, combinatorics, and machine learning to establish a new mathematical paradigm for comparative genomic analysis. We evaluate its effectiveness on three tasks-genetic variant identification, phylogenetic tree analysis, and viral genome classification-typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. Across eleven datasets, CAKL outperforms five state-of-the-art sequence analysis methods, particularly in viral classification, and maintains stable predictive accuracy as dataset size increases, underscoring its scalability and robustness. This work ushers in a new era in commutative algebraic data analysis and learning.

Journal Article↗

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids↗

Application of genetic search in derivation of matrix models of peptide binding to MHC molecules.

T cells of the vertebrate immune system recognise peptides bound by major histocompatibility complex (MHC) molecules on the surface of host cells. Peptide binding to MHC molecules is necessary for immune recognition, but only a subset of peptides are capable of binding to a particular MHC molecule. Common amino acid patterns (binding motifs) have been observed in sets of peptides that bind to specific MHC molecules. Recently, matrix models for peptide/MHC interaction have been reported. These encode the rules of peptide/ MHC interactions for an individual MHC molecule as a 20 x 9 matrix where the contribution to binding of each amino acid at each position within a 9-mer peptide is quantified. The artificial intelligence techniques of genetic search and machine learning have proved to be very useful in the area of biological sequence analysis. The availability of peptide/MHC binding data can facilitate derivation of binding matrices using machine learning techniques. We performed a simulation study to determine the minimum number of peptide samples required to derive matrices, given the pre-defined accuracy of the matrix model. The matrices were derived using a genetic search. In addition, matrices for peptide binding to the human class I MHC molecules, HLA-B35 and -A24, were derived, validated by independent experimental data and compared to previously-reported matrices. The results indicate that at least 150 peptide samples are required to derive matrices of acceptable accuracy. This result is based on a maximum noise content of 5%, the availability of precise affinity measurements and that acceptable accuracy is determined by an area under the Relative Operating Characteristic curve (Aroc) of > 0.8. More than 600 peptide samples are required to derive matrices of excellent accuracy (Aroc > 0.9). Finally, we derived a human HLA-B27 binding matrix using a genetic search and 404 experimentally-tested peptides, and estimated its accuracy at Aroc > 0.88. The results of this study are expected to be of practical interest to immunologists for efficient identification of peptides as candidates for immunotherapy.

Amino Acid Sequence↗

Computer-derived nuclear features distinguish malignant from benign breast cytology.

This article describes the use of computer-based analytical techniques to define nuclear size, shape, and texture features. These features are then used to distinguish between benign and malignant breast cytology. The benign and malignant cell samples used in this study were obtained by fine needle aspiration (FNA) from a consecutive series of 569 patients: 212 with cancer and 357 with fibrocystic breast masses. Regions of FNA preparations to be analyzed were converted by a video camera to computer files that were displayed on a computer monitor. Nuclei to be analyzed were roughly outlined by an operator using a mouse. Next, the computer generated a "snake" that precisely enclosed each designated nucleus. The computer calculated 10 features for each nucleus. The ability to correctly classify samples as benign or malignant on the basis of these features was determined by inductive machine learning and logistic regression. Cross-validation was used to test the validity of the predicted diagnosis. The logistic regression cross validated classification accuracy was 96.2% and the inductive machine learning cross-validated classification accuracy was 97.5%. Our computerized system provides a probability that a sample is malignant. Should this probability fall between 30% and 70%, the sample is considered "suspicious," in the same way a visually graded FNA may be termed suspicious. All of the 128 consecutive cases obtained since the introduction of this system were correctly diagnosed, but nine benign aspirates fell into the suspicious category.(ABSTRACT TRUNCATED AT 250 WORDS)

Breast↗

A molecular map of mesenchymal tumors.

BACKGROUND: Bone and soft tissue tumors represent a diverse group of neoplasms thought to derive from cells of the mesenchyme or neural crest. Histological diagnosis is challenging due to the poor or heterogenous differentiation of many tumors, resulting in uncertainty over prognosis and appropriate therapy. RESULTS: We have undertaken a broad and comprehensive study of the gene expression profile of 96 tumors with representatives of all mesenchymal tissues, including several problem diagnostic groups. Using machine learning methods adapted to this problem we identify molecular fingerprints for most tumors, which are pathognomonic (decisive) and biologically revealing. CONCLUSION: We demonstrate the utility of gene expression profiles and machine learning for a complex clinical problem, and identify putative origins for certain mesenchymal tumors.

Gene Expression Profiling↗

Microarray-based cancer diagnosis with artificial neural networks.

In recent years, the advent of experimental methods to probe gene expression profiles of cancer on a genome-wide scale has led to widespread use of supervised machine learning algorithms to characterize these profiles. The main applications of these analysis methods range from assigning functional classes of previously uncharacterized genes to classification and prediction of different cancer tissues. This article surveys the application of machine learning algorithms to classification and diagnosis of cancer based on expression profiles. To exemplify the important issues of the classification procedure, the emphasis of this article is on one such method, namely artificial neural networks. In addition, methods to extract genes that are important for the performance of a classifier, as well as the influence of sample selection on prediction results are discussed.

Algorithms↗

The automatic discovery of structural principles describing protein fold space.

The study of protein structure has been driven largely by the careful inspection of experimental data by human experts. However, the rapid determination of protein structures from structural-genomics projects will make it increasingly difficult to analyse (and determine the principles responsible for) the distribution of proteins in fold space by inspection alone. Here, we demonstrate a machine-learning strategy that automatically determines the structural principles describing 45 folds. The rules learnt were shown to be both statistically significant and meaningful to protein experts. With the increasing emphasis on high-throughput experimental initiatives, machine-learning and other automated methods of analysis will become increasingly important for many biological problems.

Algorithms↗

Tumour class prediction and discovery by microarray-based DNA methylation analysis.

Aberrant DNA methylation of CpG sites is among the earliest and most frequent alterations in cancer. Several studies suggest that aberrant methylation occurs in a tumour type-specific manner. However, large-scale analysis of candidate genes has so far been hampered by the lack of high throughput assays for methylation detection. We have developed the first microarray-based technique which allows genome-wide assessment of selected CpG dinucleotides as well as quantification of methylation at each site. Several hundred CpG sites were screened in 76 samples from four different human tumour types and corresponding healthy controls. Discriminative CpG dinucleotides were identified for different tissue type distinctions and used to predict the tumour class of as yet unknown samples with high accuracy using machine learning techniques. Some CpG dinucleotides correlate with progression to malignancy, whereas others are methylated in a tissue-specific manner independent of malignancy. Our results demonstrate that genome-wide analysis of methylation patterns combined with supervised and unsupervised machine learning techniques constitute a powerful novel tool to classify human cancers.

Algorithms↗

KAVAS-2: Knowledge Acquisition, Visualization and Assessment System.

The objective of KAVAS-2 is the development of a tool, named KAVIAR, with which domain experts can make their knowledge explicit. It contains components for (computer assisted) knowledge elicitation and for machine learning. A key issue in KAVAS is the assessment of the quality of the classification and domain models built. Various quality measures are available and implemented in KAVIAR to assess the quality of models, specifically those developed from data bases by machine learning techniques.

Computer Simulation↗

ABC stenosis morphology classification and outcome of coronary angioplasty: reassessment with computing techniques.

BACKGROUND: The American College of Cardiology/American Heart Association (ACC/AHA) stenosis morphology classification (MC) stratifies coronary lesions for probability of success and complications after coronary angioplasty (PTCA). Modern computing techniques were used to evaluate the individual predictive value of MC in random PTCA cases. METHODS AND RESULTS: MC was attributed to the target lesions by consensus of 2 observers. The predictive value regarding procedural success (PS) and major adverse cardiac events (MACE) of MC was analyzed by conventional logistic regression analyses and by inductive machine learning models. The study was adequately powered for the methods applied with 325 target lesions of 250 cases. Overall, PS decreased and MACE increased from type A to type C lesions. Regression analysis identified no single factor as predictive. Logistic regression showed an error rate of 42%. Machine learning techniques achieved an individual predictive error of only 10%, which could be further reduced to 2% by addition of parameters. For PS, MC parameters showed a high ranking for building the model. For MACE, variables of the medical history showed more impact. CONCLUSIONS: MC per se cannot individually predict PS or MACE. However, when all MC parameters are integrated together with additional lesion-specific and history variables, a high individual predictive value can be achieved. This technique may be clinically helpful for risk stratification in the catheterization laboratory and improvement of classification systems in interventional cardiology.

Algorithms↗

Representation for discovery of protein motifs.

There are several dimensions and levels of complexity in which information on protein motifs may be available. For example, one-dimensional sequence motifs may be associated with secondary structure identifiers. Alternatively, three-dimensional information on polypeptide segments may be used to induce prototypical three-dimensional structure templates. This paper surveys various representations encountered in the protein motif discovery literature. Many of the representations are based on incompatible semantics, making difficult the comparison and combination of previous results. To make better use of machine learning techniques and to provide for an integrated knowledge representation framework, a general representation language--in which all types of motifs can be encoded and given a uniform semantics--is required. In this paper we propose such a model, called a spatial description logic, and present a machine learning approach based on the model.

Amino Acids↗

Automated epiluminescence microscopy--tissue counter analysis using CART and 1-NN in the diagnosis of Melanoma.

BACKGROUND/PURPOSE: In tissue counter analysis, digital images are overlayed with regularly distributed measuring masks (elements) of equal size and shape, and the digital contents (grey level, colour and texture parameters) of each element are used for statistical analysis. In this study we assessed the applicability of tissue counter analysis and machine learning algorithms on tumour segmentation and diagnostic discrimination of benign and malignant melanocytic skin lesions. METHODS: A total of 369 standardised dermatoscopic images (93 melanomas, 276 benign nevi) were evaluated. The Classification and Regression Tree (CART) analysis was performed in order to differentiate between melanocytic skin lesions and surrounding skin. Instance-based learning (1-NN) was tested for differentiating between benign and malignant tumour elements. For diagnostic assessment, only the percentage of elements suggestive for malignancy in each lesion was used. RESULTS: Evaluation of a total of 369 melanocytic skin lesions showed a suitable segmentation of the tumour portion in 97.6%. When instance-based learning was applied to an independent test set, a threshold value of 27.4% of elements suggestive for malignancy recognised 35 out of 35 melanomas and 100 out of 101 nevi (sensitivity 100%, specificity 99%, positive predictive value 97.2%, negative predictive value 100%). CONCLUSION: Tissue counter analysis combined with machine learning algorithms turned out to be a useful method for diagnostic purposes in epiluminescence microscopy.

Algorithms↗

Using Bayesian networks in the construction of a bi-level multi-classifier. A case study using intensive care unit patients data.

Combining the predictions of a set of classifiers has shown to be an effective way to create composite classifiers that are more accurate than any of the component classifiers. There are many methods for combining the predictions given by component classifiers. We introduce a new method that combine a number of component classifiers using a Bayesian network as a classifier system given the component classifiers predictions. Component classifiers are standard machine learning classification algorithms, and the Bayesian network structure is learned using a genetic algorithm that searches for the structure that maximises the classification accuracy given the predictions of the component classifiers. Experimental results have been obtained on a datafile of cases containing information about ICU patients at Canary Islands University Hospital. The accuracy obtained using the presented new approach statistically improve those obtained using standard machine learning methods.

Algorithms↗

Proteomics-Driven Strategies for Proximity-Inducing Drug Discovery.

In recent years, proximity-inducing drugs have emerged as a novel therapeutic modality that induces or stabilizes protein-protein interactions, especially by recruiting effector proteins to specific target proteins, thereby achieving functions beyond traditional inhibitors. The potential of proximity-inducing drugs extends beyond targeted protein degradation (TPD), as studies have demonstrated their ability to regulate biological processes such as signal transduction, gene transcription, chromatin regulation, and protein trafficking by modulating protein interaction networks. Rational discovery of proximity-inducing drugs requires clarifying their effects on protein-protein interactions, determining drug selectivity, and developing suitable ligands for drug construction. Proteomics has become a central technology in drug discovery, enabling global identification of the direct drug targets and systematic characterization of proteome-wide downstream responses. This provides a more refined map of drug mechanisms. In parallel, advances in machine learning applied to proteomic data, together with the expansion of proteome-wide ligandability maps, are further accelerating the discovery and optimization of proximity-inducing drugs. This review summarizes recent advances of proximity-inducing drugs, with a particular emphasis on how proteomics facilitates target space expansion, drug efficacy optimization, and ligandability discovery, alongside the emerging contributions of machine learning. Collectively, these insights aim to support the rational development of next-generation proximity-inducing drugs.

Drug Discovery↗

Dietary Polyphenol Acteoside-Related Molecular Signatures in Clear Cell Renal Cell Carcinoma: Multi-Omics Profiling and Functional Validation of IMPDH1.

Clear cell renal cell carcinoma (ccRCC) is characterized by substantial metabolic and molecular heterogeneity, but the disease-relevant programs associated with acteoside, a dietary polyphenol, remain poorly understood. We integrated predicted acteoside targets with bulk, single-cell, and spatial transcriptomic data from ccRCC and combined molecular subtyping with cross-cohort machine-learning analysis. Acteoside-related signatures were preferentially enriched in malignant compartments and increased with tumor grade and stage. Consensus clustering identified two molecular subtypes with distinct biological and clinical features. C1 was associated with immune activation, metabolic activity, and more favorable survival, whereas C2 showed greater genomic instability, reduced renal epithelial differentiation, and poorer outcomes. We further benchmarked multiple machine-learning strategies and established a 10-gene prognostic model that retained predictive performance across independent cohorts, with IMPDH1 emerging as the strongest risk-associated feature. Functional experiments confirmed the biological relevance of IMPDH1: its knockdown suppressed ccRCC cell proliferation, DNA synthesis, colony formation, and migration, whereas overexpression produced the opposite effects. Together, these findings indicate that acteoside-related molecular signatures capture clinically relevant heterogeneity in ccRCC and provide a framework for linking dietary-polyphenol-related molecular space with tumor biology. The identification and functional validation of IMPDH1 further highlight its potential importance in ccRCC progression.

IMPDH1↗

Visual management of large scale data mining projects.

This paper describes a unified framework for visualizing the preparations for, and results of, hundreds of machine learning experiments. These experiments were designed to improve the accuracy of enzyme functional predictions from sequence, and in many cases were successful. Our system provides graphical user interfaces for defining and exploring training datasets and various representational alternatives, for inspecting the hypotheses induced by various types of learning algorithms, for visualizing the global results, and for inspecting in detail results for specific training sets (functions) and examples (proteins). The visualization tools serve as a navigational aid through a large amount of sequence data and induced knowledge. They provided significant help in understanding both the significance and the underlying biological explanations of our successes and failures. Using these visualizations it was possible to efficiently identify weaknesses of the modular sequence representations and induction algorithms which suggest better learning strategies. The context in which our data mining visualization toolkit was developed was the problem of accurately predicting enzyme function from protein sequence data. Previous work demonstrated that approximately 6% of enzyme protein sequences are likely to be assigned incorrect functions on the basis of sequence similarity alone. In order to test the hypothesis that more detailed sequence analysis using machine learning techniques and modular domain representations could address many of these failures, we designed a series of more than 250 experiments using information-theoretic decision tree induction and naive Bayesian learning on local sequence domain representations of problematic enzyme function classes. In more than half of these cases, our methods were able to perfectly discriminate among various possible functions of similar sequences. We developed and tested our visualization techniques on this application.

Alcohol Dehydrogenase↗

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals↗

Integrative TWAS and multi-omics analyses prioritize HSPE1 as a candidate risk gene for bipolar disorder with immune cell-specific regulatory evidence.

BACKGROUND: Bipolar disorder (BD) is a severe psychiatric disorder associated with substantial disability. Although genome-wide association studies have identified multiple BD-associated loci, the underlying genes and mechanisms remain incompletely understood. METHODS: We integrated a European-ancestry BD genome-wide association dataset with cross-tissue and tissue-specific transcriptome-wide association studies (TWAS) and complementary gene-based analysis. Candidate genes were further evaluated using differential expression analysis, consensus clustering, immune infiltration analysis, machine learning, summary-data-based Mendelian randomization, Mendelian randomization using single-cell expression quantitative trait locus data, single-nucleus transcriptomics, phenome-wide association analysis, and virtual screening. RESULTS: The integrative analyses prioritized 37 candidate genes. Peripheral-blood differential-expression analysis identified 14 genes that remained significant after FDR correction, and their expression profiles separated BD samples into two expression-defined clusters. Machine-learning analysis selected UNC50, LMAN2L, LYG2, HSPE1, and KANSL3 for an exploratory classification nomogram. SMR associated genetically predicted higher HSPE1 expression with increased BD risk in two blood eQTL datasets. Cell-type-specific analyses indicated HSPE1-related associations in T-cell and natural killer cell subsets, while single-nucleus analysis descriptively showed higher HSPE1 expression in medial thalamic T cells from BD samples. PheWAS identified no genome-wide significant associations for HSPE1, whereas virtual screening identified candidate compounds with favorable predicted docking scores against the HSPE1 structure. CONCLUSION: This integrative multi-omics study identified HSPE1 as a candidate BD risk gene with immune-cell-related regulatory evidence, providing insight into BD pathogenesis and supporting functional validation.

Humans↗