PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Design of nanobody targeting SARS-CoV-2 spike glycoprotein using CDR-grafting assisted by molecular simulation and machine learning.

The design of proteins capable effectively binding to specific protein targets is crucial for developing therapies, diagnostics, and vaccine candidates for viral infections. Here, we introduce a complementarity-determining region (CDR) grafting approach for designing nanobodies (Nbs) that target specific epitopes, with the aid of computer simulation and machine learning. As a proof-of-concept, we designed, evaluated, and characterized a high-affinity Nb against the spike protein of SARS-CoV-2, the causative agent of the COVID-19 pandemic. The designed Nb, referred to as Nb Ab.2, was synthesized and displayed high-affinity for both the purified receptor-binding domain protein and to the virus-like particle, demonstrating affinities of 9 nM and 60 nM, respectively, as measured with microscale thermophoresis. Circular dichroism showed the designed protein's structural integrity and its proper folding, whereas molecular dynamics simulations provided insights into the internal dynamics of Nb Ab.2. This study shows that our computational pipeline can be used to efficiently design high-affinity Nbs with diagnostic and prophylactic potential, which can be tailored to tackle different viral targets.

Spike Glycoprotein, Coronavirus↗

Machine learning identifies ac4C-related prognostic signature and TUBA1C as therapeutic target in COAD.

To explore the role of N4-acetylcytidine (ac4C)-related genes (acRGs) in colon adenocarcinoma (COAD) and identify reliable prognostic biomarkers and potential therapeutic targets. Multi-source transcriptomic datasets (TCGA-COAD, GSE39582, GSE17536) and single-cell RNA-seq data were analyzed. Ten machine learning algorithms were integrated to construct an acRG-based prognostic signature (acRGBS). Immune microenvironment (TME) and genomic profiling were performed, with in vitro functional experiments validating TUBA1C's role. acRGBS, comprising four hub genes (SARAF, CDC42SE2, TSPYL2, TUBA1C), effectively stratified COAD patients into high- and low-risk groups with distinct survival outcomes and was an independent prognostic factor. High-risk patients exhibited increased genomic instability and immunosuppressive TME, while low-risk patients had favorable immunotherapy response. TUBA1C was overexpressed in COAD cells, and its knockdown inhibited proliferation/migration and induced apoptosis. The acRGBS is a robust prognostic tool for COAD, and TUBA1C serves as a candidate therapeutic target, providing new insights for personalized COAD management.

Humans↗

Machine learning for sub-population assessment: evaluating the C-section rate of different physician practices.

We apply machine learning to the problem of subpopulation assessment for Caesarian Section. In subpopulation assessment, we are interested in making predictions not for a single patient, but for groups of patients. Typically, in any large population, different subpopulations will have different "outcome" rates. In our example, the C-section rate of a population of 22,176 expectant mothers is 16.8%; yet, the 17 physician groups that serve this population have vastly different group C-section rates, ranging from 11% to 23%. The ultimate goal of subpopulation assessment is to determine if these variations in the observed rates can be attributed to (a) variations in intrinsic risk of the patient sub-populations (i.e. some groups contain more "high-risk C-section" patients), or (b) differences in physician practice (i.e. some groups do more C-sections). Our results indicate that although there is some variation in intrinsic risk, there is also much variation in physician practice.

Artificial Intelligence↗

A study on several machine-learning methods for classification of malignant and benign clustered microcalcifications.

In this paper, we investigate several state-of-the-art machine-learning methods for automated classification of clustered microcalcifications (MCs). The classifier is part of a computer-aided diagnosis (CADx) scheme that is aimed to assisting radiologists in making more accurate diagnoses of breast cancer on mammograms. The methods we considered were: support vector machine (SVM), kernel Fisher discriminant (KFD), relevance vector machine (RVM), and committee machines (ensemble averaging and AdaBoost), of which most have been developed recently in statistical learning theory. We formulated differentiation of malignant from benign MCs as a supervised learning problem, and applied these learning methods to develop the classification algorithm. As input, these methods used image features automatically extracted from clustered MCs. We tested these methods using a database of 697 clinical mammograms from 386 cases, which included a wide spectrum of difficult-to-classify cases. We analyzed the distribution of the cases in this database using the multidimensional scaling technique, which reveals that in the feature space the malignant cases are not trivially separable from the benign ones. We used receiver operating characteristic (ROC) analysis to evaluate and to compare classification performance by the different methods. In addition, we also investigated how to combine information from multiple-view mammograms of the same case so that the best decision can be made by a classifier. In our experiments, the kernel-based methods (i.e., SVM, KFD, and RVM) yielded the best performance (Az = 0.85, SVM), significantly outperforming a well-established, clinically-proven CADx approach that is based on neural network (Az = 0.80).

Algorithms↗

DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins.

INTRODUCTION: The fundamental cellular processes, including transcriptional regulation, chromatin organization, and genome maintenance, are regulated by DNA-binding proteins (DBPs). Mutations in DBPs can alter protein-DNA interactions, leading to tumor development. However, identifying such driver mutations remains a major challenge due to limitations of experimental approaches. METHODS: We have trained a machine learning model, DBP-CanPred, to identify driver mutations in DBPs. We used the sequence-derived evolutionary features, as well as structure-based features such as mutation-perturbed structural descriptors. RESULTS: We evaluated DBP-CanPred using a curated test set, achieving an AU-ROC of 0.86 and a balanced accuracy of 0.79. Further analysis based on substitution-type showed consistent performance across different categories, especially higher performance on charged residues. In addition, we applied the model on an independent dataset and identified potential driver mutations with high confidence scores. DISCUSSION: The study contributes to understanding mutation patterns in DNA-binding proteins and supports variant interpretation in cancer research.

DNA-binding proteins↗

Deep assessment of machine learning techniques using patient treatment in acute abdominal pain in children.

Learning from patient records may aid knowledge acquisition and decision making. Existing inductive machine learning (ML) systems such us NewId, CN2, C4.5 and AQ15 learn from past case histories using symbolic and/or numeric values. These systems learn symbolic rules (IF... THEN like) which link an antecedent set of clinical factors to a consequent class or decision. This paper compares the learning performance of alternative ML systems with each other and with respect to a novel approach using logic minimization, called LML, to learn from data. Patient cases were taken from the archives of the Paediatric Surgery Clinic of the University Hospital of Crete, Heraklion, Greece. Comparison of ML system performance is based both on classification accuracy and on informal expert assessment of learned knowledge.

Abdomen, Acute↗

Knowledge-based generation of machine learning experiments: learning with DNA crystallography data.

Though it has been possible in the past to learn to predict DNA hydration patterns from crystallographic data, there is ambiguity in the choice of training data (both in terms of the relevant set of cases and the features needed to represent them), which limits the usefulness of standard learning techniques. Thus, we have developed a knowledge-based system to generate machine learning experiments for inducing DNA hydration pattern classifiers. The system takes as input (1) a set of classified training examples described by a large set of attributes and (2) information about a set of learning experiments that have already been run. It outputs a new learning experiment, namely a (not necessarily proper) subset of the input examples represented by a new set of features. Domain specific and domain independent knowledge is used to suggest subsets of training examples from suspected subpopulations, transform attributes in the training data or generate new ones, and choose interesting ways to substitute one experiment's set of attributes with another. Automatic hydration pattern predictors are of both theoretical and practical interest to DNA crystallographers, because they can speed up a labor intensive process, and because the extracted rules add to the knowledge of what determines DNA hydration.

Artificial Intelligence↗

Aberrant mucin expression and keratinization distinguishing severe from mild asthma revealed by interpretable machine learning.

Type 2 (T2) immune cells dominate the airways of patients with mild-moderate asthma (MMA) with a more complex type 1 (T1)-T2 mixed immune response evident in treatment-refractory severe asthma (SA). We hypothesized that comparing the transcriptomes of the airway epithelium of patients with SA and MMA would reveal molecular signatures associated with more severe disease in the context of a complex immune response. Using our interpretable machine learning tool, SLIDE, meaningful latent factors (context-specific gene co-expression networks) were revealed that distinguished SA from MMA. Unexpectedly, an aberrant high expression of normally host-protective, membrane-tethered, and IFN-inducible mucins, MUC1 and MUC4, was identified in SA. Gene networks in the significant latent factors discriminating SA from MMA corresponded to enrichment of a keratinization program in SA airways. Keratinization was marked by increased expression of the stress keratin KRT16, signifying squamous metaplasia suggesting adaptive reprogramming of the airway epithelium in response to chronic stress. These mucins and KRT16 were inversely associated with lung function in 2 separate asthma cohorts. Imaging of endobronchial biopsies revealed significantly higher KRT16 protein expression in SA compared with MMA that strongly correlated with MUC1 protein expression. Our study identifies dysregulated host-protective and maladaptive repair responses in SA distinguishing from MMA.

Humans↗

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged ≥65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Classifying "kinase inhibitor-likeness" by using machine-learning methods.

By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued--support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GA/kNN), and recursive partitioning (RP)--proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.

Algorithms↗

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning↗

Effects of information and machine learning algorithms on word sense disambiguation with small datasets.

Current approaches to word sense disambiguation use (and often combine) various machine learning techniques. Most refer to characteristics of the ambiguity and its surrounding words and are based on thousands of examples. Unfortunately, developing large training sets is burdensome, and in response to this challenge, we investigate the use of symbolic knowledge for small datasets. A naïve Bayes classifier was trained for 15 words with 100 examples for each. Unified Medical Language System (UMLS) semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in nine experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. To investigate this large variance, we performed several follow-up evaluations, testing additional algorithms (decision tree and neural network), and gold standards (per expert), but the results did not significantly differ. However, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators. We conclude that neither algorithm nor individual human behavior cause these large differences, but that the structure of the UMLS Metathesaurus (used to represent senses of ambiguous words) contributes to inaccuracies in the gold standard, leading to varied performance of word sense disambiguation techniques.

Algorithms↗

Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression data.

BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains.

Animals↗

A comparative study of machine-learning methods to predict the effects of single nucleotide polymorphisms on protein function.

MOTIVATION: The large volume of single nucleotide polymorphism data now available motivates the development of methods for distinguishing neutral changes from those which have real biological effects. Here, two different machine-learning methods, decision trees and support vector machines (SVMs), are applied for the first time to this problem. In common with most other methods, only non-synonymous changes in protein coding regions of the genome are considered. RESULTS: In detailed cross-validation analysis, both learning methods are shown to compete well with existing methods, and to out-perform them in some key tests. SVMs show better generalization performance, but decision trees have the advantage of generating interpretable rules with robust estimates of prediction confidence. It is shown that the inclusion of protein structure information produces more accurate methods, in agreement with other recent studies, and the effect of using predicted rather than actual structure is evaluated. AVAILABILITY: Software is available on request from the authors.

Algorithms↗

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-α/NF-κB, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans↗

Seeing and Feeling DNA Methylation: Single-Molecule Biophysics Meets Machine Learning.

DNA methylation at 5-methylcytosine (5mC) is crucial for embryonic development and cellular function, while aberrant patterns strongly drive disease onset and progression. Its reversible nature offers substantial therapeutic potential, emphasizing the need for precise, context-specific genome wide 5mC mapping. Conventional techniques such as bisulfite sequencing and ensemble biosensor assays are hindered by DNA degradation, amplification bias, high cost, and inability to resolve single-molecule structural and mechanical effects of methylation. This review examines advances in single-molecule biophysical methods (nanopore sensing, smFRET, optical/magnetic tweezers, and AFM) that provide direct, label-free/minimally invasive 5mC detection, along with quantitative insights into DNA conformation, mechanics, and protein-DNA interactions. These techniques complement traditional methylome mapping by linking genomic localization to molecular mechanisms. Emerging machine-learning approaches are revolutionizing analysis, particularly in nanopore sensing, while promising applications in smFRET, tweezers, and AFM address throughput and reproducibility challenges. Their convergence promises scalable, high-resolution epigenetic profiling, advancing precision epigenomics toward clinical application.

DNA Methylation↗

Exploring an Intermediate Colorectal Cancer Screening Test Based on Stool Proteomics and Machine Learning for Optimizing the Selection of Patients for Colonoscopy Identified From FIT.

The fecal immunochemical test (FIT) for detecting fecal occult blood, used alone or in combination with other stool biomarkers, has been demonstrated to be effective in the context of colorectal cancer (CRC) screening programs. However, FIT yields a significant proportion of false positives leading to unnecessary colonoscopies. In this study, we have investigated whether leftover FIT stool samples could be repurposed for proteomics analysis as a triage step for patients before recommending colonoscopy. High-throughput mass spectrometry analyses on a set of 141 FIT-positive samples (50 controls with no lesion, 45 with advanced adenomas and 46 with CRC) in combination with machine learning tools were used. Results showed that with a specificity ≥90%, a large proportion of the false FIT positives could be identified thus providing an efficient strategy for reducing unnecessary colonoscopies. Furthermore, CRC cases were also precisely predicted to be true positives, thus providing an approach for prioritizing patients for colonoscopy. In conclusion, this study demonstrates the feasibility of using proteomics for analysis of leftover FIT stool samples as an intermediate step to triage patients selected for colonoscopy in CRC screening programs.

Humans↗