PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

A novel structure-based encoding for machine-learning applied to the inference of SH3 domain specificity.

MOTIVATION: Unravelling the rules underlying protein-protein and protein-ligand interactions is a crucial step in understanding cell machinery. Peptide recognition modules (PRMs) are globular protein domains which focus their binding targets on short protein sequences and play a key role in the frame of protein-protein interactions. High-throughput techniques permit the whole proteome scanning of each domain, but they are characterized by a high incidence of false positives. In this context, there is a pressing need for the development of in silico experiments to validate experimental results and of computational tools for the inference of domain-peptide interactions. RESULTS: We focused on the SH3 domain family and developed a machine-learning approach for inferring interaction specificity. SH3 domains are well-studied PRMs which typically bind proline-rich short sequences characterized by the PxxP consensus. The binding information is known to be held in the conformation of the domain surface and in the short sequence of the peptide. Our method relies on interaction data from high-throughput techniques and benefits from the integration of sequence and structure data of the interacting partners. Here, we propose a novel encoding technique aimed at representing binding information on the basis of the domain-peptide contact residues in complexes of known structure. Remarkably, the new encoding requires few variables to represent an interaction, thus avoiding the 'curse of dimension'. Our results display an accuracy >90% in detecting new binders of known SH3 domains, thus outperforming neural models on standard binary encodings, profile methods and recent statistical predictors. The method, moreover, shows a generalization capability, inferring specificity of unknown SH3 domains displaying some degree of similarity with the known data.

Algorithms↗

Classification of substrates and inhibitors of P-glycoprotein using unsupervised machine learning approach.

P-glycoprotein (P-gp), a drug efflux pump, affects the bioavailability of therapeutic drugs and plays a potentially important role in clinical drug-drug interactions. Classification of candidate drugs as substrates or inhibitors of the carrier protein is of crucial importance in drug development. Accurate classification is difficult to achieve due to two major factors: i. The extreme diversity of substrates and the presence of multiple binding sites complicate the understanding of the mechanisms behind and hinder the development of a true, conclusive quantitative structure-activity relationship (QSAR) for P-gp substrates. ii. Both inhibitors and substrates interact with the same binding site of P-gp, as a result, it is not surprising that both share many common structural features. In this work, an unsupervised machine learning approach based on the Kohonen self-organizing maps (SOM) was explored, which incorporated a predefined set of physicochemical descriptors encoding the key molecular properties capable of discerning a substrate from an inhibitor. The SOM model can discriminate between substrates and inhibitors with an average accuracy of 82.3%. The current results show that the SOM-based method provides a potential in silico model for virtual screening.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

The machine-learning classifier ALLCatchR2 identifies 20 T-ALL subtypes across cohorts and age groups.

T-cell acute lymphoblastic leukemia (T-ALL) comprises molecularly diverse subtypes, but robust cross-cohort validations and operational gene-expression definitions are lacking. To establish a gene-expression-anchored framework for T-ALL subtyping, we aggregated 2314 transcriptomes (15 cohorts, age: 0.8-90.8 years). An extended unsupervised approach defined 17 main clusters and 3 subclusters in samples with high blast fractions. Supervised analyses added an overarching immature T-ALL (early T cell precursor [ETP]-like) definition and resolved the LMO2 &#x3b3;&#x3b4;-like subtype. All clusters contained samples from at least two cohorts. Characteristic genomic driver enrichments were consistent across cohorts, while gene-expression clusters did not correspond exclusively to single driver events but also reflected developmental origins. A machine-learning classifier based on ALLCatchR, our B-cell acute lymphoblastic leukemia (B-ALL) classifier, identified these 20 transcriptomic subtypes and the immature T-ALL (ETP-like) signature with 0.995-1.0 accuracy in a validation set (n&#x2009;=&#x2009;203). Testing the classifier on a second hold-out data set (n&#x2009;=&#x2009;265 samples) showed that 92.7% of predictions matched with corresponding driver alterations. Across all samples, 83.2% of cases received high-confidence predictions, 7.3% candidate predictions, and 9.5% remained unclassified, largely because of low blast fractions. We identified a novel gene-expression cluster markedly enriched (P&#x2009;<&#x2009;0.001) for clonal hematopoiesis mutations (IDH2 R140Q, DNMT3A) and a stem-/progenitor cell-like gene expression. This novel clonal hematopoiesis-related T-ALL subtype was observed in six cohorts and accounted for 8.9% of adults and 39.5% of patients aged >50 years. We extended&#xa0;ALLCatchR into ALLCatchR2, a free R package that now enables B-/T-lineage separation, gene-expression subtyping, blast estimation, and developmental annotation to harmonize T-ALL classification across studies and clinical contexts.

Journal Article↗

Acquiring background knowledge for machine learning using function decomposition: a case study in rheumatology.

Domain or background knowledge is often needed in order to solve difficult problems of learning medical diagnostic rules. Earlier experiments have demonstrated the utility of background knowledge when learning rules for early diagnosis of rheumatic diseases. A particular form of background knowledge comprising typical co-occurrences of several groups of attributes was provided by a medical expert. This paper explores the possibility of automating the process of acquiring background knowledge of this kind and studies the utility of such methods in the problem domain of rheumatic diseases. A method based on function decomposition is proposed that identifies typical co-occurrences for a given set of attributes. The method is evaluated by comparing the typical co-occurrences it identifies as well as their contribution to the performance of machine learning algorithms, to the ones provided by a medical expert.

Algorithms↗

Quantitative assessment of the fingerprint evidential value using machine learning.

Fingerprints as physical evidence have long supported criminal investigation and adjudication. In practice, however, fingerprint identification relies mainly on examiners' experience. Furthermore, expert opinions tend to be categorical, even though the opinions with the same conclusion could differ substantially in evidential strength. To quantitatively assess fingerprint evidential value, this study proposes a machine learning-based framework as an interpretable decision-support tool. A lightweight residual one-dimensional convolutional neural network was constructed, incorporating channel recalibration and a similarity-driven attention mechanism to learn adaptive contribution weights for different matched minutiae (minutiae for short). Controlled experiments revealed that the predicted evidential value increased with the number of minutiae and was significantly influenced by the quality of minutiae. With 10 minutiae, the mean predicted scores were 4.49, 7.00, and 9.09 for blurred, moderately blurred, and clear minutiae, respectively. Multiple regression analysis indicated that replacing a pair of blurred minutiae with a pair of clear minutiae increased the score by 0.492, whereas replacing it with a pair of moderately blurred minutiae increased the score by only 0.216. By mapping predicted scores to graded levels of evidential strength, the framework contributes to a paradigm shift from categorical expert opinions to graded ones, helping courts evaluate fingerprint evidence more scientifically.

Humans↗

A one-layer recurrent neural network for support vector machine learning.

This paper presents a one-layer recurrent neural network for support vector machine (SVM) learning in pattern classification and regression. The SVM learning problem is first converted into an equivalent formulation, and then a one-layer recurrent neural network for SVM learning is proposed. The proposed neural network is guaranteed to obtain the optimal solution of support vector classification and regression. Compared with the existing two-layer neural network for the SVM classification, the proposed neural network has a low complexity for implementation. Moreover, the proposed neural network can converge exponentially to the optimal solution of SVM learning. The rate of the exponential convergence can be made arbitrarily high by simply turning up a scaling parameter. Simulation examples based on benchmark problems are discussed to show the good performance of the proposed neural network for SVM learning.

Journal Article↗

Measuring Cell Dimensions in Fission Yeast Using Machine Learning.

In fission yeast (Schizosaccharomyces pombe), cell length is a crucial indicator of cell cycle progression. Microscopy screens that examine the effect of agents or genotypes suspected of altering genomic or metabolic stability and thus cell size are crucial for studying disruptions to cell cycle dynamics. This method is based on using an automated cell segmentation algorithm to measure S. pombe cells imaged by brightfield (BF) microscopy methods. PhotoPhenosizer (PP) is a machine learning-based tool designed for automated cell measuring and dimensional analysis of morphology frequency distributions. Integration of this method into large-scale pipelines for tracking cell dimension change streamlines morphological measurements, which facilitates the examination of cellular responses to genomic and metabolic stresses. In this protocol, we use PP to observe the effect of genomic instability on cell size dynamics over a 12-day chronological lifespan assay. Our results show that relative to wild-type cells, a replication stress mutant shows larger cells during chronological aging in excess glucose media. Our results are consistent with activation of checkpoints that regulate cell morphology in response to DNA damage. This method's application highlights the relevance of its incorporation in experimental routines that require large-scale image processing and its adoption by users with routine needs in S. pombe molecular research projects.

Schizosaccharomyces↗

Building an asynchronous web-based tool for machine learning classification.

Various unsupervised and supervised learning methods including support vector machines, classification trees, linear discriminant analysis and nearest neighbor classifiers have been used to classify high-throughput gene expression data. Simpler and more widely accepted statistical tools have not yet been used for this purpose, hence proper comparisons between classification methods have not been conducted. We developed free software that implements logistic regression with stepwise variable selection as a quick and simple method for initial exploration of important genetic markers in disease classification. To implement the algorithm and allow our collaborators in remote locations to evaluate and compare its results against those of other methods, we developed a user-friendly asynchronous web-based application with a minimal amount of programming using free, downloadable software tools. With this program, we show that classification using logistic regression can perform as well as other more sophisticated algorithms, and it has the advantages of being easy to interpret and reproduce. By making the tool freely and easily available, we hope to promote the comparison of classification methods. In addition, we believe our web application can be used as a model for other bioinformatics laboratories that need to develop web-based analysis tools in a short amount of time and on a limited budget.

Algorithms↗

Identification of MHC Ligands Through Allele-Guided Isolation Combined With Machine Learning for Improved MHC Assignment Using ARDisplay-I.

The isolation of major histocompatibility complex (MHC) ligands and subsequent analysis by mass spectrometry is considered the gold standard for defining targets for T cell-based immunotherapies. However, as many targets of high tumor specificity are only presented at low abundance on the cell surface of tumor cells, the efficient isolation of these peptides is crucial for their successful detection. Here, we demonstrate how optimizing the MHC ligand isolation strategy, based on both the presenting MHC alleles and the individual peptide level, enhances the identification of specific MHC ligands. This ideally acknowledges not only the hydrophobicity but also the post-translational modifications of the respective MHC ligands. To further improve the identification and characterization of MHC ligands, we developed an MHC class I ligand prediction algorithm (ARDisplay-I) that outperforms current state-of-the-art tools when benchmarked against competitors such as netMHCpan 4.1, MixMHCpred, or MHCflurry. Implementing these strategies can augment the development of T cell receptor-based therapies by improving the identification of novel immunotherapy targets and enriching the resources available in the computational immunology field through a superior MHC presentation prediction algorithm.

Ligands↗

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans↗

Extraction of gene-disease relations from Medline using domain dictionaries and machine learning.

We describe a system that extracts disease-gene relations from Medline. We constructed a dictionary for disease and gene names from six public databases and extracted relation candidates by dictionary matching. Since dictionary matching produces a large number of false positives, we developed a method of machine learning-based named entity recognition (NER) to filter out false recognitions of disease/gene names. We found that the performance of relation extraction is heavily dependent upon the performance of NER filtering and that the filtering improves the precision of relation extraction by 26.7% at the cost of a small reduction in recall.

Animals↗

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i)&#xa0;perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii)&#xa0;introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans↗

Construction of a molecular diagnostic system for neurogenic rosacea by combining transcriptome sequencing and machine learning.

Patients with neurogenic rosacea (NR) frequently demonstrate pronounced neurological manifestations, often unresponsive to conventional therapeutic approaches. A molecular-level understanding and diagnosis of this patient cohort could significantly guide clinical interventions. In this study, we amalgamated our sequencing data (n&#x2009;=&#x2009;46) with a publicly accessible database (n&#x2009;=&#x2009;38) to perform an unsupervised cluster analysis of the integrated dataset. The eighty-four rosacea patients were partitioned into two distinct clusters. Neurovascular biomarkers were found to be elevated in cluster 1 compared to cluster 2. Pathways in cluster 1 were predominantly involved in neurotransmitter synthesis, transmission, and functionality, whereas cluster 2 pathways were centered on inflammation-related processes. Differential gene expression analysis and WGCNA were employed to delineate the characteristic gene sets of the two clusters. Subsequently, a diagnostic model was constructed from the identified gene sets using linear regression methodologies. The model's C index, comprising genes PNPLA3, CUX2, PLIN2, and HMGCR, achieved a remarkable value of 0.9683, with an area under the curve (AUC) for the training cohort's nomogram of 0.9376. Clinical characteristics from our dataset (n&#x2009;=&#x2009;46) were assessed by three seasoned dermatologists, forming the NR validation cohort (NR, n&#x2009;=&#x2009;18; non-neurogenic rosacea, n&#x2009;=&#x2009;28). Upon application of our model to NR diagnosis, the model's AUC value reached 0.9023. Finally, potential therapeutic candidates for both patient groups were predicted via the Connectivity Map. In summation, this study unveiled two clusters with unique molecular phenotypes within rosacea, leading to the development of a precise diagnostic model instrumental in NR diagnosis.

Humans↗

Discovering hidden candidate plastic-degrading enzymes: Combined multi-omics and machine learning strategy.

Plastic pollution poses a major threat to the stability of natural ecosystems as well as human health. Microbial enzymes have long been considered a potential resource for targeted biodegradation but, except for a few successful cases, the discovery of efficient enzymes has proved challenging. Aiming to accelerate the process, we propose an approach combining metagenomics, metatranscriptomics and semi-supervised learning that selects promising plastic-degrading candidate enzymes from the proteome of relevant microorganisms. Tested on a dataset of over 10,000 microbial proteins, ranking models consistently prioritize known plastic-degrading enzymes, achieving an area under the cumulative distribution function curve above 0.96, with leave-one-family-out cross-validation indicating that performance is largely retained across protein families. As a case study, this work focuses on mixed microbial cultures exposed for extended periods to polyethylene, polyethylene terephthalate, and polyurethane substrates. The prevalent species after selective enrichment were functionally characterized, finding Rhodococcus aetherivorans as the most relevant species in two of the five cultures under investigation. Among the top-ranked proteins, several have high structural similarity with known enzymes despite not being identified by sequence similarity search. Moreover, according to metatranscriptomics results, several of these enzymes were found to be expressed at the same level or above that of annotated enzymes, suggesting that they may have functional relevance. Overall, this work highlights the potential of integrating multi-omics with data-driven methods for enzyme discovery and for accelerating the development of biotechnological solutions to plastic pollution.

Biodegradation, Environmental↗

Multi&#x2011;omics identification of a novel signature for serous ovarian carcinoma in the context of 3P medicine and based on twelve programmed cell death patterns: a multi-cohort machine learning study.

BACKGROUND: Predictive, preventive, and personalized medicine (PPPM/3PM) is a strategy aimed at improving the prognosis of cancer, and programmed cell death (PCD) is increasingly recognized as a potential target in cancer therapy and prognosis. However, a PCD-based predictive model for serous ovarian carcinoma (SOC) is lacking. In the present study, we aimed to establish a cell death index (CDI)-based model using PCD-related genes. METHODS: We included 1254 genes from 12 PCD patterns in our analysis. Differentially expressed genes (DEGs) from the Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) were screened. Subsequently, 14 PCD-related genes were included in the PCD-gene-based CDI model. Genomics, single-cell transcriptomes, bulk transcriptomes, spatial transcriptomes, and clinical information from TCGA-OV, GSE26193, GSE63885, and GSE140082 were collected and analyzed to verify the prediction model. RESULTS: The CDI was recognized as an independent prognostic risk factor for patients with SOC. Patients with SOC and a high CDI had lower survival rates and poorer prognoses than those with a low CDI. Specific clinical parameters and the CDI were combined to establish a nomogram that accurately assessed patient survival. We used the PCD-genes model to observe differences between high and low CDI groups. The results showed that patients with SOC and a high CDI showed immunosuppression and hardly benefited from immunotherapy; therefore, trametinib_1372 and BMS-754807 may be potential therapeutic agents for these patients. CONCLUSIONS: The CDI-based model, which was established using 14 PCD-related genes, accurately predicted the tumor microenvironment, immunotherapy response, and drug sensitivity of patients with SOC. Thus this model may help improve the diagnostic and therapeutic efficacy of PPPM.

Humans↗

Development and validation of a machine learning prognostic model based on an epigenomic signature in patients with pancreatic ductal adenocarcinoma.

BACKGROUND: In Pancreatic Ductal Adenocarcinoma (PDAC), current prognostic scores are unable to fully capture the biological heterogeneity of the disease. While some approaches investigating the role of multi-omics in PDAC are emerging, the analysis of methylation data is under exploited. MATERIALS AND METHODS: We analyzed CpG sites from two publicly available datasets, the TCGA-PAAD used as discovery set and the CPTAC-PDA as external test set. Single mutations and co-mutation of KRAS and TP53 genes were identified as targets, and differentially methylated CpG sites (DMC) were detected accordingly. We trained and validated Random Forest (RF) models to predict each target. Area Under the Receiver Operating Characteristic curve (AUROC) and Area Under the Precision-Recall curve (AUPRC) were used as performance metrics. Then, we performed consensus clustering from the DMCs to identify novel patients' profiles. Finally, we trained and validated a combination of eXtreme Gradient Boosting (XGB) and tree models to select an epigenomic prognostic determinant. RESULTS: From 598 DMCs extracted, an RF model predicted KRAS and TP53 co-mutation on the external test set with AUROC of 0.77 and AUPRC of 0.87. The consensus clustering allowed us to identify 4 clusters (C1, C2, C3, and C4) of patients. The C4 cluster captured a subgroup of patients with favorable Overall Survival (OS) with respect to others. The XGB model perfectly predicted C4 vs other clusters on the discovery set. In both cohorts, patients were stratified into two risk groups according to methylation levels of cg16854533, individuated as the most important CpG site. CONCLUSION: We analyzed methylation data to develop a classifier for the TP53 and KRAS mutational status. Four prognostic clusters were pointed out and a prognostic model using a CpG site was validated in an independent cohort. Our results evidence that the proposed use of methylation data facilitates risk stratification for PDAC.

Humans↗