PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Feature selection and the class imbalance problem in predicting protein function from sequence.

When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence features. However, there are two issues to consider before successful functional prediction can take place: identifying discriminatory features, and overcoming the challenge of a large imbalance in the training data. We show that by applying feature subset selection followed by undersampling of the majority class, significantly better support vector machine (SVM) classifiers are generated compared with standard machine learning approaches. As well as revealing that the features selected could have the potential to advance our understanding of the relationship between sequence and function, we also show that undersampling to produce fully balanced data significantly improves performance. The best discriminating ability is achieved using SVMs together with feature selection and full undersampling; this approach strongly outperforms other competitive learning algorithms. We conclude that this combined approach can generate powerful machine learning classifiers for predicting protein function directly from sequence.

Algorithms↗

Combining artificial neural networks and transrectal ultrasound in the diagnosis of prostate cancer.

Arguably the most important step in the prognosis of prostate cancer is early diagnosis. More than 1 million transrectal ultrasound (TRUS)-guided prostate needle biopsies are performed annually in the United States, resulting in the detection of 200,000 new cases per year. Unfortunately, the urologist's ability to diagnose prostate cancer has not kept pace with therapeutic advances; currently, many men are facing the need for prostate biopsy with the likelihood that the result will be inconclusive. This paper will focus on the tools available to assist the clinician in predicting the outcome of the prostate needle biopsy. We will examine the use of "machine learning" models (artificial intelligence), in the form of artificial neural networks (ANNs), to predict prostate biopsy outcomes using prebiopsy variables. Currently, six validated predictive models are available. Of these, five are machine learning models, and one is based on logistic regression. The role of ANNs in providing valuable predictive models to be used in conjunction with TRUS appears promising. In the few studies that have compared machine learning to traditional statistical methods, ANN and logistic regression appear to function equivalently when predicting biopsy outcome. With the introduction of more complex prebiopsy variables, ANNs are in a commanding position for use in predictive models. Easy and immediate physician access to these models will be imperative if their full potential is to be realized.

Biopsy↗

A wavelet-based data pre-processing analysis approach in mass spectrometry.

Recently, mass spectrometry analysis has a become an effective and rapid approach in detecting early-stage cancer. To identify proteomic patterns in serum to discriminate cancer patients from normal individuals, machine-learning methods, such as feature selection and classification, have already been involved in the analysis of mass spectrometry (MS) data with some success. However, the performance of existing machine learning methods for MS data analysis still needs improving. The study in this paper proposes a wavelet-based pre-processing approach to MS data analysis. The approach applies wavelet-based transforms to MS data with the aim of de-noising the data that are potentially contaminated in acquisition. The effects of the selection of wavelet function and decomposition level on the de-noising performance have also been investigated in this study. Our comparative experimental results demonstrate that the proposed de-noising pre-processing approach has potentials to remove possible noise embedded in MS data, which can lead to improved performance for existing machine learning methods in cancer detection.

Algorithms↗

Learning the relationship between patient geometry and beam intensity in breast intensity-modulated radiotherapy.

Intensity modulated radiotherapy (IMRT) has become an effective tool for cancer treatment with radiation. However, even expert radiation planners still need to spend a substantial amount of time adjusting IMRT optimization parameters in order to get a clinically acceptable plan. We demonstrate that the relationship between patient geometry and radiation intensity distributions can be automatically inferred using a variety of machine learning techniques in the case of two-field breast IMRT. Our experiments show that given a small number of human-expert-generated clinically acceptable plans, the machine learning predictions produce equally acceptable plans in a matter of seconds. The machine learning approach has the potential for greater benefits in sites where the IMRT planning process is more challenging or tedious.

Artificial Intelligence↗

Application of artificial intelligence in audiology.

In this paper, machine learning methods based on artificial intelligence theory are applied to the computer-aided decision making of some otoneurological diseases, for example Ménière's disease. Three methods explored are decision trees, genetic algorithms and neural networks. By using such a machine learning method, the decision-making program is trained with a representative training set of cases and tested with another set. The machine learning methods are useful also for our otoneurological expert system, One, which is based on a pattern recognition approach. The methods are able to differentiate most of the cases tested between the six diseases included, provided that a sufficiently large training set is available.

Algorithms↗

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

Transductive reliability estimation for medical diagnosis.

In the past decades, machine learning (ML) tools have been successfully used in several medical diagnostic problems. While they often significantly outperform expert physicians (in terms of diagnostic accuracy, sensitivity, and specificity), they are mostly not being used in practice. One reason for this is that it is difficult to obtain an unbiased estimation of diagnose's reliability. We discuss how reliability of diagnoses is assessed in medical decision-making and propose a general framework for reliability estimation in machine learning, based on transductive inference. We compare our approach with a usual (machine learning) probabilistic approach as well as with classical stepwise diagnostic process where reliability of diagnose is presented as its post-test probability. The proposed transductive approach is evaluated on several medical datasets from the University of California (UCI) repository as well as on a practical problem of clinical diagnosis of the coronary artery disease (CAD). In all cases, significant improvements over existing techniques are achieved.

Artificial Intelligence↗

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms↗

GANN: genetic algorithm neural networks for the detection of conserved combinations of features in DNA.

BACKGROUND: The multitude of motif detection algorithms developed to date have largely focused on the detection of patterns in primary sequence. Since sequence-dependent DNA structure and flexibility may also play a role in protein-DNA interactions, the simultaneous exploration of sequence- and structure-based hypotheses about the composition of binding sites and the ordering of features in a regulatory region should be considered as well. The consideration of structural features requires the development of new detection tools that can deal with data types other than primary sequence. RESULTS: GANN (available at http://bioinformatics.org.au/gann) is a machine learning tool for the detection of conserved features in DNA. The software suite contains programs to extract different regions of genomic DNA from flat files and convert these sequences to indices that reflect sequence and structural composition or the presence of specific protein binding sites. The machine learning component allows the classification of different types of sequences based on subsamples of these indices, and can identify the best combinations of indices and machine learning architecture for sequence discrimination. Another key feature of GANN is the replicated splitting of data into training and test sets, and the implementation of negative controls. In validation experiments, GANN successfully merged important sequence and structural features to yield good predictive models for synthetic and real regulatory regions. CONCLUSION: GANN is a flexible tool that can search through large sets of sequence and structural feature combinations to identify those that best characterize a set of sequences.

Algorithms↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article↗

A new hybrid method based on fuzzy-artificial immune system and k-nn algorithm for breast cancer diagnosis.

The use of machine learning tools in medical diagnosis is increasing gradually. This is mainly because the effectiveness of classification and recognition systems has improved in a great deal to help medical experts in diagnosing diseases. Such a disease is breast cancer, which is a very common type of cancer among woman. As the incidence of this disease has increased significantly in the recent years, machine learning applications to this problem have also took a great attention as well as medical consideration. This study aims at diagnosing breast cancer with a new hybrid machine learning method. By hybridizing a fuzzy-artificial immune system with k-nearest neighbour algorithm, a method was obtained to solve this diagnosis problem via classifying Wisconsin Breast Cancer Dataset (WBCD). This data set is a very commonly used data set in the literature relating the use of classification systems for breast cancer diagnosis and it was used in this study to compare the classification performance of our proposed method with regard to other studies. We obtained a classification accuracy of 99.14%, which is the highest one reached so far. The classification accuracy was obtained via 10-fold cross validation. This result is for WBCD but it states that this method can be used confidently for other breast cancer diagnosis problems, too.

Algorithms↗

Machine learning-driven spleen imaging and genomics uncover a splenic connection to coronary artery disease.

Despite advances in managing traditional risk factors, coronary artery disease (CAD) remains the leading cause of mortality. Circulating hematopoietic cells influence risk for CAD separately from traditional risk factors, but the role of a key regulating organ, the spleen, is unknown. The understudied spleen is a representation of the hematopoietic system optimally suited for unbiased radiologic investigations toward mechanistic insights. Here, we leveraged deep learning to extract 107 splenic radiomic features from abdominal magnetic resonance imaging (MRI) scans of 42,059 UK Biobank participants and of 2745 Mass General Brigham Biobank (MGBB) participants. Of these, 10 features from UK Biobank were associated with CAD. Genome-wide association analysis of CAD-associated features identified 219 loci, including 9p21. Variants at 9p21, the strongest yet mechanistically elusive CAD locus, were associated with splenic features such as run-length nonuniformity, reflecting heterogeneity of continuous texture regions. Research MRI findings were consistent internally, but external clinical validation highlighted challenges in translating analyses of abdominal MRI scans to routine clinical practice because of variability in imaging protocols and greater clinical heterogeneity among patients. Our study, combining deep learning with genomics, presents a framework to uncover potential splenic involvement in CAD and emphasizes translational gaps between research and clinical radiomics.

Humans↗

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning↗

Comparing expert systems for identifying chest x-ray reports that support pneumonia.

We compare the performance of four computerized methods in identifying chest x-ray reports that support acute bacterial pneumonia. Two of the computerized techniques are constructed from expert knowledge, and two learn rules and structure from data. The two machine learning systems perform as well as the expert constructed systems. All of the computerized techniques perform better than a baseline keyword search and a lay person, and perform as well as a physician. We conclude that machine learning can be used to identify chest x-ray reports that support pneumonia.

Algorithms↗

Why neural networks should not be used for HIV-1 protease cleavage site prediction.

UNLABELLED: Several papers have been published where nonlinear machine learning algorithms, e.g. artificial neural networks, support vector machines and decision trees, have been used to model the specificity of the HIV-1 protease and extract specificity rules. We show that the dataset used in these studies is linearly separable and that it is a misuse of nonlinear classifiers to apply them to this problem. The best solution on this dataset is achieved using a linear classifier like the simple perceptron or the linear support vector machine, and it is straightforward to extract rules from these linear models. We identify key residues in peptides that are efficiently cleaved by the HIV-1 protease and list the most prominent rules, relating them to experimental results for the HIV-1 protease. MOTIVATION: Understanding HIV-1 protease specificity is important when designing HIV inhibitors and several different machine learning algorithms have been applied to the problem. However, little progress has been made in understanding the specificity because nonlinear and overly complex models have been used. RESULTS: We show that the problem is much easier than what has previously been reported and that linear classifiers like the simple perceptron or linear support vector machines are at least as good predictors as nonlinear algorithms. We also show how sets of specificity rules can be generated from the resulting linear classifiers. AVAILABILITY: The datasets used are available at http://www.hh.se/staff/bioinf/

Algorithms↗

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan&#x2013;Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans↗

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans↗