PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Multi‑omics identification of a novel signature for serous ovarian carcinoma in the context of 3P medicine and based on twelve programmed cell death patterns: a multi-cohort machine learning study.

BACKGROUND: Predictive, preventive, and personalized medicine (PPPM/3PM) is a strategy aimed at improving the prognosis of cancer, and programmed cell death (PCD) is increasingly recognized as a potential target in cancer therapy and prognosis. However, a PCD-based predictive model for serous ovarian carcinoma (SOC) is lacking. In the present study, we aimed to establish a cell death index (CDI)-based model using PCD-related genes. METHODS: We included 1254 genes from 12 PCD patterns in our analysis. Differentially expressed genes (DEGs) from the Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) were screened. Subsequently, 14 PCD-related genes were included in the PCD-gene-based CDI model. Genomics, single-cell transcriptomes, bulk transcriptomes, spatial transcriptomes, and clinical information from TCGA-OV, GSE26193, GSE63885, and GSE140082 were collected and analyzed to verify the prediction model. RESULTS: The CDI was recognized as an independent prognostic risk factor for patients with SOC. Patients with SOC and a high CDI had lower survival rates and poorer prognoses than those with a low CDI. Specific clinical parameters and the CDI were combined to establish a nomogram that accurately assessed patient survival. We used the PCD-genes model to observe differences between high and low CDI groups. The results showed that patients with SOC and a high CDI showed immunosuppression and hardly benefited from immunotherapy; therefore, trametinib_1372 and BMS-754807 may be potential therapeutic agents for these patients. CONCLUSIONS: The CDI-based model, which was established using 14 PCD-related genes, accurately predicted the tumor microenvironment, immunotherapy response, and drug sensitivity of patients with SOC. Thus this model may help improve the diagnostic and therapeutic efficacy of PPPM.

Humans↗

Development and validation of a machine learning prognostic model based on an epigenomic signature in patients with pancreatic ductal adenocarcinoma.

BACKGROUND: In Pancreatic Ductal Adenocarcinoma (PDAC), current prognostic scores are unable to fully capture the biological heterogeneity of the disease. While some approaches investigating the role of multi-omics in PDAC are emerging, the analysis of methylation data is under exploited. MATERIALS AND METHODS: We analyzed CpG sites from two publicly available datasets, the TCGA-PAAD used as discovery set and the CPTAC-PDA as external test set. Single mutations and co-mutation of KRAS and TP53 genes were identified as targets, and differentially methylated CpG sites (DMC) were detected accordingly. We trained and validated Random Forest (RF) models to predict each target. Area Under the Receiver Operating Characteristic curve (AUROC) and Area Under the Precision-Recall curve (AUPRC) were used as performance metrics. Then, we performed consensus clustering from the DMCs to identify novel patients' profiles. Finally, we trained and validated a combination of eXtreme Gradient Boosting (XGB) and tree models to select an epigenomic prognostic determinant. RESULTS: From 598 DMCs extracted, an RF model predicted KRAS and TP53 co-mutation on the external test set with AUROC of 0.77 and AUPRC of 0.87. The consensus clustering allowed us to identify 4 clusters (C1, C2, C3, and C4) of patients. The C4 cluster captured a subgroup of patients with favorable Overall Survival (OS) with respect to others. The XGB model perfectly predicted C4 vs other clusters on the discovery set. In both cohorts, patients were stratified into two risk groups according to methylation levels of cg16854533, individuated as the most important CpG site. CONCLUSION: We analyzed methylation data to develop a classifier for the TP53 and KRAS mutational status. Four prognostic clusters were pointed out and a prognostic model using a CpG site was validated in an independent cohort. Our results evidence that the proposed use of methylation data facilitates risk stratification for PDAC.

Humans↗

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

Cross-Platform Proteomics and Machine Learning Algorithms Nominate Plasma Biomarkers of Stroke Diagnosis.

BACKGROUND: Blood-based biomarkers for stroke subtyping could improve triage in emergency settings. We used cross-platform proteomics to identify plasma biomarkers differentiating major stroke diagnostic groups. METHODS: We conducted a case-control study using 2 biorepositories. Plasma was collected in the emergency department from adults with suspected stroke before therapeutic intervention. Differentially enriched proteins were identified across acute ischemic stroke, intracerebral hemorrhage, transient ischemic attack, and stroke mimics using SomaScan discovery proteomics (Grady). Differentially enriched proteins were nominated using pairwise and multigroup comparisons and adjusted for clinical covariates. Protein panels were created using least absolute shrinkage and selection operator logistic regression. Internal validation used repeated nested cross-validation (rCV) and targeted mass spectrometry (MS), while external validation used data-independent acquisition &#xa0;mass spectrometry in an independent cohort (Yale). RESULTS: We included 100 subjects (40 with acute ischemic stroke, 20 with intracerebral hemorrhage, 20 with transient ischemic attack, 20 with stroke mimics) in discovery and 80 subjects (20 per group) in external validation cohorts. SomaScan quantified 7307 proteins, of which 61 differentiated stroke subtypes. We identified 7 protein classifiers for acute ischemic stroke (rCV-area under the curve, 0.82 [95% CI, 0.78-0.86]), 6 for intracerebral hemorrhage (rCV-area under the curve, 0.70 [95% CI, 0.64-0.76]), 8 for transient ischemic attack (rCV-area under the curve, 0.78 [95% CI, 0.73-0.84]), and 7 for stroke mimics (rCV-area under the curve, 0.81 [95% CI, 0.77-0.86]). Targeted proteomics internally validated 11 proteins, and data-independent acquisition-mass spectrometry externally validated 32 proteins, including VTN (vitronectin), PLG (plasminogen), and S100A9 as top stroke mimics, transient ischemic attack, and intracerebral hemorrhage classifiers. CONCLUSIONS: This study highlights plasma proteomics as a valuable tool for discovering protein biomarkers of stroke diagnosis. These findings support further validation in larger, multicenter cohorts to facilitate biomarker-guided stroke diagnosis in acute care.

Humans↗

The cerebellum: a neuronal learning machine?

Comparison of two seemingly quite different behaviors yields a surprisingly consistent picture of the role of the cerebellum in motor learning. Behavioral and physiological data about classical conditioning of the eyelid response and motor learning in the vestibulo-ocular reflex suggests that (i) plasticity is distributed between the cerebellar cortex and the deep cerebellar nuclei; (ii) the cerebellar cortex plays a special role in learning the timing of movement; and (iii) the cerebellar cortex guides learning in the deep nuclei, which may allow learning to be transferred from the cortex to the deep nuclei. Because many of the similarities in the data from the two systems typify general features of cerebellar organization, the cerebellar mechanisms of learning in these two systems may represent principles that apply to many motor systems.

Animals↗

Prediction of beta-turns with learning machines.

The support vector machine approach was introduced to predict the beta-turns in proteins. The overall self-consistency rate by the re-substitution test for the training or learning dataset reached 100%. Both the training dataset and independent testing dataset were taken from Chou [J. Pept. Res. 49 (1997) 120]. The success prediction rates by the jackknife test for the beta-turn subset of 455 tetrapeptides and non-beta-turn subset of 3807 tetrapeptides in the training dataset were 58.1 and 98.4%, respectively. The success rates with the independent dataset test for the beta-turn subset of 110 tetrapeptides and non-beta-turn subset of 30,231 tetrapeptides were 69.1 and 97.3%, respectively. The results obtained from this study support the conclusion that the residue-coupled effect along a tetrapeptide is important for the formation of a beta-turn.

Artificial Intelligence↗

Monitoring of complex industrial bioprocesses for metabolite concentrations using modern spectroscopies and machine learning: application to gibberellic acid production.

Two rapid vibrational spectroscopic approaches (diffuse reflectance-absorbance Fourier transform infrared [FT-IR] and dispersive Raman spectroscopy), and one mass spectrometric method based on in vacuo Curie-point pyrolysis (PyMS), were investigated in this study. A diverse range of unprocessed, industrial fed-batch fermentation broths containing the fungus Gibberella fujikuroi producing the natural product gibberellic acid, were analyzed directly without a priori chromatographic separation. Partial least squares regression (PLSR) and artificial neural networks (ANNs) were applied to all of the information-rich spectra obtained by each of the methods to obtain quantitative information on the gibberellic acid titer. These estimates were of good precision, and the typical root-mean-square error for predictions of concentrations in an independent test set was <10% over a very wide titer range from 0 to 4925 ppm. However, although PLSR and ANNs are very powerful techniques they are often described as "black box" methods because the information they use to construct the calibration model is largely inaccessible. Therefore, a variety of novel evolutionary computation-based methods, including genetic algorithms and genetic programming, were used to produce models that allowed the determination of those input variables that contributed most to the models formed, and to observe that these models were predominantly based on the concentration of gibberellic acid itself. This is the first time that these three modern analytical spectroscopies, in combination with advanced chemometric data analysis, have been compared for their ability to analyze a real commercial bioprocess. The results demonstrate unequivocally that all methods provide very rapid and accurate estimates of the progress of industrial fermentations, and indicate that, of the three methods studied, Raman spectroscopy is the ideal bioprocess monitoring method because it can be adapted for on-line analysis.

Algorithms↗

Machine learning approaches to lung cancer prediction from mass spectra.

We addressed the problem of discriminating between 24 diseased and 17 healthy specimens on the basis of protein mass spectra. To prepare the data, we performed mass to charge ratio (m/z) normalization, baseline elimination, and conversion of absolute peak height measures to height ratios. After preprocessing, the major difficulty encountered was the extremely large number of variables (1676 m/z values) versus the number of examples (41). Dimensionality reduction was treated as an integral part of the classification process; variable selection was coupled with model construction in a single ten-fold cross-validation loop. We explored different experimental setups involving two peak height representations, two variable selection methods, and six induction algorithms, all on both the original 1676-mass data set and on a prescreened 124-mass data set. Highest predictive accuracies (1-2 off-sample misclassifications) were achieved by a multilayer perceptron and Naïve Bayes, with the latter displaying more consistent performance (hence greater reliability) over varying experimental conditions. We attempted to identify the most discriminant peaks (proteins) on the basis of scores assigned by the two variable selection methods and by neural network based sensitivity analysis. These three scoring schemes consistently ranked four peaks as the most relevant discriminators: 11683, 1403, 17350 and 66107.

Algorithms↗

A shape-based machine learning tool for drug design.

Building predictive models for iterative drug design in the absence of a known target protein structure is an important challenge. We present a novel technique, Compass, that removes a major obstacle to accurate prediction by automatically selecting conformations and alignments of molecules without the benefit of a characterized active site. The technique combines explicit representation of molecular shape with neural network learning methods to produce highly predictive models, even across chemically distinct classes of molecules. We apply the method to predicting human perception of musk odor and show how the resulting models can provide graphical guidance for chemical modifications.

Algorithms↗

On the optimization of classes for the assignment of unidentified reading frames in functional genomics programmes: the need for machine learning.

At present, the assignment of function to novel genes uncovered by the systematic genome-sequencing programmes is a problem. Many studies anticipate that this can be achieved by analysing patterns of gene expression via the transcriptome, proteome and metabolome. Thus, functional genomics is, in part, an exercise in pattern classification. Because many genes have known functional classes, the problem of predicting their functional class is a supervised learning problem. However, most pattern classification methods that have been applied to the problem have been unsupervised clustering methods. Consequently, the best classification tools have not always been used. Furthermore, the present functional classes are suboptimal and new unsupervised clustering methods are needed to improve them. Better-structured functional classes will facilitate the prediction of biochemically testable functions.

Animals↗

Relating clinical and neurophysiological assessment of spasticity by machine learning.

Spasticity following spinal cord injury (SCI) is most often assessed clinically using a five-point Ashworth score (AS). A more objective assessment of altered motor control may be achieved by using a comprehensive protocol based on a surface electromyographic (sEMG) activity recorded from thigh and leg muscles. However, the relationship between the clinical and neurophysiological assessments is still unknown. In this paper we employ three different classification methods to investigate this relationship. The experimental results indicate that, if the appropriate set of sEMG features is used, the neurophysiological assessment is related to clinical findings and can be used to predict the AS. A comprehensive sEMG assessment may be proven useful as an objective method of evaluating the effectiveness of various interventions and for follow-up of SCI patients.

Artificial Intelligence↗

Modeling of human cytochrome p450-mediated drug metabolism using unsupervised machine learning approach.

We developed a computational algorithm for evaluating the possibility of cytochrome P450-mediated metabolic transformations that xenobiotics molecules undergo in the human body. First, we compiled a database of known human cytochrome P-450 substrates, products, and nonsubstrates for 38 enzyme-specific groups (total of 2200 compounds). Second, we determined the cytochrome-mediated metabolic reactions most typical for each group and examined the substrates and products of these reactions. To assess the probability of P450 transformations of novel compounds, we built a nonlinear quantitative structure-metabolism relationships (QSMR) model based on Kohonen self-organizing maps (SOM). This neural network QSMR model incorporated a predefined set of physicochemical descriptors encoding the key molecular properties that define the metabolic fate of individual molecules. Isozyme-specific groups of substrate molecules were visualized, thus facilitating prediction of tissue-specific metabolism. The developed algorithm can be used in early stages of drug discovery as an efficient tool for the assessment of human metabolism and toxicity of novel compounds in designing discovery libraries and in lead optimization.

Algorithms↗

Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.

Diffuse large B-cell lymphoma (DLBCL), the most common lymphoid malignancy in adults, is curable in less than 50% of patients. Prognostic models based on pre-treatment characteristics, such as the International Prognostic Index (IPI), are currently used to predict outcome in DLBCL. However, clinical outcome models identify neither the molecular basis of clinical heterogeneity, nor specific therapeutic targets. We analyzed the expression of 6,817 genes in diagnostic tumor specimens from DLBCL patients who received cyclophosphamide, adriamycin, vincristine and prednisone (CHOP)-based chemotherapy, and applied a supervised learning prediction method to identify cured versus fatal or refractory disease. The algorithm classified two categories of patients with very different five-year overall survival rates (70% versus 12%). The model also effectively delineated patients within specific IPI risk categories who were likely to be cured or to die of their disease. Genes implicated in DLBCL outcome included some that regulate responses to B-cell-receptor signaling, critical serine/threonine phosphorylation pathways and apoptosis. Our data indicate that supervised learning classification techniques can predict outcome in DLBCL and identify rational targets for intervention.

Antineoplastic Combined Chemotherapy Protocols↗

Structure-activity relationships derived by machine learning: the use of atoms and their bond connectivities to predict mutagenicity by inductive logic programming.

We present a general approach to forming structure-activity relationships (SARs). This approach is based on representing chemical structure by atoms and their bond connectivities in combination with the inductive logic programming (ILP) algorithm PROGOL. Existing SAR methods describe chemical structure by using attributes which are general properties of an object. It is not possible to map chemical structure directly to attribute-based descriptions, as such descriptions have no internal organization. A more natural and general way to describe chemical structure is to use a relational description, where the internal construction of the description maps that of the object described. Our atom and bond connectivities representation is a relational description. ILP algorithms can form SARs with relational descriptions. We have tested the relational approach by investigating the SARs of 230 aromatic and heteroaromatic nitro compounds. These compounds had been split previously into two subsets, 188 compounds that were amenable to regression and 42 that were not. For the 188 compounds, a SAR was found that was as accurate as the best statistical or neural network-generated SARs. The PROGOL SAR has the advantages that it did not need the use of any indicator variables handcrafted by an expert, and the generated rules were easily comprehensible. For the 42 compounds, PROGOL formed a SAR that was significantly (P < 0.025) more accurate than linear regression, quadratic regression, and back-propagation. This SAR is based on an automatically generated structural alert for mutagenicity.

Algorithms↗

Machine learning methods applied on dental fear and behavior management problems in children.

The etiologies of dental fear and dental behavior management problems in children were investigated in a database of information on 2,257 Swedish children 4-6 and 9-11 years old. The analyses were performed using computerized inductive techniques within the field of artificial intelligence. The database held information regarding dental fear levels and behavior management problems, which were defined as outcomes, i.e. dependent variables. The attributes, i.e. independent variables, included data on dental health and dental treatments, information about parental dental fear, general anxiety, socioeconomic variables, etc. The data contained both numerical and discrete variables. The analyses were performed using an inductive analysis program (XpertRule Analyser, Attar Software Ltd, Lancashire, UK) that presents the results in a hierarchic diagram called a knowledge tree. The importance of the different attributes is represented by their position in this diagram. The results show that inductive methods are well suited for analyzing multifactorial and complex relationships in large data sets, and are thus a useful complement to multivariate statistical techniques. The knowledge trees for the two outcomes, dental fear and behavior management problems, were very different from each other, suggesting that the two phenomena are not equivalent. Dental fear was found to be more related to non-dental variables, whereas dental behavior management problems seemed connected to dental variables.

Artificial Intelligence↗

Application of metabolomics to plant genotype discrimination using statistics and machine learning.

MOTIVATION: Metabolomics is a post genomic technology which seeks to provide a comprehensive profile of all the metabolites present in a biological sample. This complements the mRNA profiles provided by microarrays, and the protein profiles provided by proteomics. To test the power of metabolome analysis we selected the problem of discrimating between related genotypes of Arabidopsis. Specifically, the problem tackled was to discrimate between two background genotypes (Col0 and C24) and, more significantly, the offspring produced by the crossbreeding of these two lines, the progeny (whose genotypes would differ only in their maternally inherited mitichondia and chloroplasts). OVERVIEW: A gas chromotography--mass spectrometry (GCMS) profiling protocol was used to identify 433 metabolites in the samples. The metabolomic profiles were compared using descriptive statistics which indicated that key primary metabolites vary more than other metabolites. We then applied neural networks to discriminate between the genotypes. This showed clearly that the two background lines can be discrimated between each other and their progeny, and indicated that the two progeny lines can also be discriminated. We applied Euclidean hierarchical and Principal Component Analysis (PCA) to help understand the basis of genotype discrimination. PCA indicated that malic acid and citrate are the two most important metabolites for discriminating between the background lines, and glucose and fructose are two most important metabolites for discriminating between the crosses. These results are consistant with genotype differences in mitochondia and chloroplasts.

Algorithms↗