PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Predicting food taste with bound-driven optimization.

The prediction of sensory attributes from ingredient-level formulations is an emerging challenge at the intersection of food science and artificial intelligence. We address the fundamental question of whether the taste of a food can be predicted from its ingredients by treating recipes as composite materials. We apply Hashin-Shtrikman (HS) and Reuss-Voigt (RV) bounds, techniques originally developed for elastic moduli, as a null-hypothesis additive baseline for five taste dimensions (sweetness, sourness, bitterness, umami, saltiness) on a curated dataset of 70 recipes decomposed into 115 distinct ingredients scored against a library of 209 ingredient-level taste references with trained-panel ground truth. This baseline systematically under-predicts perceived taste: 77% of actual taste values exceeded the HS upper bound, with the exceedance rate ranging from 26% (bitterness) to 97% (saltiness). We traced this gap to specific processing chemistry (Maillard reactions, caramelization, evaporative concentration, protein hydrolysis, and nucleotide synergy) and introduced a hybrid model that augments the HS baseline with eight chemistry-proxy features encoding these mechanisms. Our results show that our interpretable hybrid model eliminates the systematic bias and reduces mean absolute error by 27%-62% for sweetness, sourness, umami, and saltiness while using only 10 interpretable features, achieving performance comparable to a black-box Lasso regression on 115 per-ingredient features. We further demonstrate constrained inverse design via Differential Evolution, recovering ingredient formulations that match target taste profiles subject to compositional bounds. Our work demonstrates how key chemical processes during food preparation can inform and augment physics-based and machine learning models, providing a quantitative fingerprint of processing chemistry's contribution to taste perception and paving the way for model-driven food formulation with targeted sensory characteristics.

Composite material bounds↗

Predicting the efficacy of short oligonucleotides in antisense and RNAi experiments with boosted genetic programming.

MOTIVATION: Both small interfering RNAs (siRNAs) and antisense oligonucleotides can selectively block gene expression. Although the two methods rely on different cellular mechanisms, these methods share the common property that not all oligonucleotides (oligos) are equally effective. That is, if mRNA target sites are picked at random, many of the antisense or siRNA oligos will not be effective. Algorithms that can reliably predict the efficacy of candidate oligos can greatly reduce the cost of knockdown experiments, but previous attempts to predict the efficacy of antisense oligos have had limited success. Machine learning has not previously been used to predict siRNA efficacy. RESULTS: We develop a genetic programming based prediction system that shows promising results on both antisense and siRNA efficacy prediction. We train and evaluate our system on a previously published database of antisense efficacies and our own database of siRNA efficacies collected from the literature. The best models gave an overall correlation between predicted and observed efficacy of 0.46 on both antisense and siRNA data. As a comparison, the best correlations of support vector machine classifiers trained on the same data were 0.40 and 0.30, respectively.

Algorithms↗

Prediction of CTL epitopes using QM, SVM and ANN techniques.

Cytotoxic T lymphocyte (CTL) epitopes are potential candidates for subunit vaccine design for various diseases. Most of the existing T cell epitope prediction methods are indirect methods that predict MHC class I binders instead of CTL epitopes. In this study, a systematic attempt has been made to develop a direct method for predicting CTL epitopes from an antigenic sequence. This method is based on quantitative matrix (QM) and machine learning techniques such as Support Vector Machine (SVM) and Artificial Neural Network (ANN). This method has been trained and tested on non-redundant dataset of T cell epitopes and non-epitopes that includes 1137 experimentally proven MHC class I restricted T cell epitopes. The accuracy of QM-, ANN- and SVM-based methods was 70.0, 72.2 and 75.2%, respectively. The performance of these methods has been evaluated through Leave One Out Cross-Validation (LOOCV) at a cutoff score where sensitivity and specificity was nearly equal. Finally, both machine-learning methods were used for consensus and combined prediction of CTL epitopes. The performances of these methods were evaluated on blind dataset where machine learning-based methods perform better than QM-based method. We also demonstrated through subgroup analysis that our methods can discriminate between T-cell epitopes and MHC binders (non-epitopes). In brief this method allows prediction of CTL epitopes using QM, SVM, ANN approaches. The method also facilitates prediction of MHC restriction in predicted T cell epitopes.

Algorithms↗

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning↗

Facial attractiveness: beauty and the machine.

This work presents a novel study of the notion of facial attractiveness in a machine learning context. To this end, we collected human beauty ratings for data sets of facial images and used various techniques for learning the attractiveness of a face. The trained predictor achieves a significant correlation of 0.65 with the average human ratings. The results clearly show that facial beauty is a universal concept that a machine can learn. Analysis of the accuracy of the beauty prediction machine as a function of the size of the training data indicates that a machine producing human-like attractiveness rating could be obtained given a moderately larger data set.

Algorithms↗

Decision tree-based formation of consensus protein secondary structure prediction.

MOTIVATION: Prediction of protein secondary structure provides information that is useful for other prediction methods like fold recognition and ab initio 3D prediction. A consensus prediction constructed from the output of several methods should yield more reliable results than each of the individual methods. METHOD: We present an approach that reveals subtle but systematic differences in the output of different secondary structure prediction methods allowing the derivation of coherent consensus predictions. The method uses a machine learning technique that builds decision trees from existing data. RESULTS: The first results of our analysis show that consensus prediction of protein secondary structure may be improved both quantitatively and qualitatively.

Algorithms↗

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1 = 0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n = 305) demonstrated generalizability (AUROC = 0.8105; F1 = 0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC) = 0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article↗

Searching for functional sites in protein structures.

An ability to assign protein function from protein structure is important for structural genomics consortia. The complex relationship between protein fold and function highlights the necessity of looking beyond the global fold of a protein to specific functional sites. Many computational methods have been developed that address this issue. These include evolutionary trace methods, methods that involve the calculation and assessment of maximal superpositions, methods based on graph theory, and methods that apply machine learning techniques. Such function prediction techniques have been applied to the identification of enzyme catalytic triads and DNA-binding motifs.

Binding Sites↗

Integrative multi-omics profiling deciphers tumor microenvironment heterogeneity and immunotherapy vulnerabilities in lung neuroendocrine carcinomas.

INTRODUCTION: Lung neuroendocrine carcinomas (Lu-NECs) are rare, highly aggressive lung tumors with poor prognosis and limited therapeutic options. Understanding the tumor immune microenvironment (TIME) is crucial towards personalized therapeutic strategies. OBJECTIVES: This study aims to systematically characterize the heterogeneity and complexity of the TIME in Lu-NECs by integrating proteomic, transcriptomic, and genomic data. METHODS: We performed comprehensive immune-proteomic profiling of 76 Lu-NECs across diverse histopathological subtypes to elucidate intra-tumoral TIME heterogeneity at the proteomic level. Validation was conducted in multiple independent cohorts, including 112 Lu-NECs using immunohistochemistry, 147 Lu-NECs, and 17 small cell lung carcinoma samples using transcriptomics. We integrated proteomic, transcriptomic, genomic, and clinical data to assess molecular, immunological, and clinical features, as well as therapeutic vulnerabilities across different immune subtypes. RESULTS: We delineated the immuno-proteomic landscape of Lu-NECs and identified two major immuno-proteomic clusters with distinct immunological, molecular, and clinical characteristics. IPC1 was characterized by high immune cell infiltration, while IPC2 exhibited sparse immune cell presence. Genomic analysis revealed distinct mutational patterns, with IPC1 showing a higher incidence of APOBEC-associated mutation signatures and IPC2 being enriched for mutations associated with defective DNA mismatch repair and tobacco-related mutagens. Functional analyses indicated that IPC1 was related to immune and oncogenic signaling activity, whereas IPC2 was associated with cancer stemness and proliferation-related features. Furthermore, IPC1 and IPC2 demonstrated histological subtype-specific clinical benefits from postoperative chemotherapy. Finally, we developed a machine learning model (iPROM) to predict Lu-NECs immune classification and improve risk stratification, which was validated across multiple independent cohorts. CONCLUSIONS: This study advances the understanding of the tumor immune microenvironment in Lu-NECs through multi-omics characterization and highlights potential personalized therapeutic vulnerabilities tailored to the specific immune landscapes of Lu-NECs.

Humans↗

Machine learning in sedimentation modelling.

The paper presents machine learning (ML) models that predict sedimentation in the harbour basin of the Port of Rotterdam. The important factors affecting the sedimentation process such as waves, wind, tides, surge, river discharge, etc. are studied, the corresponding time series data is analysed, missing values are estimated and the most important variables behind the process are chosen as the inputs. Two ML methods are used: MLP ANN and M5 model tree. The latter is a collection of piece-wise linear regression models, each being an expert for a particular region of the input space. The models are trained on the data collected during 1992-1998 and tested by the data of 1999-2000. The predictive accuracy of the models is found to be adequate for the potential use in the operational decision making.

Algorithms↗

Saturating the eQTL map in Drosophila: Genome-wide patterns of cis and trans regulation of transcriptional variation in outbred populations.

Most genetic polymorphisms associated with complex traits are found in non-coding regions of the genome. Characterizing their effect presents a formidable challenge, and expression quantitative trait locus (eQTLs) mapping has been a key approach to do so. As comprehensive eQTL maps are available only for a few species, here we developed the Drosophila outbred synthetic population (Dros-OSP) and used it to characterize the landscape of transcriptional regulation in Drosophila melanogaster. We collected head and body transcriptomes and genomes from 1,286 outbred flies and mapped local and distant eQTLs for 98% of the genes. We characterized the network organization of the transcriptome across tissues and described the properties of local and distal eQTLs in terms of genetic diversity, heritability, connectivity, and pleiotropy. These results provide new insights into the genetic basis of transcriptional regulation in the fruit fly and offer a new mapping resource that will expand the possibilities currently available for the Drosophila community.

Animals↗

The problem of bias in training data in regression problems in medical decision support.

This paper describes a bias problem encountered in a machine learning approach to outcome prediction in anticoagulant drug therapy. The outcome to be predicted is a measure of the clotting time for the patient; this measure is continuous and so the prediction task is a regression problem. Artificial neural networks (ANNs) are a powerful mechanism for learning to predict such outcomes from training data. However, experiments have shown that an ANN is biased towards values more commonly occurring in the training data and is thus, less likely to be correct in predicting extreme values. This issue of bias in training data in regression problems is similar to the associated problem with minority classes in classification. However, this bias issue in classification is well documented and is an on-going area of research. In this paper, we consider stratified sampling and boosting as solutions to this bias problem and evaluate them on this outcome prediction problem and on two other datasets. Both approaches produce some improvements with boosting showing the most promise.

Bias↗

Structure-activity relationships derived by machine learning: the use of atoms and their bond connectivities to predict mutagenicity by inductive logic programming.

We present a general approach to forming structure-activity relationships (SARs). This approach is based on representing chemical structure by atoms and their bond connectivities in combination with the inductive logic programming (ILP) algorithm PROGOL. Existing SAR methods describe chemical structure by using attributes which are general properties of an object. It is not possible to map chemical structure directly to attribute-based descriptions, as such descriptions have no internal organization. A more natural and general way to describe chemical structure is to use a relational description, where the internal construction of the description maps that of the object described. Our atom and bond connectivities representation is a relational description. ILP algorithms can form SARs with relational descriptions. We have tested the relational approach by investigating the SARs of 230 aromatic and heteroaromatic nitro compounds. These compounds had been split previously into two subsets, 188 compounds that were amenable to regression and 42 that were not. For the 188 compounds, a SAR was found that was as accurate as the best statistical or neural network-generated SARs. The PROGOL SAR has the advantages that it did not need the use of any indicator variables handcrafted by an expert, and the generated rules were easily comprehensible. For the 42 compounds, PROGOL formed a SAR that was significantly (P < 0.025) more accurate than linear regression, quadratic regression, and back-propagation. This SAR is based on an automatically generated structural alert for mutagenicity.

Algorithms↗

Structural models of osteogenesis imperfecta-associated variants in the COL1A1 gene.

Osteogenesis imperfecta (OI) is a genetic disease in which the most common mutations result in substitutions for glycine residues in the triple helical domain of the chains of type I collagen. Currently there is no way to use sequence information to predict the clinical OI phenotype. However, structural models coupled with biophysical and machine learning methods may be able to predict sequences that, when mutated, would be associated with more severe forms of OI. To build appropriate structural models, we have applied a high throughput molecular dynamic approach. Homotrimeric peptides covering 57 positions in which mutations are associated with OI were simulated both with and without mutations. Our models revealed structural differences that occur with different substituting amino acids. When mutations were introduced, we observed a decrease in helix stability, as caused by fewer main chain backbone hydrogen bonds, and an increase in main chain root mean square deviation and specifically bound water molecules.

Collagen Type I↗

Modular DAG-RNN architectures for assembling coarse protein structures.

We develop and test machine learning methods for the prediction of coarse 3D protein structures, where a protein is represented by a set of rigid rods associated with its secondary structure elements (alpha-helices and beta-strands). First, we employ cascades of recursive neural networks derived from graphical models to predict the relative placements of segments. These are represented as discretized distance and angle maps, and the discretization levels are statistically inferred from a large and curated dataset. Coarse 3D folds of proteins are then assembled starting from topological information predicted in the first stage. Reconstruction is carried out by minimizing a cost function taking the form of a purely geometrical potential. We show that the proposed architecture outperforms simpler alternatives and can accurately predict binary and multiclass coarse maps. The reconstruction procedure proves to be fast and often leads to topologically correct coarse structures that could be exploited as a starting point for various protein modeling strategies. The fully integrated rod-shaped protein builder (predictor of contact maps + reconstruction algorithm) can be accessed at http://distill.ucd.ie/.

Algorithms↗

Ecological Filtering by Tuber Compartments Shapes Stable Core Microbiomes That Underpin Potato Plant Growth Across Environments.

Harnessing plant microbiomes for sustainable agriculture requires understanding not only whether they can boost crop performance, but also how ecological processes govern their assembly, stability, and functional contributions across environments. While we previously showed that seed tuber microbiomes can predict potato vigour using machine learning, it remained unclear how ecological processes shape tuber microbiome stability and functionality across host genotypes, tuber compartments, soil types, and years. Here, we analyzed the national-scale dataset of 240 field-collected potato seedlots, spanning six genotypes, two soil types, and two growing years, with a focus on the spatially distinct heel and eye compartments of the potato tuber. By profiling over 1200 bacterial and fungal communities and linking microbiome composition to plant performance, we show that plant genotype and tuber compartment are the strongest determinants of microbial diversity and composition. Compartment-specific enrichment of functional traits revealed spatial partitioning of microbial functions, with organic compound conversion and nitrogen cycling dominant in the heel, and energy metabolism enriched in the eye. Applying a macroecological abundance-occupancy framework, we identified a stable core microbiome of bacterial and fungal taxa that persisted across all environments and years. These core members were more strongly associated with plant growth-related traits than non-core taxa, and core taxa in different tuber compartments showed distinct correlations with taxa of potential pathogenic relevance. Together, our findings demonstrate that tuber compartments act as ecological filters that structure persistent, functionally specialised microbiomes linked to plant growth-related traits across environments. By providing an ecological and functional framework for compartment-resolved, stable core microbiomes, this study advances mechanistic understanding of plant-microbe interactions and identifies stable microbial partners as promising targets for improving potato resilience and productivity.

Journal Article↗

Allostatic load is associated with symptoms in chronic fatigue syndrome patients.

OBJECTIVES: To further explore the relationship between chronic fatigue syndrome (CFS) and allostatic load (AL), we conducted a computational analysis involving 43 patients with CFS and 60 nonfatigued, healthy controls (NF) enrolled in a population-based case-control study in Wichita (KS, USA). We used traditional biostatistical methods to measure the association of high AL to standardized measures of physical and mental functioning, disability, fatigue and general symptom severity. We also used nonlinear regression technology embedded in machine learning algorithms to learn equations predicting various CFS symptoms based on the individual components of the allostatic load index (ALI). METHODS: An ALI was computed for all study participants using available laboratory and clinical data on metabolic, cardiovascular and hypothalamic-pituitary-adrenal (HPA) axis factors. Physical and mental functioning/impairment was measured using the Medical Outcomes Study 36-item Short Form Health Survey (SF-36); current fatigue was measured using the 20-item multidimensional fatigue inventory (MFI); frequency and intensity of symptoms was measured using the 19-item symptom inventory (SI). Genetic programming, a nonlinear regression technique, was used to learn an ensemble of different predictive equations rather just than a single one. Statistical analysis was based on the calculation of the percentage of equations in the ensemble that utilized each input variable, producing a measure of the 'utility' of the variable for the predictive problem at hand. Traditional biostatistics methods include the median and Wilcoxon tests for comparing the median levels of subscale scores obtained on the SF-36, the MFI and the SI summary score. RESULTS: Among CFS patients, but not controls, a high level of AL was significantly associated with lower median values (indicating worse health) of bodily pain, physical functioning and general symptom frequency/intensity. Using genetic programming, the ALI was determined to be a better predictor of these three health measures than any subcombination of ALI components among cases, but not controls.

Adult↗

Multiclass cancer classification using gene expression profiling and probabilistic neural networks.

Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 or more. In addition, microarray data exhibit a high degree of noise. Most of the discussed methods do not adequately address the problem of dimensionality and noise. Furthermore, although machine learning and data mining methods are based on statistics, most such techniques do not address the biologist's requirement for sound mathematical confidence measures. Finally, most machine learning and data mining classification methods fail to incorporate misclassification costs, i.e. they are indifferent to the costs associated with false positive and false negative classifications. In this paper, we present a probabilistic neural network (PNN) model that addresses all these issues. The PNN model provides sound statistical confidences for its decisions, and it is able to model asymmetrical misclassification costs. Furthermore, we demonstrate the performance of the PNN for multiclass gene expression data sets. Here, we compare the performance of the PNN with two machine learning methods, a decision tree and a neural network. To assess and evaluate the performance of the classifiers, we use a lift-based scoring system that allows a fair comparison of different models. The PNN clearly outperformed the other models. The results demonstrate the successful application of the PNN model for multiclass cancer classification.

Artificial Intelligence↗