PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Genome-wide Association Studies of the Pathogenic Sphingosine-1-Phosphate Gene in Ulcerative Colitis.

BACKGROUND: Ulcerative colitis (UC) is a chronic inflammatory bowel disease that can lead to malignancies over time. Sphingosine-1-phosphate (S1P) receptor signaling affects lymphocyte trafficking and vascular integrity, influencing intestinal inflammation. This study aimed to identify S1P-related key genes in UC. METHODS: Differentially expressed genes (DEGs) between the UC and control groups were analyzed in the GSE87473 (training) dataset. Genes overlapping between the DEGs and S1P-related genes were considered candidate genes. These genes were incorporated into machine learning algorithms and subjected to expression analysis to identify key genes. Gene functions were determined through a gene–gene interaction network, enrichment analysis, and immune cell infiltration analysis. In addition, transcription factor–mRNA and mRNA–miRNA–lncRNA networks were constructed. Finally, reverse transcription–quantitative polymerase chain reaction (RT-qPCR) was performed to evaluate the expression of key candidate genes in UC and control tissues. RESULTS: This study identified two key genes (SPHK2 and SPNS2) associated with UC. Notably, SPHK2 expression was lower and SPNS2 expression was higher in the UC group in both training and validation datasets and in clinical UC tissues (RT-qPCR). The area under the curve values of SPHK2 and SPNS2 exceeded 0.7 in both datasets, indicating that the genes had good diagnostic efficacy for UC. Consistently, the nomogram showed that the two genes had promising diagnostic value in UC. SPHK2 and SPNS2 were found to be localized to the plasma membrane. The correlations of the two genes with different immune cells showed significantly opposite trends. In particular, SPHK2 had the strongest positive correlation with M2 macrophages (r = 0.6) and the strongest negative correlation with neutrophils. Moreover, mRNA–miRNA–lncRNA and transcription factor– mRNA networks of the key genes were constructed. CONCLUSION: This study suggests that SPHK2 and SPNS2 are key genes associated with UC, highlighting their potential as effective diagnostic biomarkers.

Humans↗

ORBIT: Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space for cancer driver gene identification.

Accurate identification of cancer driver genes is crucial for precision oncology but remains challenging due to the complexity of integrating heterogeneous data and modeling dynamic biological systems. To address these limitations, we propose ORBIT (Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space). Our framework synergistically fuses multi-omics profiles with functional network data using a context-adaptive graph reweighting mechanism to capture cancer-specific dynamics. The model employs a bi-prototype contrastive learning strategy within hyperbolic space, which aligns gene representations around distinct driver and non-driver semantic anchors while preserving the intrinsic hierarchy of biological networks. Comprehensive evaluations demonstrate that ORBIT achieves highly competitive stability in pan-cancer analysis while consistently outperforming state-of-the-art methods in cancer-specific predictions. Furthermore, functional enrichment analysis confirms that the model effectively segregates core cancer pathways, and drug sensitivity profiling validates the clinical relevance of the identified drivers. By integrating hyperbolic geometry with context-adaptive learning, ORBIT offers a robust and interpretable paradigm for precision medicine. The source codes and datasets are publicly accessible at https://github.com/spcho-dev/ORBIT.

Humans↗

Reconstructing muscle activation during normal walking: a comparison of symbolic and connectionist machine learning techniques.

One symbolic (rule-based inductive learning) and one connectionist (neural network) machine learning technique were used to reconstruct muscle activation patterns from kinematic data measured during normal human walking at several speeds. The activation patterns (or desired outputs) consisted of surface electromyographic (EMG) signals from the semitendinosus and vastus medialis muscles. The inputs consisted of flexion and extension angles measured at the hip and knee of the ipsilateral leg, their first and second derivatives, and bilateral foot contact information. The training set consisted of data from six trials, at two different speeds. The testing set consisted of data from two additional trials (one at each speed), which were not in the training set. It was possible to reconstruct the muscular activation at both speeds using both techniques. Timing of the reconstructed signals was accurate. The integrated value of the activation bursts was less accurate. The neural network gave a continuous output, whereas the rule-based inductive learning rule tree gave a quantised activation level. The advantage of rule-based inductive learning was that the rules used were both explicit and comprehensible, whilst the rules used by the neural network were implicit within its structure and not easily comprehended. The neural network was able to reconstruct the activation patterns of both muscles from one network, whereas two separate rule sets were needed for the rule-based technique. It is concluded that machine learning techniques, in comparison to explicit inverse muscular skeletal models, show good promise in modelling nearly cyclic movements such as locomotion at varying walking speeds.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learning.

MOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780.

Chromatin↗

HIDE: a new hybrid environment for the design of custom-made hip prosthesis.

This technical note describes a new software environment (HIPCOM design environment, HIDE) for the design of custom-made total hip replacements. These devices are frequently designed using general-purpose mechanical computer-aided design (CAD) programs using a set of bone contours extracted from the computer tomography (CT) images as anatomical reference. On the contrary, the HIDE system was developed to let the operator directly design the stem shape onto the CT images in a single-step operation. The operator can directly import CT data in DICOM format or use special functions to reconvert to a digital stack, the CT images printed on a radiological film. Once the stack of CT images is loaded, the operator can design the implant shape by imposing control sections directly on the CT images. The interpolation of these control sections produces the basic 3D shape of the custom-made stem. The shape is then exported to the CAD-computer-aided manufacturing (CAM) program to refine the design and to generate the part program to manufacture the implant with a CNC tooling machine. Using HIDE, the duration of design steps it affected was reduced by more than 50% with respect to the standard method in use at the manufacturer site. HIDE also improved the accuracy and the repeatability of the whole procedure. The learning curve became flat after only ten cases. These good results were achieved because of the integration of the vectorial description of the prosthetic component with the raster description of the CT data that allowed the designer to use all details available in the CT images.

Computer-Aided Design↗

ProMeta: a meta-learning framework for robust disease diagnosis and prediction from plasma proteomics.

MOTIVATION: The plasma proteome offers a dynamic window of human health, capturing the real-time intersections between genetics and physiology. However, the application of deep learning to proteomics is currently hindered by a reliance on large-scale labeled datasets, rendering standard models ineffective for rare or novel diseases where patient samples are inherently scarce. RESULTS: Here, we present ProMeta, a meta-learning framework designed to enable robust disease modeling under extreme data restrictions. By integrating knowledge-guided pathway encoding with bi-level meta-optimization, ProMeta projects unstructured proteomic profiles into biologically interpretable functional tokens. This architecture allows the model to learn a global initialization containing transferable biological priors from biobank-scale data, facilitating rapid adaptation to novel tasks. Through comprehensive benchmark experiments, ProMeta consistently outperformed transfer learning and traditional machine learning baselines in both disease diagnosis and prediction tasks. In the most challenging 4-shot scenarios (utilizing only 2 cases and 2 controls), the model achieved robust generalization with an average AUROC of ∼0.69, representing a 24.6% relative improvement over the best-performing baseline methods. Mechanistic investigation revealed that ProMeta disentangles cases from controls in the latent space prior to task-specific adaptation, confirming the acquisition of universal biological rules rather than rote memorization. Furthermore, gradient-based interpretation identified disease-specific protein biomarkers and functional pathways consistent with known pathophysiology. Collectively, ProMeta overcomes the data-scarcity bottleneck in precision medicine, providing a scalable, interpretable framework for characterizing the full spectrum of human diseases, particularly for rare conditions lacking extensive clinical cohorts. AVAILABILITY AND IMPLEMENTATION: The source code of ProMeta is available at GitHub (https://github.com/lihan97/ProMeta).

Proteomics↗

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology↗

Integrated metagenomic and metabolomic analysis identifies severity-specific inflammatory and metabolic signatures in post-stroke depression.

Post-stroke depression (PSD) is a common complication that significantly impacts patient prognosis. This study aimed to systematically characterize the associations among gut microbial ecology, metabolic profiles, and inflammatory responses across different severities of PSD. We conducted metagenomic sequencing, non-targeted metabolomics, and serum cytokine analysis (IL-1β, IL-6, IL-10, IL-18, TNF-α, IFN-γ, and CRP) in 91 patients with varying degrees of PSD and non-PSD controls. Bioinformatics analyzes were employed to construct multi-omics association networks and machine learning models. Results indicated that PSD patients exhibited significantly increased gut microbiota alpha-diversity, suggesting dysbiosis. Mild depression was characterized by compensatory neural signaling activation, whereas the moderate depression group exhibited abnormalities in tryptophan/indole metabolism, oxidative stress-related metabolic imbalances, and functional decompensation. Further analyzes suggested that Alistipes, Blautia_A, Evtepia gabavorous, and Lachnospira were associated with inflammatory features, GABA-related metabolic alterations, aromatic amino acid/indole metabolism, and lipid-amino acid metabolism, respectively. Under a more rigorous 10-fold cross-validation framework, the performance of different multi-omics combination models showed heterogeneity; however, some combinations still demonstrated superior discriminatory ability compared to single-omics approaches. This study provides multi-omics clues suggesting associations between different PSD severity levels and features such as increased Alistipes abundance, reduced antioxidant capacity, and altered tryptophan metabolism. It provides candidate biomarker combinations that may be useful for PSD stratification and suggests that the gut microbiome may represent a potential target for future PSD intervention. In summary, PSD may be associated with dynamic alterations along the "gut-brain-inflammation-metabolism" axis. These findings provide integrated evidence for microbial, metabolic, and inflammatory abnormalities across different PSD severity levels, but still require validation in larger samples, longitudinal cohorts, and mechanistic studies.

Humans↗

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans↗

Can host genetics transform the sustainable control of tropical theileriosis? Insights from the Tick-Theileria interface.

Tropical theileriosis, caused by the tick-transmitted apicomplexan parasite Theileria annulata, remains a major constraint on cattle production across North Africa, the Mediterranean basin, the Middle East and South Asia. Current control depends on acaricides, the theilericidal drug buparvaquone and live attenuated schizont vaccines, but acaricide resistance, buparvaquone-resistance mutations and the logistical demands of vaccination are eroding the sustainability of these tools. Host genetics offers a complementary and durable alternative. Indigenous Bos indicus breeds are consistently more resistant to ticks and tolerate T. annulata infection better than exotic Bos taurus cattle, and this advantage has a measurable heritable component. Unlike previous reviews, which treat tick resistance, T. annulata immunobiology and livestock genomic selection as separate subjects, we integrate all three and assess host genetics specifically against the failure modes of current control. We review the tick, parasite and host interface, the evidence for natural resistance, and the genetic and immunological mechanisms involved, including signal-regulatory protein, bovine major histocompatibility complex class II and inflammatory pathway genes. We then assess whether genomic selection, multi-omics, machine learning and gene editing can translate these mechanisms into resistant cattle, and we weigh the biological, economic and infrastructural barriers to implementation. The evidence indicates that host genetics will not replace existing control but could reduce reliance on acaricides and chemotherapy. That contribution remains prospective rather than demonstrated: no resistance marker for T. annulata has yet been validated, prediction accuracies are moderate and transfer poorly between breeds, and no endemic production system has implemented selection for resistance.

Animals↗

ABC stenosis morphology classification and outcome of coronary angioplasty: reassessment with computing techniques.

BACKGROUND: The American College of Cardiology/American Heart Association (ACC/AHA) stenosis morphology classification (MC) stratifies coronary lesions for probability of success and complications after coronary angioplasty (PTCA). Modern computing techniques were used to evaluate the individual predictive value of MC in random PTCA cases. METHODS AND RESULTS: MC was attributed to the target lesions by consensus of 2 observers. The predictive value regarding procedural success (PS) and major adverse cardiac events (MACE) of MC was analyzed by conventional logistic regression analyses and by inductive machine learning models. The study was adequately powered for the methods applied with 325 target lesions of 250 cases. Overall, PS decreased and MACE increased from type A to type C lesions. Regression analysis identified no single factor as predictive. Logistic regression showed an error rate of 42%. Machine learning techniques achieved an individual predictive error of only 10%, which could be further reduced to 2% by addition of parameters. For PS, MC parameters showed a high ranking for building the model. For MACE, variables of the medical history showed more impact. CONCLUSIONS: MC per se cannot individually predict PS or MACE. However, when all MC parameters are integrated together with additional lesion-specific and history variables, a high individual predictive value can be achieved. This technique may be clinically helpful for risk stratification in the catheterization laboratory and improvement of classification systems in interventional cardiology.

Algorithms↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Global microbial DNA signatures of temperature and nutrient limitation across ecosystems.

Microbial genomes continuously adapt to environmental conditions, but identifying universal signatures of adaptation remains challenging. Here we show that environmental temperature can be accurately predicted across ecosystems from DNA composition alone (R2&#x2009;=&#x2009;0.75), using tetranucleotide frequencies from 1,235 marine and soil metagenomes and a machine learning approach. This predictive signal was also apparent within individual taxa, consistent with a fundamental temperature-associated signature. By contrast, GC content exhibited opposite correlations with temperature in soil (positive) and marine (negative) environments. This phenomenon was probably driven by differences in nutrient availability, as GC content increases with nutrients while nutrients decrease with temperature in marine samples. By integrating these observations, we identified specific tetranucleotides, with 50% GC, that displayed consistent and robust temperature correlations across environments and may have contributed to the stability of predictions. This work highlights metagenome-wide DNA-temperature associations, relevant for understanding microbial community responses to global changes.

Journal Article↗

Neurotoxic lesions of the rat perirhinal cortex fail to disrupt the acquisition or performance of tests of allocentric spatial memory.

Rats with neurotoxic lesions of the perirhinal cortex (n = 9) were compared with sham controls (n = 14) on a working memory task in the radial arm maze. Rats were trained under varying levels of proactive interference and with different retention intervals. Finally, performance was assessed when the maze was switched to a novel room. None of these manipulations differentially impaired rats with perirhinal lesions. Rats were next trained on delayed matching-to-place in the water maze. Even with retention delays of 30 min, there was no evidence of a deficit. Although interactions between the perirhinal cortex and hippocampus may be important for integrating object-place information, the perirhinal cortex is often not necessary for tasks that selectively tax allocentric spatial memory.

Animals↗

Integration of Genetic Information to Improve Brain Age Gap Estimation Models in the UK Biobank.

Neurodegeneration occurs when the body's central nervous system becomes impaired as a person ages, which can happen at an accelerated pace. Neurodegeneration impairs quality of life, affecting essential functions, including memory and the ability to self-care. Genetics play an important role in neurodegeneration and longevity. Brain age gap estimation (BrainAGE) is a biomarker that quantifies the difference between a machine learning model-predicted biological age of the brain and the true chronological age for healthy subjects; however, a large portion of the variance remains unaccounted for in these models, attributed to individual differences. This study focuses on predicting the BrainAGE more accurately, aided by genetic information associated with neurodegeneration. To achieve this, a BrainAGE model was developed based on MRI measures, and then the associated genes were determined with a Genome-Wide Association Study. Subsequently, genetic information was incorporated into the models. The incorporation of genetic information yielded improvements in the model performances by 7% to 12%, showing that the incorporation of genetic information can notably reduce unexplained variance. This work helps to define new ways of determining persons susceptible to neurological aging decline and reveals genes for targeted precision medicine therapies.

Humans↗

Real-time computing without stable states: a new framework for neural computation based on perturbations.

A key challenge for neural modeling is to explain how a continuous stream of multimodal input from a rapidly changing environment can be processed by stereotypical recurrent circuits of integrate-and-fire neurons in real time. We propose a new computational model for real-time computing on time-varying input that provides an alternative to paradigms based on Turing machines or attractor neural networks. It does not require a task-dependent construction of neural circuits. Instead, it is based on principles of high-dimensional dynamical systems in combination with statistical learning theory and can be implemented on generic evolved or found recurrent circuitry. It is shown that the inherent transient dynamics of the high-dimensional dynamical system formed by a sufficiently large and heterogeneous neural circuit may serve as universal analog fading memory. Readout neurons can learn to extract in real time from the current state of such recurrent neural circuit information about current and past inputs that may be needed for diverse tasks. Stable internal states are not required for giving a stable output, since transient internal states can be transformed by readout neurons into stable target outputs due to the high dimensionality of the dynamical system. Our approach is based on a rigorous computational model, the liquid state machine, that, unlike Turing machines, does not require sequential transitions between well-defined discrete internal states. It is supported, as the Turing machine is, by rigorous mathematical results that predict universal computational power under idealized conditions, but for the biologically more realistic scenario of real-time processing of time-varying inputs. Our approach provides new perspectives for the interpretation of neural coding, the design of experiments and data analysis in neurophysiology, and the solution of problems in robotics and neurotechnology.

Action Potentials↗

A generalized hidden Markov model for the recognition of human genes in DNA.

We present a statistical model of genes in DNA. A Generalized Hidden Markov Model (GHMM) provides the framework for describing the grammar of a legal parse of a DNA sequence (Stormo & Haussler 1994). Probabilities are assigned to transitions between states in the GHMM and to the generation of each nucleotide base given a particular state. Machine learning techniques are applied to optimize these probabilities using a standardized training set. Given a new candidate sequence, the best parse is deduced from the model using a dynamic programming algorithm to identify the path through the model with maximum probability. The GHMM is flexible and modular, so new sensors and additional states can be inserted easily. In addition, it provides simple solutions for integrating cardinality constraints, reading frame constraints, "indels", and homology searching. The description and results of an implementation of such a gene-finding model, called Genie, is presented. The exon sensor is a codon frequency model conditioned on windowed nucleotide frequency and the preceding codon. Two neural networks are used, as in (Brunak, Engelbrecht, & Knudsen 1991), for splice site prediction. We show that this simple model performs quite well. For a cross-validated standard test set of 304 genes [ftp:@www-hgc.lbl.gov/pub/genesets] in human DNA, our gene-finding system identified up to 85% of protein-coding bases correctly with a specificity of 80%. 58% of exons were exactly identified with a specificity of 51%. Genie is shown to perform favorably compared with several other gene-finding systems.

Chromosomes, Human↗

Decision tree-based formation of consensus protein secondary structure prediction.

MOTIVATION: Prediction of protein secondary structure provides information that is useful for other prediction methods like fold recognition and ab initio 3D prediction. A consensus prediction constructed from the output of several methods should yield more reliable results than each of the individual methods. METHOD: We present an approach that reveals subtle but systematic differences in the output of different secondary structure prediction methods allowing the derivation of coherent consensus predictions. The method uses a machine learning technique that builds decision trees from existing data. RESULTS: The first results of our analysis show that consensus prediction of protein secondary structure may be improved both quantitatively and qualitatively.

Algorithms↗