PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗

Analysis of end-stage renal disease mediated by cuproptosis-related genes.

OBJECTIVE: The complex pathophysiological mechanism of end-stage renal disease (ESRD) has not been fully understood. Cuproptosis is a newly discovered type of programmed cell death. Therefore, this study attempts to clarify the relationship between cuproptosis-related genes (CRGs) and the phenotype of ESRD. MATERIALS AND METHODS: The National Center for Biological Information Gene Expression Omnibus database was applied to obtain the GSE37171 dataset comprising whole-genome microarray analysis of peripheral blood samples. A 3 : 1 case-control design was employed with 75 ESRD patients and 20 healthy controls who were frequency-matched for age, sex, and ethnicity. Based on differentially expressed genes (DEGs) and genes related to cuproptosis, CRGs were identified. Thereafter, we explored two different subpopulations based on the cuproptosis gene and analyzed their expression and immune infiltration. Genes specific to the CRG cluster were identified through the weighted gene co-expression network analysis algorithm, and the best prediction model was determined and verified by four machine learning methods. RESULTS: The study identified 14 differentially expressed CRGs, among which ATP7B, SLC31A1, LIAS, LIPT1, DLD, MTF1, CDKN2A, DBT, and DLST had relatively high expression levels in the ESRD samples. Compared with the control group, expression levels of FDX1, DLAT, PDHA1, PDHB, and GLS were significantly lower in the ESRD group, and CRGs played a key role in the regulation of immune infiltration in ESRD. Two cuproptosis-related molecular clusters were identified in the ESRD samples. Cluster2 was more correlated with the immune infiltration of ESRD. By analyzing the intersection points between CRG cluster and key genes of ESRD, a total of 888 specific DEGs were identified. Functional differences related to specific DEGs were further explored using gene set variation analysis. Five significant genes (SMC5, USP47, USP53, AGA, and DMXL1) were identified by the support vector machine model as key predictors for ESRD disease risk, achieving an area under the curve (AUC) of 1.00 in internal validation. However, external validation in independent cohorts is required prior to clinical application. Individual gene analysis showed an AUC > 0.81 in discriminating ESRD patients from healthy controls, and the expression of all 5 genes in ESRD patients was significantly lower than in the control group. CONCLUSION: This study clarified the relationship between CRGs and the phenotype of ESRD, analyzed their specific roles in the immune microenvironment, and obtained a predictive model, providing new insights for the study of its potential therapeutic targets.

Humans↗

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal↗

Biological Parts in Yeast Synthetic Biology: From Regulatory Elements to Predictive Design Platforms.

Yeasts, particularly Saccharomyces cerevisiae, are important eukaryotic chassis for synthetic biology because of their tractable genetics, versatile toolkits, and broad utility in metabolic engineering and functional genomics. Progress in this field has been driven by biological parts that enable programmable control of gene expression and cellular behavior. Early efforts focused mainly on promoters, terminators, and other regulatory elements for tuning individual genes. However, as engineering expanded to multigene pathways, genetic circuits, and dynamic regulatory systems, the limits of part-centric design became clear. Part performance is often shaped by genomic context, chromatin state, host physiology, and interactions with other components, which restricts modularity and predictability. In response, yeast synthetic biology is shifting toward integrated design frameworks combining multilayer regulation, standardized assembly, automated experimentation, and computational modeling. This review provides an integrated perspective on the evolution of biological parts across DNA-, RNA-, and protein-level regulation, connecting these advances with assembly frameworks, biofoundries, and machine learning to trace the trajectory from part-centric engineering toward predictive, system-level design in yeast synthetic biology.

Biofoundry↗

RNA secondary structure prediction from sequence alignments using a network of k-nearest neighbor classifiers.

We present a machine learning method (a hierarchical network of k-nearest neighbor classifiers) that uses an RNA sequence alignment in order to predict a consensus RNA secondary structure. The input to the network is the mutual information, the fraction of complementary nucleotides, and a novel consensus RNAfold secondary structure prediction of a pair of alignment columns and its nearest neighbors. Given this input, the network computes a prediction as to whether a particular pair of alignment columns corresponds to a base pair. By using a comprehensive test set of 49 RFAM alignments, the program KNetFold achieves an average Matthews correlation coefficient of 0.81. This is a significant improvement compared with the secondary structure prediction methods PFOLD and RNAalifold. By using the example of archaeal RNase P, we show that the program can also predict pseudoknot interactions.

Algorithms↗

A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.

We have introduced a new method of protein secondary structure prediction which is based on the theory of support vector machine (SVM). SVM represents a new approach to supervised pattern classification which has been successfully applied to a wide range of pattern recognition problems, including object recognition, speaker identification, gene function prediction with microarray expression profile, etc. In these cases, the performance of SVM either matches or is significantly better than that of traditional machine learning approaches, including neural networks.The first use of the SVM approach to predict protein secondary structure is described here. Unlike the previous studies, we first constructed several binary classifiers, then assembled a tertiary classifier for three secondary structure states (helix, sheet and coil) based on these binary classifiers. The SVM method achieved a good performance of segment overlap accuracy SOV=76.2 % through sevenfold cross validation on a database of 513 non-homologous protein chains with multiple sequence alignments, which out-performs existing methods. Meanwhile three-state overall per-residue accuracy Q(3) achieved 73.5 %, which is at least comparable to existing single prediction methods. Furthermore a useful "reliability index" for the predictions was developed. In addition, SVM has many attractive features, including effective avoidance of overfitting, the ability to handle large feature spaces, information condensing of the given data set, etc. The SVM method is conveniently applied to many other pattern classification tasks in biology.

Computer Simulation↗

Logistic regression model: an assessment of variability of predictions.

Risk prediction models available for cardiovascular prevention are statistical or based on machine learning methods. This paper investigates whether the logistic regression method can be considered as reference for validation of other methods. In order to test the stability of the predictions using this method, we performed two types of analyses on 50 random training and test samples drawn from the same database. In first analyses three models were obtained by forced entry of different sets of four variables. In second analyses, models were built with increasing number of predictive variables. The predictive performance was assessed by the area under the ROC curve. Although across-samples variability is low for a given model, it is large enough to lead to wrong conclusions when comparing different prediction methods. We also suggest that a low events-per-variable ratio alters the stability of a model's coefficients but does not affect the variability of prediction performance.

Area Under Curve↗

SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput structural proteomics.

High-throughput structural proteomics is expected to generate considerable amounts of data on the progress of structure determination for many proteins. For each protein this includes information about cloning, expression, purification, biophysical characterization and structure determination via NMR spectroscopy or X-ray crystallography. It will be essential to develop specifications and ontologies for standardizing this information to make it amenable to retrospective analysis. To this end we created the SPINE database and analysis system for the Northeast Structural Genomics Consortium. SPINE, which is available at bioinfo.mbb.yale.edu/nesg or nesg.org, is specifically designed to enable distributed scientific collaboration via the Internet. It was designed not just as an information repository but as an active vehicle to standardize proteomics data in a form that would enable systematic data mining. The system features an intuitive user interface for interactive retrieval and modification of expression construct data, query forms designed to track global project progress and external links to many other resources. Currently the database contains experimental data on 985 constructs, of which 740 are drawn from Methanobacterium thermoautotrophicum, 123 from Saccharomyces cerevisiae, 93 from Caenorhabditis elegans and the remainder from other organisms. We developed a comprehensive set of data mining features for each protein, including several related to experimental progress (e.g. expression level, solubility and crystallization) and 42 based on the underlying protein sequence (e.g. amino acid composition, secondary structure and occurrence of low complexity regions). We demonstrate in detail the application of a particular machine learning approach, decision trees, to the tasks of predicting a protein's solubility and propensity to crystallize based on sequence features. We are able to extract a number of key rules from our trees, in particular that soluble proteins tend to have significantly more acidic residues and fewer hydrophobic stretches than insoluble ones. One of the characteristics of proteomics data sets, currently and in the foreseeable future, is their intermediate size ( approximately 500-5000 data points). This creates a number of issues in relation to error estimation. Initially we estimate the overall error in our trees based on standard cross-validation. However, this leaves out a significant fraction of the data in model construction and does not give error estimates on individual rules. Therefore, we present alternative methods to estimate the error in particular rules.

Animals↗

Uncovering the genetic architecture of ME/CFS: a precision approach reveals impact of rare monogenic variation.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a disabling and heterogeneous disorder lacking validated biomarkers or targeted therapies. Clinical variability and elusive pathophysiology hinder progress toward effective diagnostics and treatment. Core symptoms include persistent fatigue, post-exertional malaise, unrefreshing sleep, cognitive dysfunction, and pain. We tested whether an individualized, “n-of-1” genomic and transcriptomic framework combined with comprehensive, participant-informed phenotyping could reveal molecular signatures unique to each patient. METHODS: Clinical-grade whole-genome sequencing was conducted in 31 affected individuals from 25 families, with RNA-seq performed on a subset (16 affected, 7 unaffected) using blood samples. Machine-learning assisted variant triage, transcript-aware damage prediction, and expert review identified pathogenic or likely pathogenic variants in 8 of 25 probands (32%) and 12 of 31 affected individuals (39%). RESULTS: Findings revealed marked genetic heterogeneity, including large-effect rare and more common variants. Implicated pathways included ATP generation, oxidative phosphorylation, fatty acid oxidation; regulation of glycolysis, amino acid and lipid turnover; ion and solute homeostasis; synaptic signaling, excitability, oxygen transport, and muscle integrity, resilience, and post-exertional recovery; previously implicated processes. Plausible modifiers influencing disease onset, severity, and relapsing–remitting patterns and possibly explaining intrafamilial variability and inconsistent findings across studies, were also identified. Despite gene-level diversity, downstream effects converged on impaired energy production, reduced stress resilience, and vulnerability to post-exertional metabolic failure; disruptions consistent with core ME/CFS symptoms of exertional intolerance, cognitive fog, and fatigue. CONCLUSIONS: Our findings support the hypothesis that at least a subset of ME/CFS cases represent distinct molecular disorders that converge on shared physiological pathways. Validation in larger, more diverse cohorts will be essential to test this hypothesis and establish generalizability, but increase size alone is unlikely to resolve causation in a disorder defined by rarity, heterogeneity, and molecular complexity. We suggest that progress will require experimental designs that integrate individual-level genomic data with deep, participant-informed deep phenotyping, capturing the combined effects of rare and common variants and environmental modifiers on disease expression and progression. We believe that an individualized precision medicine framework will uncover molecular drivers and modifiers of ME/CFS previously obscured by heterogeneity, enabling biologically informed stratification, improved trial design, biomarker discovery, and targeted interventions in this historically neglected condition.

Humans↗

Combination of computational techniques and RNAi reveal targets in Anopheles gambiae for malaria vector control.

Increasing reports of insecticide resistance continue to hamper the gains of vector control strategies in curbing malaria transmission. This makes identifying new insecticide targets or alternative vector control strategies necessary. CLassifier of Essentiality AcRoss EukaRyote (CLEARER), a leave-one-organism-out cross-validation machine learning classifier for essential genes, was used to predict essential genes in Anopheles gambiae and selected predicted genes experimentally validated. The CLEARER algorithm was trained on six model organisms: Caenorhabditis elegans, Drosophila melanogaster, Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Schizosaccharomyces pombe, and employed to identify essential genes in An. gambiae. Of the 10,426 genes in An. gambiae, 1,946 genes (18.7%) were predicted to be Cellular Essential Genes (CEGs), 1716 (16.5%) to be Organism Essential Genes (OEGs), and 852 genes (8.2%) to be essential as both OEGs and CEGs. RNA interference (RNAi) was used to validate the top three highly expressed non-ribosomal predictions as probable vector control targets, by determining the effect of these genes on the survival of An. gambiae G3 mosquitoes. In addition, the effect of knockdown of arginase (AGAP008783) on Plasmodium berghei infection in mosquitoes was evaluated, an enzyme we computationally inferred earlier to be essential based on chokepoint analysis. Arginase and the top three genes, AGAP007406 (Elongation factor 1-alpha, Elf1), AGAP002076 (Heat shock 70kDa protein 1/8, HSP), AGAP009441 (Elongation factor 2, Elf2), had knockdown efficiencies of 91%, 75%, 63%, and 61%, respectively. While knockdown of HSP or Elf2 significantly reduced longevity of the mosquitoes (p<0.0001) compared to control groups, Elf1 or arginase knockdown had no effect on survival. However, arginase knockdown significantly reduced P. berghei oocytes counts in the midgut of mosquitoes when compared to LacZ-injected controls. The study reveals HSP and Elf2 as important contributors to mosquito survival and arginase as important for parasite development, hence placing them as possible targets for vector control.

Animals↗

Evaluating the C-section rate of different physician practices: using machine learning to model standard practice.

The C-section rate of a population of 22,175 expectant mothers is 16.8%; yet the 17 physician groups that serve this population have vastly different group C-section rates, ranging from 13% to 23%. Our goal is to determine retrospectively if the variations in the observed rates can be attributed to variations in the intrinsic risk of the patient sub-populations (i.e. some groups contain more "high-risk C-section" patients), or differences in physician practice (i.e. some groups do more C-sections). We apply machine learning to this problem by training models to predict standard practice from retrospective data. We then use the models of standard practice to evaluate the C-section rate of each physician practice. Our results indicate that although there is variation in intrinsic risk among the groups, there also is much variation in physician practice.

Artificial Intelligence↗

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms↗

Synthetic community Hi-C benchmarking provides a baseline for virus-host inferences.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least ten studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z &#x2265; 0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone - i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Genomics↗

Benchmarking with synthetic communities provides a baseline for virus-host inferences from Hi-C proximity linking.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least 10 studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z&#x2009;&#x2265;&#x2009;0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone, i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Benchmarking↗

A machine learning information retrieval approach to protein fold recognition.

MOTIVATION: Recognizing proteins that have similar tertiary structure is the key step of template-based protein structure prediction methods. Traditionally, a variety of alignment methods are used to identify similar folds, based on sequence similarity and sequence-structure compatibility. Although these methods are complementary, their integration has not been thoroughly exploited. Statistical machine learning methods provide tools for integrating multiple features, but so far these methods have been used primarily for protein and fold classification, rather than addressing the retrieval problem of fold recognition-finding a proper template for a given query protein. RESULTS: Here we present a two-stage machine learning, information retrieval, approach to fold recognition. First, we use alignment methods to derive pairwise similarity features for query-template protein pairs. We also use global profile-profile alignments in combination with predicted secondary structure, relative solvent accessibility, contact map and beta-strand pairing to extract pairwise structural compatibility features. Second, we apply support vector machines to these features to predict the structural relevance (i.e. in the same fold or not) of the query-template pairs. For each query, the continuous relevance scores are used to rank the templates. The FOLDpro approach is modular, scalable and effective. Compared with 11 other fold recognition methods, FOLDpro yields the best results in almost all standard categories on a comprehensive benchmark dataset. Using predictions of the top-ranked template, the sensitivity is approximately 85, 56, and 27% at the family, superfamily and fold levels respectively. Using the 5 top-ranked templates, the sensitivity increases to 90, 70, and 48%.

Algorithms↗

A computational approach toward label-free protein quantification using predicted peptide detectability.

We propose here a new concept of peptide detectability which could be an important factor in explaining the relationship between a protein's quantity and the peptides identified from it in a high-throughput proteomics experiment. We define peptide detectability as the probability of observing a peptide in a standard sample analyzed by a standard proteomics routine and argue that it is an intrinsic property of the peptide sequence and neighboring regions in the parent protein. To test this hypothesis we first used publicly available data and data from our own synthetic samples in which quantities of model proteins were controlled. We then applied machine learning approaches to demonstrate that peptide detectability can be predicted from its sequence and the neighboring regions in the parent protein with satisfactory accuracy. The utility of this approach for protein quantification is demonstrated by peptides with higher detectability generally being identified at lower concentrations over those with lower detectability in the synthetic protein mixtures. These results establish a direct link between protein concentration and peptide detectability. We show that for each protein there exists a level of peptide detectability above which peptides are detected and below which peptides are not detected in an experiment. We call this level the minimum acceptable detectability for identified peptides (MDIP) which can be calibrated to predict protein concentration. Triplicate analysis of a biological sample showed that these MDIP values are consistent among the three data sets.

Algorithms↗

Drug discovery using support vector machines. The case studies of drug-likeness, agrochemical-likeness, and enzyme inhibition predictions.

Support Vector Machines (SVM) is a powerful classification and regression tool that is becoming increasingly popular in various machine learning applications. We tested the ability of SVM, in comparison with well-known neural network techniques, to predict drug-likeness and agrochemical-likeness for large compound collections. For both kinds of data, SVM outperforms various neural networks using the same set of descriptors. We also used SVM for estimating the activity of Carbonic Anhydrase II (CA II) enzyme inhibitors and found that the prediction quality of our SVM model is better than that reported earlier for conventional QSAR. Model characteristics and data set features were studied in detail.

Agrochemicals↗

Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.

BACKGROUND: The number of protein sequences deriving from genome sequencing projects is outpacing our knowledge about the function of these proteins. With the gap between experimentally characterized and uncharacterized proteins continuing to widen, it is necessary to develop new computational methods and tools for functional prediction. Knowledge of catalytic sites provides a valuable insight into protein function. Although many computational methods have been developed to predict catalytic residues and active sites, their accuracy remains low, with a significant number of false positives. In this paper, we present a novel method for the prediction of catalytic sites, using a carefully selected, supervised machine learning algorithm coupled with an optimal discriminative set of protein sequence conservation and structural properties. RESULTS: To determine the best machine learning algorithm, 26 classifiers in the WEKA software package were compared using a benchmarking dataset of 79 enzymes with 254 catalytic residues in a 10-fold cross-validation analysis. Each residue of the dataset was represented by a set of 24 residue properties previously shown to be of functional relevance, as well as a label {+1/-1} to indicate catalytic/non-catalytic residue. The best-performing algorithm was the Sequential Minimal Optimization (SMO) algorithm, which is a Support Vector Machine (SVM). The Wrapper Subset Selection algorithm further selected seven of the 24 attributes as an optimal subset of residue properties, with sequence conservation, catalytic propensities of amino acids, and relative position on protein surface being the most important features. CONCLUSION: The SMO algorithm with 7 selected attributes correctly predicted 228 of the 254 catalytic residues, with an overall predictive accuracy of more than 86%. Missing only 10.2% of the catalytic residues, the method captures the fundamental features of catalytic residues and can be used as a "catalytic residue filter" to facilitate experimental identification of catalytic residues for proteins with known structure but unknown function.

Algorithms↗