PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “machine learning prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Rule generation for protein secondary structure prediction with support vector machines and decision tree.

Support vector machines (SVMs) have shown strong generalization ability in a number of application areas, including protein structure prediction. However, the poor comprehensibility hinders the success of the SVM for protein structure prediction. The explanation of how a decision made is important for accepting the machine learning technology, especially for applications such as bioinformatics. The reasonable interpretation is not only useful to guide the "wet experiments," but also the extracted rules are helpful to integrate computational intelligence with symbolic AI systems for advanced deduction. On the other hand, a decision tree has good comprehensibility. In this paper, a novel approach to rule generation for protein secondary structure prediction by integrating merits of both the SVM and decision tree is presented. This approach combines the SVM with decision tree into a new algorithm called SVM_ DT, which proceeds in three steps. This algorithm first trains an SVM. Then, a new training set is generated through careful selection from the output of the SVM. Finally, the obtained training set is used to train a decision tree learning system and to extract the corresponding rule sets. The results of the experiments of protein secondary structure prediction on RS126 data set show that the comprehensibility of SVM_DT is much better than that of the SVM. Moreover, the generalization ability of SVM_DT is better than that of C4.5 decision trees and is similar to that of the SVM. Hence, SVM_DT can be used not only for prediction, but also for guiding biological experiments.

Algorithms↗

Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression data.

BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains.

Animals↗

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive↗

Decoding the distribution, structure-function-redox potential relationship and recent advances in fungal laccases: a systematic approach.

Laccases, categorized as multicopper oxidases, are recognized for their multifaceted roles in ecosystems and their utility in diverse industrial applications. Laccases from higher fungi, specifically Ascomycota and Basidiomycota, have garnered significant research interest due to their elevated redox potentials and their capacity to degrade lignin in decaying wood, alongside other industrial uses. Here, we have conducted a comprehensive and systematic analysis on fungal laccases using Web of Science, Scopus, PubMed, and ScienceDirect. The genomic distribution, phylogenetic affiliation, and structural organization of laccase-encoding genes in higher fungal species were investigated, as were the catalytic mechanisms of the corresponding enzymes. Additionally, the study explores the correlation between structural domains and redox potential, as well as the impact of post-translational modifications like glycosylation on enzyme activity. Furthermore, the recent advancements in laccase engineering, employing strategies such as rational design, directed evolution, and heterologous expression are discussed. The review also explores the scope of "artificial intelligence and machine learning" in deducing the structure-function relationships, optimizing codon usage, predicting signal peptides, enhancing enzymatic performance, and developing host-specific genetic engineering techniques is also discussed for tailoring fungal laccases to meet the demands of industrial biocatalysis for improved activity and stability.

Laccase↗

Modelling the structure and function of enzymes by machine learning.

A machine learning program, GOLEM, has been applied to two problems: (1) the prediction of protein secondary structure from sequence and (2) modelling a quantitative structure-activity relationship in drug design. GOLEM takes as input observations and combines them with background knowledge of chemistry to yield rules expressed as stereochemical principles for prediction. The secondary structure prediction was explored on the alpha/alpha class of proteins; on an unrelated test set it yielded 81% accuracy. The rules from GOLEM defined patterns of residues forming alpha-helices. The system studied for drug design was the activities of trimethoprim analogues binding to E. coli dihydrofolate reductase. The GOLEM rules were a better model than standard regression approaches. More importantly, these rules described the chemical properties of the enzyme-binding site that were in broad agreement with the crystallographic structure.

Amino Acid Sequence↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article↗

Prediction of estrogen receptor agonists and characterization of associated molecular descriptors by statistical learning methods.

Specific estrogen receptor (ER) agonists have been used for hormone replacement therapy, contraception, osteoporosis prevention, and prostate cancer treatment. Some ER agonists and partial-agonists induce cancer and endocrine function disruption. Methods for predicting ER agonists are useful for facilitating drug discovery and chemical safety evaluation. Structure-activity relationships and rule-based decision forest models have been derived for predicting ER binders at impressive accuracies of 87.1-97.6% for ER binders and 80.2-96.0% for ER non-binders. However, these are not designed for identifying ER agonists and they were developed from a subset of known ER binders. This work explored several statistical learning methods (support vector machines, k-nearest neighbor, probabilistic neural network and C4.5 decision tree) for predicting ER agonists from comprehensive set of known ER agonists and other compounds. The corresponding prediction systems were developed and tested by using 243 ER agonists and 463 ER non-agonists, respectively, which are significantly larger in number and structural diversity than those in previous studies. A feature selection method was used for selecting molecular descriptors responsible for distinguishing ER agonists from non-agonists, some of which are consistent with those used in other studies and the findings from X-ray crystallography data. The prediction accuracies of these methods are comparable to those of earlier studies despite the use of significantly more diverse range of compounds. SVM gives the best accuracy of 88.9% for ER agonists and 98.1% for non-agonists. Our study suggests that statistical learning methods such as SVM are potentially useful for facilitating the prediction of ER agonists and for characterizing the molecular descriptors associated with ER agonists.

Forecasting↗

Prediction of the axillary lymph node status in mammary cancer on the basis of clinicopathological data and flow cytometry.

Axillary lymph node status is a major prognostic factor in mammary carcinoma. It is clinically desirable to predict the axillary lymph node status from data from the mammary cancer specimen. In the study, the axillary lymph node status, routine histological parameters and flow-cytometric data were retrospectively obtained from 1139 specimens of invasive mammary cancer. The ten variables: age, tumour type, tumour grade, tumour size, skin infiltration, lymphangiosis carcinomatosa, pT4 category, percentage of tumour cells in G2/M- and S-phases of the cell cycle, and ploidy index were considered as predictor variables, and the single variable lymph node metastasis pN (0 for pN0, or 1 for pN1 or pN2) was used as an output variable. A stepwise logistic regression analysis, with the axillary lymph node as a dependent variable, was used for feature selection. Only lymphangiosis carcinomatosa and tumour size proved to be significant as independent predictor variables; the other variables were non-contributory. Three paradigms with supervised learning rules (multilayer perceptron, learning vector quantisation and support vector machines) were used for the purpose of prediction. If any of these paradigms was used with the information from all ten input variables, 73% of cases could be correctly predicted, with specificity ranging from 82 to 84% and sensitivity ranging from 60 to 63%. If only the two significant input variables were used, lymphangiosis carcinomatosa and tumour diameter, the prediction accuracy was no worse. Nearly identical results were obtained by two different techniques of cross-validation (leave-one-out against ten-fold cross validation). It was concluded that: artificial neural networks can be used for risk stratification on the basis of routine data in individual cases of mammary cancer; and lymphangiosis carcinomatosa and tumour size are independent predictors of axillary lymph node metastasis in mammary cancer.

Algorithms↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

Two models for outcome prediction - a comparison of logistic regression and neural networks.

OBJECTIVES: Accurately predicting disease progress from a set of predictive variables is an important aspect of clinical work. For binary outcomes, the classical approach is to develop prognostic logistic regression (LR) models. Alternatively, machine learning algorithms were proposed with artificial neural networks (ANN) having become popular over the last decades. Although some studies have compared predictive accuracies of LR and ANN models, some concerns regarding their methodological quality have been voiced. Our comparison has the advantage of being based on two large independent data sets allowing for elaborate model development and independent validation. METHODS: From the German Stroke Database, a learning data set including 1754 prospectively recruited patients with acute ischemic stroke was used. Utilizing LR and ANN, two prognostic models were developed predicting restitution of functional independence and survival after 100 days. The resulting models were applied to classify 1470 patients with acute ischemic stroke; this test data set was collected independently from the learning data. Error fractions in the test data were determined, and differences in error fractions between the algorithms were calculated with 95% confidence intervals. RESULTS: For most prognostic models, error fractions in the test data were below 40%. There was no difference between the algorithms except for the model predicting completely versus incompletely restituted or deceased patients (difference in error fractions = 4.01% [2.10-5.96%], p = 0.0001). CONCLUSIONS: The conscientiously applied LR remains the gold standard for prognostic modelling; however, ANN can be an alternative automated "quick and easy" multivariate analysis.

Aged↗

LINC01871-Mediated Sensitivity to Cyclin-Dependent Kinase 4/6 Inhibitors in Human Breast Cancer.

Breast cancer remains the most frequently diagnosed malignancy in women, and resistance to cyclin-dependent kinase 4 and 6 (CDK4/6) inhibitors limits long-term treatment efficacy. This study aimed to identify long non-coding RNAs (lncRNAs) associated with predicted sensitivity to CDK4/6 inhibitors and to investigate their biological functions in breast cancer. Transcriptomic data from The Cancer Genome Atlas (TCGA) and drug sensitivity data from the Genomics of Drug Sensitivity in Cancer 2 (GDSC2) database were integrated, and drug sensitivity was predicted using the oncoPredict algorithm. Candidate lncRNAs were identified through differential expression analysis, weighted gene co-expression network analysis, prognostic analysis, and machine learning. The biological functions of LINC01871 were subsequently evaluated using in vitro and in vivo experiments. Sixty-two lncRNAs associated with predicted sensitivity to ribociclib and palbociclib were identified, and six core lncRNAs were selected. LINC01871 showed the highest discriminatory performance for predicted drug sensitivity. Overexpression of LINC01871 was associated with increased sensitivity of breast cancer cells to ribociclib and palbociclib, inhibition of cell proliferation, promotion of apoptosis, and suppression of nuclear factor kappa B (NF-κB) signaling. Single-cell transcriptomic analysis demonstrated high LINC01871 expression in T cells and natural killer (NK) cells, while transcriptome-based immune infiltration analyses showed that high LINC01871 expression was associated with increased immune infiltration. These findings identify LINC01871 as a candidate biomarker of sensitivity to CDK4/6 inhibitors and demonstrate its tumor-suppressive effects in breast cancer. Further clinical and mechanistic studies are required to validate its predictive value and therapeutic relevance.

Humans↗

Data mining and machine learning techniques for the identification of mutagenicity inducing substructures and structure activity relationships of noncongeneric compounds.

This paper explores the utility of data mining and machine learning algorithms for the induction of mutagenicity structure-activity relationships (SARs) from noncongeneric data sets. We compare (i) a newly developed algorithm (MOLFEA) for the generation of descriptors (molecular fragments) for noncongeneric compounds with traditional SAR approaches (molecular properties) and (ii) different machine learning algorithms for the induction of SARs from these descriptors. In addition we investigate the optimal parameter settings for these programs and give an exemplary interpretation of the derived models. The predictive accuracies of models using MOLFEA derived descriptors is approximately 10-15%age points higher than those using molecular properties alone. Using both types of descriptors together does not improve the derived models. From the applied machine learning techniques the rule learner PART and support vector machines gave the best results, although the differences between the learning algorithms are only marginal. We were able to achieve predictive accuracies up to 78% for 10-fold cross-validation. The resulting models are relatively easy to interpret and usable for predictive as well as for explanatory purposes.

Algorithms↗

BhairPred: prediction of beta-hairpins in a protein from multiple alignment information using ANN and SVM techniques.

This paper describes a method for predicting a supersecondary structural motif, beta-hairpins, in a protein sequence. The method was trained and tested on a set of 5102 hairpins and 5131 non-hairpins, obtained from a non-redundant dataset of 2880 proteins using the DSSP and PROMOTIF programs. Two machine-learning techniques, an artificial neural network (ANN) and a support vector machine (SVM), were used to predict beta-hairpins. An accuracy of 65.5% was achieved using ANN when an amino acid sequence was used as the input. The accuracy improved from 65.5 to 69.1% when evolutionary information (PSI-BLAST profile), observed secondary structure and surface accessibility were used as the inputs. The accuracy of the method further improved from 69.1 to 79.2% when the SVM was used for classification instead of the ANN. The performances of the methods developed were assessed in a test case, where predicted secondary structure and surface accessibility were used instead of the observed structure. The highest accuracy achieved by the SVM based method in the test case was 77.9%. A maximum accuracy of 71.1% with Matthew's correlation coefficient of 0.41 in the test case was obtained on a dataset previously used by X. Cruz, E. G. Hutchinson, A. Shephard and J. M. Thornton (2002) Proc. Natl Acad. Sci. USA, 99, 11157-11162. The performance of the method was also evaluated on proteins used in the '6th community-wide experiment on the critical assessment of techniques for protein structure prediction (CASP6)'. Based on the algorithm described, a web server, BhairPred (http://www.imtech.res.in/raghava/bhairpred/), has been developed, which can be used to predict beta-hairpins in a protein using the SVM approach.

Algorithms↗

Potential assessment of the "support vector machine" method in forecasting ambient air pollutant trends.

Monitoring and forecasting of air quality parameters are popular and important topics of atmospheric and environmental research today due to the health impact caused by exposing to air pollutants existing in urban air. The accurate models for air pollutant prediction are needed because such models would allow forecasting and diagnosing potential compliance or non-compliance in both short- and long-term aspects. Artificial neural networks (ANN) are regarded as reliable and cost-effective method to achieve such tasks and have produced some promising results to date. Although ANN has addressed more attentions to environmental researchers, its inherent drawbacks, e.g., local minima, over-fitting training, poor generalization performance, determination of the appropriate network architecture, etc., impede the practical application of ANN. Support vector machine (SVM), a novel type of learning machine based on statistical learning theory, can be used for regression and time series prediction and have been reported to perform well by some promising results. The work presented in this paper aims to examine the feasibility of applying SVM to predict air pollutant levels in advancing time series based on the monitored air pollutant database in Hong Kong downtown area. At the same time, the functional characteristics of SVM are investigated in the study. The experimental comparisons between the SVM model and the classical radial basis function (RBF) network demonstrate that the SVM is superior to the conventional RBF network in predicting air quality parameters with different time series and of better generalization performance than the RBF model.

Air Pollutants↗

Predicting rRNA-, RNA-, and DNA-binding proteins from primary structure with support vector machines.

In the post-genome era, the prediction of protein function is one of the most demanding tasks in the study of bioinformatics. Machine learning methods, such as the support vector machines (SVMs), greatly help to improve the classification of protein function. In this work, we integrated SVMs, protein sequence amino acid composition, and associated physicochemical properties into the study of nucleic-acid-binding proteins prediction. We developed the binary classifications for rRNA-, RNA-, DNA-binding proteins that play an important role in the control of many cell processes. Each SVM predicts whether a protein belongs to rRNA-, RNA-, or DNA-binding protein class. Self-consistency and jackknife tests were performed on the protein data sets in which the sequences identity was < 25%. Test results show that the accuracies of rRNA-, RNA-, DNA-binding SVMs predictions are approximately 84%, approximately 78%, approximately 72%, respectively. The predictions were also performed on the ambiguous and negative data set. The results demonstrate that the predicted scores of proteins in the ambiguous data set by RNA- and DNA-binding SVM models were distributed around zero, while most proteins in the negative data set were predicted as negative scores by all three SVMs. The score distributions agree well with the prior knowledge of those proteins and show the effectiveness of sequence associated physicochemical properties in the protein function prediction. The software is available from the author upon request.

Amino Acid Sequence↗

Toward a Better Paradigm for Head and Neck Cancer Treatment Applying AI (HNC-TACTIC): Protocol for an International Cohort Study of Electronic Health Records.

BACKGROUND: Head and neck squamous cell carcinomas (HNSCCs) cause considerable morbidity and mortality. Multimodal treatment strategies can cause significant toxicity, and therapy options are limited for recurrent disease. Immunotherapy has emerged as a promising approach. However, patient response variability underscores the need for better predictive markers. OBJECTIVE: This study aims to use artificial intelligence to develop two predictive models in patients with HNSCC to assess (1) progression or recurrence following primary curative treatment and (2) long-term survival after immunotherapy schemes in recurrent and metastatic disease. This study will also describe the characteristics of patients with early, locally advanced, and recurrent or metastatic cancers. METHODS: This is a retrospective, observational study of data captured in electronic health records (EHRs) from participating hospitals between January 1, 2014, and December 31, 2021. This study's population comprises adults diagnosed with HNSCC at any stage. Study variables, including demographics, comorbidities, clinical variables, treatments, and outcomes, will be extracted using EHRead, a technology that applies natural language processing and machine learning to extract and analyze structured and unstructured clinical information in deidentified EHRs. Predictive models based on dynamic risk stratification for treatment response and progression or recurrence will be developed using multivariable logistic regressions, decision tree classifiers, and random forest approaches. Descriptive and outcome analyses will be shown for different anatomic subsites and stratified by stage and treatment. RESULTS: This study began enrolling sites in July 2021 and is currently ongoing. By December 2025, data from 10 centers has been collected, comprising a total of 151,934,990 EHRs from 2,159,719 patients. CONCLUSIONS: Development of predictive models using artificial intelligence will advance clinical understanding of HNSCC to improve patient outcomes.

Humans↗

Models to predict cardiovascular risk: comparison of CART, multilayer perceptron and logistic regression.

The estimate of a multivariate risk is now required in guidelines for cardiovascular prevention. Limitations of existing statistical risk models lead to explore machine-learning methods. This study evaluates the implementation and performance of a decision tree (CART) and a multilayer perceptron (MLP) to predict cardiovascular risk from real data. The study population was randomly splitted in a learning set (n = 10,296) and a test set (n = 5,148). CART and the MLP were implemented at their best performance on the learning set and applied on the test set and compared to a logistic model. Implementation, explicative and discriminative performance criteria are considered, based on ROC analysis. Areas under ROC curves and their 95% confidence interval are 0.78 (0.75-0.81), 0.78 (0.75-0.80) and 0.76 (0.73-0.79) respectively for logistic regression, MLP and CART. Given their implementation and explicative characteristics, these methods can complement existing statistical models and contribute to the interpretation of risk.

Artificial Intelligence↗

Method of predicting splice sites based on signal interactions.

BACKGROUND: Predicting and proper ranking of canonical splice sites (SSs) is a challenging problem in bioinformatics and machine learning communities. Any progress in SSs recognition will lead to better understanding of splicing mechanism. We introduce several new approaches of combining a priori knowledge for improved SS detection. First, we design our new Bayesian SS sensor based on oligonucleotide counting. To further enhance prediction quality, we applied our new de novo motif detection tool MHMMotif to intronic ends and exons. We combine elements found with sensor information using Naive Bayesian Network, as implemented in our new tool SpliceScan. RESULTS: According to our tests, the Bayesian sensor outperforms the contemporary Maximum Entropy sensor for 5' SS detection. We report a number of putative Exonic (ESE) and Intronic (ISE) Splicing Enhancers found by MHMMotif tool. T-test statistics on mouse/rat intronic alignments indicates, that detected elements are on average more conserved as compared to other oligos, which supports our assumption of their functional importance. The tool has been shown to outperform the SpliceView, GeneSplicer, NNSplice, Genio and NetUTR tools for the test set of human genes. SpliceScan outperforms all contemporary ab initio gene structural prediction tools on the set of 5' UTR gene fragments. CONCLUSION: Designed methods have many attractive properties, compared to existing approaches. Bayesian sensor, MHMMotif program and SpliceScan tools are freely available on our web site. REVIEWERS: This article was reviewed by Manyuan Long, Arcady Mushegian and Mikhail Gelfand.

Journal Article↗