PubMed Health⌕ Search

Biomedical subjects

E Uberbacher

Publications and source records attributed to E Uberbacher.

5 recordsLinked to original sources

Protein threading by PROSPECT: a prediction experiment in CASP3.

We present an analysis of the protein fold recognition experiment using PROSPECT in The Third Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP3). PROSPECT is a computer program we have recently developed for finding an optimal alignment between a protein sequence and a protein structural fold. Two unique features of PROSPECT are (a) that it guarantees to find the globally optimal sequence-structure alignment and does so in an efficient manner, when the alignment-scoring function consists of three additive terms: (i) a singleton fitness term, (ii) a pairwise contact preference term between residues that are spatially close (</=15 A between their beta-carbons) and (iii) an alignment gap penalty; and (b) that it guarantees to find the globally-optimal alignment under various constraints on the unknown protein specified by the user. In the CASP3 experiment, PROSPECT correctly identified the most similar folds for 11 targets and predicted closely-similar folds for five other targets among the 23 targets which can be classified into the category of fold-recognition problems and also had their experimentally-determined structures available. Among the 11 correctly identified folds, PROSPECT obtained good sequence-structure alignments for nine of them. On three of the five ab initio prediction problems, PROSPECT successfully located partial structures from our template library, which align accurately with the corresponding targets.

Amino Acid Sequence↗

Detection of RNA polymerase II promoters and polyadenylation sites in human DNA sequence.

Detection of RNA polymerase II promoters and polyadenylation sites helps to locate gene boundaries and can enhance accurate gene recognition and modeling in genomic DNA sequence. We describe a system which can be used to detect polyadenylation sites and thus delineate the 3' boundary of a gene, and discuss improvements to a system first described in Matis et al. (1995) [Matis S., Shah M., Mural R. J. & Uberbacher E.C. (1995) Proc. First Wrld Conf. Computat. Med., Public Hlth, Biotechnol. (Wrld Sci.) (in press).], which predicts a large subset of RNA polymerase II promoters. The promoter system used statistical matrices and distance information as inputs for a neural network which was trained to provide initial promoter recognition. The output of the network was further refined by applying rules which use the gene context information predicted by GRAIL. We have reconstructed the rule-based system which uses gene context information and significantly improved the sensitivity and selectivity of promoter detection.

Algorithms↗

Discovering the intelligence in molecular biology.

The Third International Conference on Intelligent Systems in Molecular Biology was truly an outstanding event. Computational methods in molecular biology have reached a new level of maturity and utility, resulting in many high-impact applications. The success of this meeting bodes well for the rapid and continuing development of computational methods, intelligent systems and information-based approaches for the biosciences. The basic technology, originally most often applied to 'feasibility' problems, is now dealing effectively with the most difficult real-world problems. Significant progress has been made in understanding protein-structure information, structural classification, and how functional information and the relevant features of active-site geometry can be gleaned from structures by automated computational approaches. The value and limits of homology-based methods, and the ability to classify proteins by structure in the absence of homology, have reached a new level of sophistication. New methods for covariation analysis in the folding of large structures such as RNAs have shown remarkably good results, indicating the long-term potential to understand very complicated molecules and multimolecular complexes using computational means. Novel methods, such as HMMs, context-free grammars and the uses of mutual information theory, have taken center stage as highly valuable tools in our quest to represent and characterize biological information. A focus on creative uses of intelligent systems technologies and the trend toward biological application will undoubtedly continue and grow at the 1996 ISMB meeting in St Louis.

Algorithms↗

Recognizing exons in genomic sequence using GRAIL II.

We have described an improved neural network system for recognizing protein coding regions (exons) in human genomic DNA sequences. This coding region recognition system is part of a new version of GRAIL, GRAIL II, and represents a significant improvement over the coding recognition performance of the previous GRAIL system. GRAIL II divides the process of locating exons into four steps. It first generates an exon candidate pool consisting of all possible (translation start-donor), (acceptor-donor), and (acceptor-translation stop) pairs within all open reading frames of the test sequence. The vast majority of these exon candidates are eliminated from consideration by applying a set of heuristic rules. After reducing the size of the candidate pool, GRAIL II uses three trained neural networks to evaluate the coding potential and accuracy of the edges of starting exon, internal exon and terminal exon candidates. These networks output a set of overlapping candidates for each exon which differ by their scores and position of their edges. Multiple candidates for a given exon are grouped into a cluster based on their locations relative to candidates corresponding to other exons, and the highest scoring candidate for each cluster is used as the "best" prediction of the corresponding exon. Unlike the previous GRAIL version, GRAIL II uses variable-length windows to evaluate exon candidates and its performance is nearly independent of exon length. In addition to several strong indicators of coding potential, the system uses several other types of information including scores for splice junctions, GC composition, and the properties of the regions adjacent to an exon candidate, to aid in the discrimination process. On a large set of sequences from Genbank (3), GRAIL II located 93% of all exons regardless of size with a false positive rate of 12%. Among the true positives, 62% match the actual exons exactly (the exons edges are correct to the base), and 93% match at least one edge correctly. These statistics are further improved, especially the false positive rate and accuracy of the edges, through a process of gene model construction by the Gene Assembly Program (GAP III) (4) module of GRAIL II, which uses the scored exon candidates as input and constructs optimal gene models. The gene modeling system will be described elsewhere.

Amino Acid Sequence↗