Report of the ninth international workshop on the identification of transcribed sequences.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to R Mural.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Detection of RNA polymerase II promoters and polyadenylation sites helps to locate gene boundaries and can enhance accurate gene recognition and modeling in genomic DNA sequence. We describe a system which can be used to detect polyadenylation sites and thus delineate the 3' boundary of a gene, and discuss improvements to a system first described in Matis et al. (1995) [Matis S., Shah M., Mural R. J. & Uberbacher E.C. (1995) Proc. First Wrld Conf. Computat. Med., Public Hlth, Biotechnol. (Wrld Sci.) (in press).], which predicts a large subset of RNA polymerase II promoters. The promoter system used statistical matrices and distance information as inputs for a neural network which was trained to provide initial promoter recognition. The output of the network was further refined by applying rules which use the gene context information predicted by GRAIL. We have reconstructed the rule-based system which uses gene context information and significantly improved the sensitivity and selectivity of promoter detection.
We have described an improved neural network system for recognizing protein coding regions (exons) in human genomic DNA sequences. This coding region recognition system is part of a new version of GRAIL, GRAIL II, and represents a significant improvement over the coding recognition performance of the previous GRAIL system. GRAIL II divides the process of locating exons into four steps. It first generates an exon candidate pool consisting of all possible (translation start-donor), (acceptor-donor), and (acceptor-translation stop) pairs within all open reading frames of the test sequence. The vast majority of these exon candidates are eliminated from consideration by applying a set of heuristic rules. After reducing the size of the candidate pool, GRAIL II uses three trained neural networks to evaluate the coding potential and accuracy of the edges of starting exon, internal exon and terminal exon candidates. These networks output a set of overlapping candidates for each exon which differ by their scores and position of their edges. Multiple candidates for a given exon are grouped into a cluster based on their locations relative to candidates corresponding to other exons, and the highest scoring candidate for each cluster is used as the "best" prediction of the corresponding exon. Unlike the previous GRAIL version, GRAIL II uses variable-length windows to evaluate exon candidates and its performance is nearly independent of exon length. In addition to several strong indicators of coding potential, the system uses several other types of information including scores for splice junctions, GC composition, and the properties of the regions adjacent to an exon candidate, to aid in the discrimination process. On a large set of sequences from Genbank (3), GRAIL II located 93% of all exons regardless of size with a false positive rate of 12%. Among the true positives, 62% match the actual exons exactly (the exons edges are correct to the base), and 93% match at least one edge correctly. These statistics are further improved, especially the false positive rate and accuracy of the edges, through a process of gene model construction by the Gene Assembly Program (GAP III) (4) module of GRAIL II, which uses the scored exon candidates as input and constructs optimal gene models. The gene modeling system will be described elsewhere.
Genetic evidence indicates that Oxys-6, an oxygen-sensitive mutant of Escherichia coli AB1157, is defective in the region of the hemB locus. Oxys-6 is capable of growth under aerobic conditions only if cultures are initiated at low-inoculum levels. Aerobic liquid cultures are limited to a cell density of 10(7) cells per ml by the accumulation of a metabolically produced, low-molecular-weight, heat-stable material in complex organic media. Both Oxys-6 and AB1157 cells produce the material, but only aerobic cultures of the mutant are inhibited by it. The material is produced by both intact cells and cell extracts in complex media. This reaction also occurs when the amino acid L-lysine is substituted for complex media.
The structure of the endogenous murine leukemia virus (MuLV) sequences of NIH/Swiss mice was analyzed by restriction endonuclease digestion, gel electrophoresis, and hybridization to an MuLV nucleic acid probe. Digestion of mouse DNA with certain restriction endonucleases revealed two classes of fragments. A large number of fragments (about 30) were present at a relatively low concentration, indicating that each derived from a sequence present once in the mouse genome. A smaller number of fragments (one to five) were present at a much higher concentration and must have resulted from sequences present multiple times in the mouse genome. These results indicated that the endogenous MuLV sequences represent a family of dispersed repetitive sequences. Hybridization of these same digested mouse DNAs to nucleic acid probes representing different portions of the MuLV genome allowed construction of a map of the sites where restriction endonucleases cleave the endogenous MuLV sequences. Several independent recombinant DNA clones of endogenous MuLV sequences have been isolated from C3H mice (Roblin et al., J. Virol. 43:113-126, 1982). Analysis of these sequences shows that they have the structure of MuLV proviruses. The sites at which restriction endonucleases cleave within these proviruses appeared to be similar or identical to the sites at which these nucleases cleaved within the MuLV sequences of NIH/Swiss mice. This identity was confirmed by parallel electrophoresis. We conclude that the apparently complex pattern of endogenous MuLV sequences of NIH/Swiss mice consists largely of only two kinds of provirus, each repeated multiple times at dispersed sites in the mouse genome.