PubMed HealthSearch

Biomedical subjects

G D Stormo

Publications and source records attributed to G D Stormo.

11 recordsLinked to original sources

Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods.

Comparative sequence analysis addresses the problem of RNA folding and RNA structural diversity, and is responsible for determining the folding of many RNA molecules, including 5S, 16S, and 23S rRNAs, tRNA, RNAse P RNA, and Group I and II introns. Initially this method was utilized to fold these sequences into their secondary structures. More recently, this method has revealed numerous tertiary correlations, elucidating novel RNA structural motifs, several of which have been experimentally tested and verified, substantiating the general application of this approach. As successful as the comparative methods have been in elucidating higher-order structure, it is clear that additional structure constraints remain to be found. Deciphering such constraints requires more sensitive and rigorous protocols, in addition to RNA sequence datasets that contain additional phylogenetic diversity and an overall increase in the number of sequences. Various RNA databases, including the tRNA and rRNA sequence datasets, continue to grow in number as well as diversity. Described herein is the development of more rigorous comparative analysis protocols. Our initial development and applications on different RNA datasets have been very encouraging. Such analyses on tRNA, 16S and 23S rRNA are substantiating previously proposed associations and are now beginning to reveal additional constraints on these molecules. A subset of these involve several positions that correlate simultaneously with one another, implying units larger than a basepair can be under a phylogenetic constraint.

Base Sequence

Splicing signals in Drosophila: intron size, information content, and consensus sequences.

A database of 209 Drosophila introns was extracted from Genbank (release number 64.0) and examined by a number of methods in order to characterize features that might serve as signals for messenger RNA splicing. A tight distribution of sizes was observed: while the smallest introns in the database are 51 nucleotides, more than half are less than 80 nucleotides in length, and most of these have lengths in the range of 59-67 nucleotides. Drosophila splice sites found in large and small introns differ in only minor ways from each other and from those found in vertebrate introns. However, larger introns have greater pyrimidine-richness in the region between 11 and 21 nucleotides upstream of 3' splice sites. The Drosophila branchpoint consensus matrix resembles C T A A T (in which branch formation occurs at the underlined A), and differs from the corresponding mammalian signal in the absence of G at the position immediately preceding the branchpoint. The distribution of occurrences of this sequence suggests a minimum distance between 5' splice sites and branchpoints of about 38 nucleotides, and a minimum distance between 3' splice sites and branchpoints of 15 nucleotides. The methods we have used detect no information in exon sequences other than in the few nucleotides immediately adjacent to the splice sites. However, Drosophila resembles many other species in that there is a discontinuity in A + T content between exons and introns, which are A + T rich.

Animals

Expectation maximization algorithm for identifying protein-binding sites with variable lengths from unaligned DNA fragments.

An Expectation Maximization algorithm for identification of DNA binding sites is presented. The approach predicts the location of binding regions while allowing variable length spacers within the sites. In addition to predicting the most likely spacer length for a set of DNA fragments, the method identifies individual sites that differ in spacer size. No alignment of DNA sequences is necessary. The method is illustrated by application to 231 Escherichia coli DNA fragments known to contain promoters with variable spacings between their consensus regions. Maximum-likelihood tests of the differences between the spacing classes indicate that the consensus regions of the spacing classes are not distinct. Further tests suggest that several positions within the spacing region may contribute to promoter specificity.

Algorithms

Translation initiation in Escherichia coli: sequences within the ribosome-binding site.

The translational roles of the Shine-Dalgarno sequence, the initiation codon, the space between them, and the second codon have been studied. The Shine-Dalgarno sequence UAAGGAGG initiated translation roughly four times more efficiently than did the shorter AAGGA sequence. Each Shine-Dalgarno sequence required a minimum distance to the initiation codon in order to drive translation; spacing, however, could be rather long. Initiation at AUG was more efficient than at GUG or UUG at each spacing examined; initiation at GUG was only slightly better than UUG. Translation was also affected by residues 3' to the initiation codon. The second codon can influence the rate of initiation, with the magnitude depending on the initiation codon. The data are consistent with a simple kinetic model in which a variety of rate constants contribute to the process of translation initiation.

Base Sequence

Specificity of the Mnt protein determined by binding to randomized operators.

The relative binding affinities of Mnt protein from bacteriophage P22 are determined for each possible base pair at position 17 of the operator. These are determined from the partitioning of randomized operators into bound and unbound fractions; quantitation is provided by restriction enzyme analysis. Mnt protein is found to have an unusual specificity at this position: a C.G base pair (the wild-type operator) has the highest affinity, a G.C base pair has the lowest affinity, and both orientations of A.T base pairs are intermediate and nearly equivalent. A specific binding constant and specific binding free energy are defined and shown to be directly related to the information content of the operator sequences bound to the protein, taking into account the quantitative differences in binding affinities.

Bacteriophages

Probing information content of DNA-binding sites.

An information content analysis of protein-binding sites gives a quantitative description of the specificity of the protein, independent of the mechanism of specificity. It gives useful information about the total specificity of the protein and about the individual positions within the binding sites. Information content is consistent with both thermodynamic and statistical analyses of specificity. When applied to a collection of known binding sites, the description provided may be limited by the sample size or by unknown constraints on those sites. Experimental procedures to determine the information content can give much more reliable measures. A large number of functional sites can be obtained from a much larger pool of randomized potential sites. Quantitative assays for the activity of different sites can be easily incorporated into the analysis, thereby increasing its sensitivity. Both in vitro and in vivo experiments are amenable to information content analysis.

Base Sequence

Automated kinetic assay of beta-galactosidase activity.

An automated kinetic assay for beta-galactosidase activity in Escherichia coli was developed to permit the measurement of many independent samples simultaneously. Bacteria are grown, lysed from without (by adsorption of a high multiplicity of bacteriophage T4) and assayed in microtiter plates with 96 wells. Absorbance data are collected and analyzed by computer. The growth and lysis procedure, apparatus and software used in this assay can be used for other spectrophotometric enzyme assays.

Automation

Sequence requirements of the hammerhead RNA self-cleavage reaction.

A previously well-characterized hammerhead catalytic RNA consisting of a 24-nucleotide substrate and a 19-nucleotide ribozyme was used to perform an extensive mutagenesis study. The cleavage rates of 21 different substrate mutations and 24 different ribozyme mutations were determined. Only one of the three phylogenetically conserved base pairs but all nine of the conserved single-stranded residues in the central core are needed for self cleavage. In most cases the mutations did not alter the ability of the hammerhead to assemble into a bimolecular complex. In the few cases where mutant hammerheads did not assemble, it appeared to be the result of the mutation stabilizing an alternate substrate or ribozyme secondary structure. All combinations of mutant substrate and mutant ribozyme were less active than the corresponding single mutations, suggesting that the hammerhead contains few, if any, replaceable tertiary interactions as are found in tRNA. The refined consensus hammerhead resulting from this work was used to identify potential hammerheads present in a variety of Escherichia coli gene sequences.

Base Sequence

Consensus patterns in DNA.

Matrices can provide realistic representations of protein/DNA specificity. In many cases simple mononucleotide-based matrices are adequate representations, but more complex matrices may be needed for other cases. Unlike simple consensus sequences, matrices allow for different penalties to be assessed for different changes to a binding site, a property that is essential for accurate description of a binding site pattern. When only a collection of binding site sequences is known, the best representation for the pattern is an information content formulation, based on both thermodynamic and statistical considerations. Quantitative data on relative binding affinities may be used to determine matrices that provide a best fit to the data. Matrix representations also provide an efficient method of aligning multiple sequences to identify binding site patterns that they have in common.

Base Sequence

Identification of consensus patterns in unaligned DNA sequences known to be functionally related.

We have developed a method for identifying consensus patterns in a set of unaligned DNA sequences known to bind a common protein or to have some other common biochemical function. The method is based on a matrix representation of binding site patterns. Each row of the matrix represents one of the four possible bases, each column represents one of the positions of the binding site and each element is determined by the frequency the indicated base occurs at the indicated position. The goal of the method is to find the most significant matrix--i.e. the one with the lowest probability of occurring by chance--out of all the matrices that can be formed from the set of related sequences. The reliability of the method improves with the number of sequences, while the time required increases only linearly with the number of sequences. To test this method, we analysed 11 DNA sequences containing promoters regulated by the Escherichia coli LexA protein. The matrices we found were consistent with the known consensus sequence, and could distinguish the generally accepted LexA binding sites from other DNA sequences.

Algorithms