PubMed HealthSearch

Biomedical subjects

T D Schneider

Publications and source records attributed to T D Schneider.

18 recordsLinked to original sources

Using information content and base frequencies to distinguish mutations from genetic polymorphisms in splice junction recognition sites.

Predicting the effects of nucleotide substitutions in human splice sites has been based on analysis of consensus sequences. We used a graphic representation of sequence conservation and base frequency, the sequence logo, to demonstrate that a change in a splice acceptor of hMSH2 (a gene associated with familial nonpolyposis colon cancer) probably does not reduce splicing efficiency. This confirms a population genetic study that suggested that this substitution is a genetic polymorphism. The information theory-based sequence logo is quantitative and more sensitive than the corresponding splice acceptor consensus sequence for detection of true mutations. Information analysis may potentially be used to distinguish polymorphisms from mutations in other types of transcriptional, translational, or protein-coding motifs.

Base Sequence

Redox-dependent shift of OxyR-DNA contacts along an extended DNA-binding site: a mechanism for differential promoter selection.

The redox-sensitive OxyR protein activates the transcription of antioxidant defense genes in response to oxidative stress and represses its own expression under both oxidizing and reducing conditions. Previous studies showed that OxyR-binding sites are unusually long with limited sequence similarity. Here, we report that oxidized OxyR recognizes a motif comprised of four ATAGnt elements spaced at 10 bp intervals and contacts these elements in four adjacent major grooves on one face of the DNA helix. In contrast, reduced OxyR contacts two pairs of adjacent major grooves separated by one helical turn. The two modes of binding are essential for OxyR to function as both an activator and a repressor in vivo. We propose that specific DNA recognition by an OxyR tetramer is achieved with four contacts of intermediate affinity allowing OxyR to reposition its DNA contacts and target alternate sets of promoters as the cellular redox state is altered.

Adaptation, Biological

Quantitative analysis of ribosome binding sites in E.coli.

185 clones with randomized ribosome binding sites, from position -11 to 0 preceding the coding region of beta-galactosidase, were selected and sequenced. The translational yield of each clone was determined; they varied by more than 3000-fold. Multiple linear regression analysis was used to determine the contribution to translation initiation activity of each base at each position. Features known to be important for translation initiation, such as the initiation codon, the Shine/Dalgarno sequence, the identity of the base at position -3 and the occurrence of alternative ATGs, are all found to be important quantitatively for activity. No other features are found to be of general significance, although the effects of secondary structure can be seen as outliers. A comparison to a large number of natural E.coli translation initiation sites shows the information profile to be qualitatively similar although differing quantitatively. This is probably due to the selection for good translation initiation sites in the natural set compared to the low average activity of the randomized set.

Base Sequence

Information analysis of sequences that bind the replication initiator RepA.

The replication initiator protein RepA of plasmid P1 can bind to 14 sites on the plasmid. These sites are variously used to autoregulate RepA synthesis and for initiation and control of DNA replication. Analysis of information (degree of conservation) at the sites revealed three sequence patches of high conservation. By saturation mutagenesis, the conservation at the outer two patches was found to contribute to RepA binding more critically. The guanine bases that are likely to contact RepA through the major groove were identified by methylation interference and methylation protection experiments. These bases mapped to the outer two patches and were separated by one turn of the helix. Therefore, they belong to major grooves on the same face of DNA. All backbone contacts of the protein, determined by hydroxyl radical footprinting, also mapped to the same face. We conclude from this that RepA binds to its site on one face of the DNA. Information analysis of binding sites for several prokaryotic repressors and activators, where the nature of DNA-protein contacts are known, revealed a correlation between the positions of high conservation and the positions of major grooves that faced the protein. The middle patch of high conservation in the RepA binding sites is an exception since in this region a minor groove is likely to face the protein. The simplest model for minor groove contacts suggests that in B-form DNA a T.A base-pair cannot easily be distinguished from an A.T pair by inspection of the minor groove. Yet in the RepA site, a T-->A mutation in the middle patch significantly affects binding. Therefore, the simplest models for both minor and major groove contacts are unlikely. It is possible that the patch determines the proper conformation of the site and thereby contributes to recognition indirectly.

Bacterial Proteins

Features of spliceosome evolution and function inferred from an analysis of the information at human splice sites.

An information analysis of the 5' (donor) and 3' (acceptor) sequences spanning the ends of nearly 1800 human introns has provided evidence for structural features of splice sites that bear upon spliceosome evolution and function: (1) 82% of the sequence information (i.e. sequence conservation) at donor junctions and 97% of the sequence information at acceptor junctions is confined to the introns, allowing codon choices throughout exons to be largely unrestricted. The distribution of information at intron-exon junctions is also described in detail and compared with footprints. (2) Acceptor sites are found to possess enough information to be located in the transcribed portion of the human genome, whereas donor sites possess about one bit less than the information needed to locate them independently. This difference suggests that acceptor sites are located first in humans and, having been located, reduce by a factor of two the number of alternative sites available as donors. Direct experimental evidence exists to support this conclusion. (3) The sequences of donor and acceptor splice sites exhibit a striking similarity. This suggests that the two junctions derive from a common ancestor and that during evolution the information of both sites shifted onto the intron. If so, the protein and RNA components that are found in contemporary spliceosomes, and which are responsible for recognizing donor and acceptor sequences, should also be related. This conclusion is supported by the common structures found in different parts of the spliceosome.

Base Sequence

High information conservation implies that at least three proteins bind independently to F plasmid incD repeats.

The 12 incD repeats in the F plasmid each contain about 60 bits of information, which is three times the amount of conservation that a single protein would need to distinguish the repeats from the rest of the Escherichia coli genome. This is the first reported discovery of a case of threefold excess information, and it implies that at least three proteins bind independently to the repeats. In support of this observation, other workers have shown that three polypeptides bind to this region, but only one, SopB, is known to bind independently of other factors. Identification of the other two proteins should help us to understand the mechanism of plasmid partitioning during cell division.

Base Sequence

Theory of molecular machines. I. Channel capacity of molecular machines.

Like macroscopic machines, molecular-sized machines are limited by their material components, their design, and their use of power. One of these limits is the maximum number of states that a machine can choose from. The logarithm to the base 2 of the number of states is defined to be the number of bits of information that the machine could "gain" during its operation. The maximum possible information gain is a function of the energy that a molecular machine dissipates into the surrounding medium (Py), the thermal noise energy which disturbs the machine (Ny) and the number of independently moving parts involved in the operation (dspace): Cy = dspace log2 [( Py + Ny)/Ny] bits per operation. This "machine capacity" is closely related to Shannon's channel capacity for communications systems. An important theorem that Shannon proved for communication channels also applies to molecular machines. With regard to molecular machines, the theorem states that if the amount of information which a machine gains is less than or equal to Cy, then the error rate (frequency of failure) can be made arbitrarily small by using a sufficiently complex coding of the molecular machine's operation. Thus, the capacity of a molecular machine is sharply limited by the dissipation and the thermal noise, but the machine failure rate can be reduced to whatever low level may be required for the organism to survive.

Animals

Theory of molecular machines. II. Energy dissipation from molecular machines.

Single molecules perform a variety of tasks in cells, from replicating, controlling and translating the genetic material to sensing the outside environment. These operations all require that specific actions take place. In a sense, each molecule must make tiny decisions. To make a decision, each "molecular machine" must dissipate an energy Py in the presence of thermal noise Ny. The number of binary decisions that can be made by a machine which has dspace independently moving parts is the "machine capacity" Cy = dspace log2 [(Py + Ny)/Ny]. This formula is closely related to Shannon's channel capacity for communications systems, C = W log2 [(P + N)/N]. This paper shows that the minimum amount of energy that a molecular machine must dissipate in order to gain one bit of information is epsilon min = kB T ln (2) joules/bit. This equation is derived in two distinct ways. The first derivation begins with the Second Law of Thermodynamics, which shows that the statement that there is a minimum energy dissipation is a restatement of the Second Law of Thermodynamics. The second derivation begins with the machine capacity formula, which shows that the machine capacity is also related to the Second Law of Thermodynamics. One of Shannon's theorems for communications channels is that as long as the channel capacity is not exceeded, the error rate may be made as small as desired by a sufficiently involved coding. This result also applies to the dissipation formula for molecular machines. So there is a precise upper bound on the number of choices a molecular machine can make for a given amount of energy loss. This result will be important for the design and construction of molecular computers.

Animals

Automated kinetic assay of beta-galactosidase activity.

An automated kinetic assay for beta-galactosidase activity in Escherichia coli was developed to permit the measurement of many independent samples simultaneously. Bacteria are grown, lysed from without (by adsorption of a high multiplicity of bacteriophage T4) and assayed in microtiter plates with 96 wells. Absorbance data are collected and analyzed by computer. The growth and lysis procedure, apparatus and software used in this assay can be used for other spectrophotometric enzyme assays.

Automation

Sequence logos: a new way to display consensus sequences.

A graphical method is presented for displaying the patterns in a set of aligned sequences. The characters representing the sequence are stacked on top of each other for each position in the aligned sequences. The height of each letter is made proportional to its frequency, and the letters are sorted so the most common one is on top. The height of the entire stack is then adjusted to signify the information content of the sequences at that position. From these 'sequence logos', one can determine not only the consensus sequence but also the relative frequency of bases and the information content (measured in bits) at every position in a site or sequence. The logo displays both significant residues and subtle sequence patterns.

Amino Acid Sequence

Excess information at bacteriophage T7 genomic promoters detected by a random cloning technique.

In our previous analysis of the information at binding sites on nucleic acids, we found that most of the sites examined contain the amount of information expected from their frequency in the genome. The sequences at bacteriophage T7 promoters are an exception, because they are far more conserved (35 bits of information content) than should be necessary to distinguish them from the background of the Escherichia coli genome (17 bits). To determine the information actually used by the T7 RNA polymerase, promoters were chemically synthesized with many variations and those that function well in an in vivo assay were sequenced. Our analysis shows that the polymerase uses 18 bits of information, so the sequences at phage genomic promoters have significantly more information than the polymerase needs. The excess may represent the binding site of another protein.

Base Composition

Quantitative analysis of the relationship between nucleotide sequence and functional activity.

Matrices can be used to evaluate sequences for functional activity. Multiple regression can solve for the matrix that gives the best fit between sequence evaluations and quantitative activities. This analysis shows that the best model for context effects on suppression by su2 involves primarily the two nucleotides 3' to the amber codon, and that their contributions are independent and additive. Context effects on 2AP mutagenesis also involve the two nucleotides 3' to the 2AP insertion, but their effects are not independent. In a construct for producing beta-galactosidase, the effects on translational yields of the tri-nucleotide 5' to the initiation codon are dependent on the entire triplet. Models based on these quantitative results are presented for each of the examples.

Base Sequence

Information content of binding sites on nucleotide sequences.

Repressors, polymerases, ribosomes and other macromolecules bind to specific nucleic acid sequences. They can find a binding site only if the sequence has a recognizable pattern. We define a measure of the information (R sequence) in the sequence patterns at binding sites. It allows one to investigate how information is distributed across the sites and to compare one site to another. One can also calculate the amount of information (R frequency) that would be required to locate the sites, given that they occur with some frequency in the genome. Several Escherichia coli binding sites were analyzed using these two independent empirical measurements. The two amounts of information are similar for most of the sites we analyzed. In contrast, bacteriophage T7 RNA polymerase binding sites contain about twice as much information as is necessary for recognition by the T7 polymerase, suggesting that a second protein may bind at T7 promoters. The extra information can be accounted for by a strong symmetry element found at the T7 promoters. This element may be an operator. If this model is correct, these promoters and operators do not share much information. The comparisons between R sequence and R frequency suggest that the information at binding sites is just sufficient for the sites to be distinguished from the rest of the genome.

Bacterial Proteins

Sequence landscapes.

We describe a method for representing the structure of repeating sequences in nucleic-acids, proteins and other texts. A portion of the sequence is presented at the bottom of a CRT screen. Above the sequence is its landscape, which looks like a mountain range. Each mountain corresponds to a subsequence of the sequence. At the peak of every mountain is written the number of times that the subsequence appears. A data structure called a DAWG, which can be built in time proportional to the length of the sequence, is used to construct the landscape. For the 40 thousand bases of bacteriophage T7, the DAWG can be built in 30 seconds. The time to display any portion of the landscape is less than a second. Using sequence landscapes, one can quickly locate significant repeats.

DNA, Viral

Delila system tools.

We introduce three new computer programs and associated tools of the Delila nucleic-acid sequence analysis system. The first program, Module, allows rapid transportation of new sequence analysis tools between scientists using different computers. The second program, DBpull, allows efficient access to the large nucleic-acid sequence databases being collected in the United States and Europe. The third program, Encode, provides a flexible way to process sequence data for analysis by other programs.

Base Sequence

Characterization of translational initiation sites in E. coli.

We characterize the Shine and Dalgarno sequence of 124 known gene beginnings. This information is used to make "rules" which help distinguish gene beginning from other sites in a library of over 78,000 bases of mRNA. Gene beginnings are found to have information besides the initiation codon and Shine and Dalgarno sequence which can be used to make better "rules".

Base Composition

Use of the 'Perceptron' algorithm to distinguish translational initiation sites in E. coli.

We have used a "Perceptron" algorithm to find a weighting function which distinguishes E. coli translational initiation sites from all other sites in a library of over 78,000 nucleotides of mRNA sequence. The "Perceptron" examined sequences as linear representations. The "Perceptron" is more successful at finding gene beginnings than our previous searches using "rules" (see previous paper). We note that the weighting function can find translational initiation sites within sequences that were not included in the training set.

Base Sequence