PubMed HealthSearch

Biomedical subjects

R Nussinov

Publications and source records attributed to R Nussinov.

At least 19 recordsLinked to original sources

The eukaryotic CCAAT and TATA boxes, DNA spacer flexibility and looping.

Regulation of RNA transcription in eukaryotic polymerase II promoters involves a complex assembly of protein factors. Some of the factors bind to their cognate DNA-sequence elements while others mediate between the DNA bound ones. In order to enable protein-protein interaction, their spatial positioning with respect to each other is critical. Here two DNA-sequence-elements are investigated, the CCAAT and the TATA boxes and their spacers. Whereas the position of the TATA is fixed at about -30, that of the CCAAT can vary substantially from -50 to -200. Despite the variable loop sizes, the CTF (CCAAT-binding) protein interacts--either directly or indirectly via a co-activator--with the general basal TATA-binding transcription factors. Sequence analysis of the spacers, as a function of their sizes, reveals that in the upstream regions of the spacers RR and YY are abundant. In the downstream, 3' region of the spacers RY and YR are very frequent. The DNA sequence elements and their intervening spacers are analyzed in terms of their geometry, anisotropic flexibility and local superhelical density. Our results indicate that the CCAAT and its vicinity is rigid, whereas the TATA and its surroundings is flexible. It is the large flexibility of this region in twist and in roll which allows DNA looping. General mechanistic implications for pol II promoters are discussed.

Animals

An efficient automated computer vision based technique for detection of three dimensional structural motifs in proteins.

As the number of available three dimensional coordinates of proteins increases, it is now recognized that proteins from different families and topologies are constructed from independent motifs. Detection of specific structural motifs within proteins aids in understanding their role and the mechanism of their operation. To aid in identification and use of these motifs it has become necessary to develop efficient methods for systematic scanning of structural databases. To date, methods of structural protein comparison suffer from at least one of the following limitations: (1) are not fully automated (require human intervention), (2) are limited to relatively similar structures, (3) are constrained to linear alignments of the structures, (4) are sensitive to insertions, deletions or gaps in the sequences or (5) are very time consuming. We present a method to overcome the above limitations. The method discovers and ranks every piece of structural similarity between the structures compared, thus allowing the simultaneous detection of real 3-D motifs in different domains, between domains, in active sites, surfaces etc. The method uses the Geometric Hashing Paradigm which is an efficient technique originally developed for Computer Vision. The algorithm exploits the geometrical constraints of rigid objects, it is especially geared towards recognition of partial structures in rigid objects belonging to large data bases and is straightforwardly parallelizable. Computer Vision techniques are for the first time applied to molecular structure comparison, resulting in an efficient, fully automated tool. The method has been tested in a number of cases, including comparisons of the haemoglobins, immunoglobulins, serine proteinases, calcium binding proteins, DNA binding proteins and others. In all examples our results were equivalent to the published results from previous methods and in some cases additional structural information was obtained by our method.

Amino Acid Sequence

DNA sequences at and between the GC and TATA boxes: potential DNA looping and spatial juxtapositioning of the protein factors.

Regulation of gene expression in eukaryotes involves a complex assembly of DNA recognition sequence elements and their respective protein factors. The upstream promoter/enhancer sequences are position and orientation independent. Despite their variable distances from the TATA box and transcription start site, interaction between the protein activators and TATA general transcription factors takes place, enabling induced levels of transcription initiation. Here the intervening sequences between the GC and TATA boxes are examined as functions of their lengths. Regardless of the substantial differences in the spacer sizes, similar mono and dinucleotide distributions are noted. Purine-purine base pair steps, except for AA, are more frequent at and near the GC box in the 5' ends of the loops than in their 3' ends. Pyrimidine-pyrimidine base pair steps, except for TT behave similarly. AT and TA (as well as AA and TT) are more frequent in the 3' ends of the loops near the TATA. Examination of these distributions, as well as of the sequences composing the GC and TATA boxes indicates that the DNA in the upstream part of the loop is more rigid, whereas the downstream regions are far more flexible. The flexibility of the general TATA region may afford correct spatial juxtapositioning of the proteins with respect to each other, enabling interactions between the activators and the general transcription factors.

Base Composition

Efficient detection of three-dimensional structural motifs in biological macromolecules by computer vision techniques.

Macromolecules carrying biological information often consist of independent modules containing recurring structural motifs. Detection of a specific structural motif within a protein (or DNA) aids in elucidating the role played by the protein (DNA element) and the mechanism of its operation. The number of crystallographically known structures at high resolution is increasing very rapidly. Yet, comparison of three-dimensional structures is a laborious time-consuming procedure that typically requires a manual phase. To date, there is no fast automated procedure for structural comparisons. We present an efficient O(n3) worst case time complexity algorithm for achieving such a goal (where n is the number of atoms in the examined structure). The method is truly three-dimensional, sequence-order-independent, and thus insensitive to gaps, insertions, or deletions. This algorithm is based on the geometric hashing paradigm, which was originally developed for object recognition problems in computer vision. It introduces an indexing approach based on transformation invariant representations and is especially geared toward efficient recognition of partial structures in rigid objects belonging to large data bases. This algorithm is suitable for quick scanning of structural data bases and will detect a recurring structural motif that is a priori unknown. The algorithm uses protein (or DNA) structures, atomic labels, and their three-dimensional coordinates. Additional information pertaining to the structure speeds the comparisons. The algorithm is straightforwardly parallelizable, and several versions of it for computer vision applications have been implemented on the massively parallel connection machine. A prototype version of the algorithm has been implemented and applied to the detection of substructures in proteins.

Algorithms

Trends in the 5' vs. 3' flanks of oligonucleotides in eukaryotic and prokaryotic genomes: the asymmetric roles played by cytosine and guanine.

Studies of sequence context preferences of oligonucleotides composed of (G/C)n and (A/T)m blocks (n + m = 3,4,5) unravel strong patterns. Comparisons of the 5' and 3' nearest neighbor doublets flanking these oligomers reveal the preference of (G/C)2 to be positioned immediately next to the (A/T)m block, enclosing it by (G/C) nucleotides rather than extending the (G/C)n block. That is, for a (G/C)n(A/T)m oligomer and a (G/C)2 doublet, (G/C)n(A/T)m(G/C)2 greater than (G/C)n + 2 (A/T)m. Similarly for an (A/T)m(G/C)n oligomer, (G/C)2(A/T)m(G/C)n greater than (A/T)m(G/C)n + 2. In an analogous manner, (A/T)2 flanking doublets prefer enclosing the (G/C)n blocks, although these patterns are weaker. Here we show a strong, direct relationship between the magnitude of the trends and the presence of Cs in the (G/C)n block in the (G/C)n(A/T)m oligomer, and the presence of Gs in the complementary (A/T)m(G/C)n oligomers. The trends are stronger in eukaryotic than in prokaryotic sequences. They are stronger for longer (G/C)n and shorter (A/T)m blocks. We suggest that the preference for (A/T)m to be enclosed by (G/C) rather than be flanked by them on only one side is related to DNA structure and DNA-protein interaction. Sequences of the (G/C)(A/T)(G/C) type may have more homogeneous minor groove geometry. In particular, the strong G vs. C asymmetry in the trends may be related to pyrimidine-purine junctions, possibly to CG sequences.

Cytosine

The ordering of nucleotides in the DNA: strong pyrimidine-purine patterns near homooligomer tracts.

Here, we study the frequencies of occurrence of homooligomers flanked by one base, XnU or UXn, where X = A, C, G, T and U not equal to X. Specifically, we search for preferences (or discriminations) in their nearest neighbor doublet, VV. Extensive analysis of the data base reveals striking patterns in such VVUXn or UXn VV oligomers (V = A, C, G, T). With very few exceptions, if the VV and Xn are composed of complementary nucleotides, those oligomers having a pyrimidine (Y)-purine (R) junction are preferred over those with an RY one. If the VV and Xn nucleotides are not complementary, the RY junction oligomers are preferred over their YR counterparts. These trends are observed consistently in eukaryotic and prokaryotic sequences. They are particularly striking in the YR greater than RY oligomers containing complementary nucleotides. The general preferences and discriminations described here are in the same direction as our previous results for homooligomer tracts. These recurrences, along with some additional universal "rules", aid in our understanding of the ordering of nucleotides in the DNA.

Animals

Distinct patterns in the dinucleotide nearest neighbors to G/C and A/T oligomers in eukaryotic sequences.

The eukaryotic and prokaryotic databases are scanned for potential nearest-neighbor doublet preferences at the 5' and 3' flanks of some oligomers. Here we focus on oligomers containing alternating nucleotides, i.e., UV, UVUV, and UUVV where U not equal to V. Strong, consistent trends are observed in eukaryotic sequences. A/T alternation oligomers are preferentially flanked by A/T. G/C flanks are disfavored. G/C alternation oligomers are preferentially flanked by G/C. A/T flanks are disfavored. These trends are consistent with those observed previously for homooligomer tracts (Nussinov et al. 1989a,b). G/C tracts are preferentially flanked by G/C. A/T nearest neighbors are disfavored. The reverse holds for A/T tracts. Additional patterns are described here as well. The possible origin of these DNA composition and sequence trends is discussed. These trends are suggested to stem from protein-DNA interaction constraints.

Base Sequence

RNA pseudoknots downstream of the frameshift sites of retroviruses.

RNA pseudoknot structural motifs could have implications for a wide range of biological processes of RNAs. In this study, the potential RNA pseudoknots just downstream from the known and suspected retroviral frame-shift sites were predicted in the Rous sarcoma virus, primate immunodeficiency viruses (HIV-1, HIV-2, and SIV), equine infectious anemia virus, visna virus, bovine leukemia virus, human T-cell leukemia virus (types I and II), mouse mammary tumor virus, Mason-Pfizer monkey virus, and simian SRV-1 type-D retrovirus. Also, the putative RNA pseudoknots were detected in the gag-pol overlaps of two retrotransposons of Drosophila, 17.6 and gypsy, and the mouse intracisternal A particle. For each sequence, the thermodynamic stability and statistical significance of the secondary structure involved in the predicted tertiary structure were assessed and compared. Our results show that the stem-loop structures in the pseudoknots are both thermodynamically highly stable and statistically significant relative to other such configurations that potentially occur in the gag-pol or gag-pro and pro-pol junction domains of these viruses (300 nucleotides upstream and downstream from the possible frameshift sites are included). Moreover, the structural features of the predicted pseudoknots following the frameshift site of pro-pol overlaps of the HTLV-1 and HTLV-2 retroviruses are structurally well conserved. The occurrence of eight compensatory base changes in the tertiary interaction of the two related sequences allow the conservation of their tertiary structures in spite of the sequence divergence. The results support the possible control mechanism for frameshifting proposed by Brierley et al. and Jacks et al.

Base Sequence

Compositional variations in DNA sequences.

Biologically occurring nucleotide sequences differ from randomly generated ones. Here we describe general patterns found in prokaryotic and in eukaryotic DNA. In the accompanying paper (Nussinov, 1991) we also describe DNA signals recognized by their corresponding protein factors. In particular, we focus on modes of searches for such patterns and signals and on the potential properties such sequences may possess.

DNA

Signals in DNA sequences and their potential properties.

DNA and RNA molecules contain signals which are recognized by regulatory proteins or enzymes either directly, through their nucleotide sequences or indirectly, through induced structural changes on their neighboring sequences. To date, most signal searches have been focused on specific recurrences of nucleotide sequences. Much less attention has been directed towards the structure, flexibility and hydrogen-bonding patterns that recognition elements may possess. Here we review the various methods involved in such searches. In particular, however, we also address the searches for potential properties. In this regard it is of interest to inspect the asymmetry in the distributions of complementary oligomers near biological features. Upstream of transcription initiation the frequencies of G-rich oligomers are particularly high (Nussinov, 1987a; 1990). The frequencies of C-rich oligomers are lower. A-tracts are also very frequent in these regions. This may correlate with the recent finding that guanine, but not cytosine, tracts enhance A-tract directed bends (Milton et al., 1990). Presumably A-tracts near G-tracts on the same strand may induce a structural change in the G-tracts which may enhance the bend. G-tracts may have the potential for participating in a DNA bend due to their compression of the major groove. Thus, proteins may not always be necessary to induce DNA conformational changes. This example illustrates the importance of studies of the properties of DNA oligomers in regulatory regions, and of algorithms for their detection.

Algorithms

Long range and symmetry considerations in the DNA.

The common description of DNA is that of a polymer composed of overlapping dinucleotide base pairs. Conformational, thermodynamic and flexibility calculations of the DNA are frequently based on that description. As demonstrated by the good fit with experimental data, such as the free energy of opening the double stranded DNA and the correspondence between the DNA crystal and computed structures, this approach is to a large extent a faithful reproduction of the DNA. Here I show that longer range effects should be considered as well. Statistically strong longer-range sequence patterns are described. Such sequence preferences are clearly translated in the DNA into long range structural propagation. Long range effects have also been observed in some experimental studies, like gel mobility and patterns of DNA cleavage. Long range structural effects are the likely explanation of the effect of mutations at a distance from a protein binding site. They might also aid in understanding DNA looping.

Base Composition

Speeding up the dynamic algorithm for planar RNA folding.

The simplest dynamic algorithm for planar RNA folding searches for the maximum number of base pairs. The algorithm uses O(n3) steps. The more general case, where different weights (energies) are assigned to stacked base pairs and to the various types of single-stranded region topologies, requires a considerably longer computation time because of the partial backtracking involved. Limiting the loop size reduces the running time back to O(n3). Reduction in the number of steps in the calculations of the various RNA topologies has recently been suggested, thereby improving the time behavior. Here we show how a "jumping" procedure can be used to speed up the computation, not only for the maximal number of base pairs algorithm, but for the minimal energy algorithm as well.

Algorithms

General nearest neighbor preferences in G/C oligomers interrupted by A/T: correlation with DNA structure.

The frequencies of occurrence of the 5' and 3' nearest neighbor doublets of oligonucleotides containing (G/C) and (A/T) blocks show strong trends. Specifically, the following trends are observed. Given a (G/C)n (A/T)m oligomer (where G/C)n indicates a sequence of length n composed solely of Gs and/or Cs and (A/T)m is a sequence of length m composed solely of As and/or Ts, and n = 3,2,1; m = 1,2,3) and a (G/mC)2 doublet, (G/C)n (A/T)m (G/C)2 greater than (G/C)n + 2 (A/T)m. That is the (G/C)2 doublet is preferentially located 3' of the oligomer, enclosing the (A/T)m stretch. The trends are strongest for n = 3, m = 1 and gradually weaken as the size of the (mG/C)n block decreases (with a concomitant increase of (A/T)m). (A/T)2 nearest neighbor flank preferentially encloses the (G/C)n block (to produce (A/T)2 (G/C)n (A/T)m). The (A/T)2 flank trends are weaker than the (G/C)2 flank ones. The (A/T)2 flank trends also decrease in strength as the size of the (G/C)n block decreases. The statistical significance of these trends in eukaryotes is very high. A possible correlation with DNA structural parameters, in particular groove geometry, is discussed.

Base Composition

Sequence signals in eukaryotic upstream regions.

Two DNA sequence elements are known to recur frequently upstream of eukaryotic polymerase II-transcribed genes. The TATAAA, at position -40, specifies the transcription initiation site. The GGCCAATCT is less frequent around -80. Sequence analysis of upstream regions reveals that the underlined yeast UAS2 consensus sequence, TGATTGGT, is also very frequent at -80 in higher polymerase II-transcribed animal sequences. The underlined CCAAT box and yeast UAS sequences are complementary. Structural analysis suggests some symmetry in their DNA structures. Upstream of the TATAAT-rich region there is an abundance of GC sequences. Analysis of nucleotide tracts indicates that these are preferentially flanked by their complementary nucleotides with a pyrimidine-purine junction, i.e., TTAN, CCGn, CnGG, TnAA. Here, I discuss DNA structural consideration in upstream regions along with protein readout of the major and minor groove information content. These sequence-structure aspects are put in the general context of protein (factors)-DNA (elements) recognition and regulation.

Animals

Sequence dependence of DNA conformational flexibility.

By using conformational free energy calculations, we have studied the sequence dependence of flexibility and its anisotropy along various conformational variables of DNA base pairs. The results show the AT base step to be very flexible along the twist coordinate. On the other hand, homonucleotide steps, GG(CC) and AA(TT), are among the most rigid sequences. For the roll motion that would correspond to a bend, the TA step is most flexible, while the GG(CC) step is least flexible. The flexibility of roll is quite anisotropic; the ratio of fluctuations toward the major and minor grooves is the largest for the GC step and the smallest for the AA(TT) and CG steps. Propeller twisting of base pairs is quite flexible, especially of A.T base pairs; propeller twist can reach 19 degrees by thermal fluctuation. We discuss the effect of electrostatic parameters, comparison with available experimental results, and biological relevance of these results.

Base Sequence

Distinct patterns in homooligomer tract sequence context in prokaryotic and eukaryotic DNA.

The distributions of the junction sequences of homooligomer tracts of various lengths have been examined in prokaryotic DNA sequences and compared with those of eukaryotes. The general trends in the nearest and next to nearest neighbors to the tracts are similar for both groups. In both prokaryotes and eukaryotes A/T runs are preferentially flanked on either the 5' or the 3' ends by A and/or T. G/C runs are preferentially flanked by G and/or C. There is discrimination against A/T runs flanked by G or C and G/C runs flanked by A or T. However, whereas the distribution of prokaryotic homooligomer tract junction sequences was quite homogeneous, large variations were observed in the 5-fold larger eukaryotic database, increasing in magnitude from tracts of length 2 to 3 to 4 base pairs long. Possible DNA conformational implications and in particular DNA curvature and packaging aspects of prokaryotes and eukaryotes are discussed.

Base Sequence

Tree graphs of RNA secondary structures and their comparisons.

To facilitate comparison of RNA secondary structures each structure is represented as an ordered labeled tree. Several alternate secondary structures yielding a set of trees can be computed for any given RNA molecule (sequence). Frequently recurring subtrees are searched in this set of trees. The consensus structure motifs are then selected and used to construct a secondary structure model of the RNA. Given the difficulties involved in RNA secondary structure calculations, this procedure may significantly improve our predictive capabilities. In addition, the change of secondary structures between two different RNA sequences is described as a transformation of ordered trees. The transferable ratio of tree A from tree B is defined as a proportion of the largest common subtrees in trees A and B occurring in tree A. The method is applied to the study of the mechanism of human alpha 1 globin pre-mRNA splicing. In the study, two tentative splicing mechanisms, A and B, with different orders of intron excision from alpha 1 globin pre-mRNA have been stimulated. A possible relationship between the structural features of the secondary structures and the order of intron excision in the pathway of precursor splicing of human alpha 1 globin is discussed.

Base Sequence