PubMed HealthSearch

Biomedical subjects

S Y Le

Publications and source records attributed to S Y Le.

At least 19 recordsLinked to original sources

Conserved tertiary structure elements in the 5' untranslated region of human enteroviruses and rhinoviruses.

A combination of comparative sequence analysis and thermodynamic methods reveals the conservation of tertiary structure elements in the 5' untranslated region (UTR) of human enteroviruses and rhinoviruses. The predicted common structural elements occur in the 3' end of a segment that is critical for internal ribosome binding, termed "ribosome landing pad" (RLP), of polioviruses. Base pairings between highly conserved 17-nucleotide (nt) and 21-nt sequences in the 5' UTR of human enteroviruses and rhinoviruses constitute a predicted pseudoknot that is significantly more stable than those that can be formed from a large set of randomly shuffled sequences. A conserved single-stranded polypyrimidine tract is located between two conserved tertiary elements. R. Nicholson, J. Pelletier, S.-Y. Le, and N. Sonenberg (1991, J. Virol. 65, 5886-5894) demonstrated that the point mutations of 3-nt UUU out of an essential 4-nt pyrimidine stretch sequence UUUC abolished translation. Structural analysis of the mutant sequence indicates that small point mutations within the short polypyrimidine sequence would destroy the tertiary interaction in the predicted, highly ordered structure. The proposed common tertiary structure can offer experimentalists a model upon which to extend the interpretations for currently available data. Based on these structural features possible base-pairing models between human enteroviruses and 18 S rRNA and between human rhinoviruses and 18 S rRNA are proposed. The proposed common structure implicates a biological function for these sequences in translational initiation.

Base Composition

A procedure for RNA pseudoknot prediction.

The RNA pseudoknot has been proposed as a significant structural motif in a wide range of biological processes of RNAs. A pseudoknot involves intramolecular pairing of bases in a hairpin loop with bases outside the stem of the loop to form a second stem and loop region. In this study, we propose a method for searching and predicting pseudoknots that are likely to have functional meaning. In our procedure, the orthodox hairpin structure involved in the pseudoknot is required to be both statistically significant and relatively stable to the others in the sequence. The bases outside the stem of the hairpin loop in the predicted pseudoknot are not entangled with any formation of a highly stable secondary structure in the sequence. Also, the predicted pseudoknot is significantly more stable than those that can be formed from a large set of scrambled sequences under the assumption that the energy contribution from a pseudoknot is proportional to the size of second loop region and planar energy contribution from second stem region. A number of functional pseudoknots that have been reported before can be identified and predicted from their sequences by our method.

Algorithms

Nucleosome fractionation by mercury affinity chromatography. Contrasting distribution of transcriptionally active DNA sequences and acetylated histones in nucleosome fractions of wild-type yeast cells and cells expressing a histone H3 gene altered to encode a cysteine 110 residue.

A technique for the separation of transcriptionally active and inactive nucleosomes by mercury affinity chromatography has been applied to study the nucleosomal distribution of DNA sequences from the GAL1, ACT1, HIS4, MAT alpha, and HMRa genes of yeast. In mammalian cells, the method has been shown to separate active from inactive nucleosomes and to fractionate the active nucleosomes into two classes, one retained on the mercury column because of salt-labile associations with certain thiol-reactive non-histone proteins, and the other bound by covalent linkage of the cysteine 110 thiol groups of histone H3 molecules to the mercurated support. The first class of nucleosomes is elutable in 0.5 M NaCl; the second is displaced by 10 mM dithiothreitol (DTT) (Walker, J., Chen, T. A., Sterner, R., Berger, M., Winston, F., and Allfrey, V.G. (1990) J. Biol. Chem. 265, 5736-5746). We show that, in wild-type yeast cells, in which histone H3 lacks cysteinyl residues, very little DNA and a negligible complement of nucleosomes appear in the DTT-eluate, confirming the requirement for the H3-thiols in the mercury-binding reaction. Moreover, the DTT-eluted fraction is seriously deficient in the actively transcribed GAL1, ACT1, HIS4, and MAT alpha DNA sequences. Site-directed mutagenesis was employed to create an H3 gene containing a cysteine codon in place of the alanine codon at position 110 of the yeast H3 amino acid sequence. A strain was constructed containing the mutant histone H3 gene instead of the normal H3 gene. Subsequent fractionations of the mutant nucleosomes by mercury-affinity chromatography revealed a characteristic nucleosome peak in the DTT-eluted fraction. Its content of transcribed GAL1, ACT1, and HIS4 DNA sequences was 20- to 500-fold higher than that of the corresponding DTT-eluted fraction of wild-type yeast. Although this result is in accord with the finding that, in mammalian cells, the thiol groups of histone H3 become accessible when nucleosomes "unfold" during transcription, we find that nucleosomes containing the GAL1 DNA sequences of the yeast H3-mutant also bind to the mercury column when that gene is not being expressed. We conclude that many yeast nucleosomes are maintained in a "primed," potentially active state, possibly due to the very high constitutive levels of acetylation of the core histones. However, the nucleosomes of the HMRa gene, which is not expressed in a MAT alpha yeast strain, are virtually absent from the DTT-eluted nucleosome fractions of the H3-mutant cells, indicating that prolonged silencing of the gene is accompanied by compaction and loss of H3-thiol reactivity of its nucleosomes.

Acetylation

RNA pseudoknots downstream of the frameshift sites of retroviruses.

RNA pseudoknot structural motifs could have implications for a wide range of biological processes of RNAs. In this study, the potential RNA pseudoknots just downstream from the known and suspected retroviral frame-shift sites were predicted in the Rous sarcoma virus, primate immunodeficiency viruses (HIV-1, HIV-2, and SIV), equine infectious anemia virus, visna virus, bovine leukemia virus, human T-cell leukemia virus (types I and II), mouse mammary tumor virus, Mason-Pfizer monkey virus, and simian SRV-1 type-D retrovirus. Also, the putative RNA pseudoknots were detected in the gag-pol overlaps of two retrotransposons of Drosophila, 17.6 and gypsy, and the mouse intracisternal A particle. For each sequence, the thermodynamic stability and statistical significance of the secondary structure involved in the predicted tertiary structure were assessed and compared. Our results show that the stem-loop structures in the pseudoknots are both thermodynamically highly stable and statistically significant relative to other such configurations that potentially occur in the gag-pol or gag-pro and pro-pol junction domains of these viruses (300 nucleotides upstream and downstream from the possible frameshift sites are included). Moreover, the structural features of the predicted pseudoknots following the frameshift site of pro-pol overlaps of the HTLV-1 and HTLV-2 retroviruses are structurally well conserved. The occurrence of eight compensatory base changes in the tertiary interaction of the two related sequences allow the conservation of their tertiary structures in spite of the sequence divergence. The results support the possible control mechanism for frameshifting proposed by Brierley et al. and Jacks et al.

Base Sequence

Predicting common foldings of homologous RNAs.

A new approach is proposed for determining common RNA secondary structures within a set of homologous RNAs. The approach is a combination of phylogenetic and thermodynamic methods which is based on the prediction of optimal and suboptimal secondary structures, topological similarity searches and phylogenetic comparative analysis. The optimal and suboptimal RNA secondary structures are predicted by energy minimization. Structural comparison of the predicted RNA secondary structures is used to find conserved structures that are topologically similar in all these homologous RNAs. The validity of the conserved structural elements found is then checked by phylogenetic comparison of the sequences. This procedure is used to predict common structures of ribonuclease P (RNAase P) RNAs.

Algorithms

Detection of unusual RNA folding regions in HIV and SIV sequences.

We have developed a method for detecting more stable and significant folding regions relative to others in the sequence. The algorithm is based on the calculation of the lowest free energy of RNA secondary structures and Monte Carlo simulation. For any given RNA segment, the stability and statistical significance of RNA folding are assessed by two measures: the stability score and the significance score. The stability score measures the degree of thermodynamic stability of the segment between all possible biological segments in the RNA sequence. The significance score characterizes the specific arrangement of the nucleotides in the segment that could imply a structural role for the sequence information. Using these two measures, we are able to detect a series of distinct folding regions where highly stable and statistically significant secondary structures occur in human immunodeficiency virus (HIV) and simian immunodeficiency virus (SIV) sequences.

Algorithms

Structural and functional analysis of the ribosome landing pad of poliovirus type 2: in vivo translation studies.

The naturally uncapped genomic and mRNAs of poliovirus initiate translation by an internal ribosome-binding mechanism. The mRNA 5' untranslated region (UTR) of poliovirus is approximately 750 nucleotides in length and has seven to eight (depending on the serotype) AUG codons upstream of the initiator AUG. The sequence required for internal ribosome binding has been termed the ribosome landing pad (RLP). To better understand the mechanisms of internal initiation, we have determined the boundaries and critical elements of the RLP of poliovirus type 2 (Lansing strain) in vivo. By using deletion analysis, we demonstrate the existence of a core RLP in the poliovirus mRNA 5' UTR whose boundaries are between nucleotides 134 and 155 at the 5' end and nucleotides 556 and 585 at the 3' end. Sequences flanking the core RLP affect translational activity. The importance of several stem-loop structures in the RLP for internal initiation has been determined. Mutation of the phylogenetically conserved loop sequences in the proximal stem-loop structure of the RLP (stem-loop structure III; nucleotides 127 to 165) abolished internal translation. However, deletion of the second stem-loop in the RLP (stem-loop structure IV; nucleotides 189 to 223) reduced internal translation by only 50%. Internal deletions encompassing nucleotides 240 to 300, 350 to 380, or 450 to 480, predicted to disrupt stem-loop structure V and possibly VI, also abrogated internal initiation. Small point mutations within a short polypyrimidine sequence, highly conserved among all picornaviruses, abolished translation. A conservation of distance between the conserved polypyrimidine tract and a downstream AUG could play an important role in the mechanism of internal initiation.

Animals

Common structures of the 5' non-coding RNA in enteroviruses and rhinoviruses. Thermodynamical stability and statistical significance.

A total of 4051 suboptimal secondary structures are predicted by folding the 5' non-coding region of ten polioviruses, five human rhinoviruses and three coxsackieviruses using our new suboptimal folding algorithm for the prediction of both optimal and suboptimal RNA secondary structures. A comparative analysis of these RNA secondary structures reveals the conservation of common secondary structure that can be supported by phylogenetic data. The thermodynamic stability and statistical significance of these predicted, conserved helical elements are assessed and significant structure motifs in the 5' non-coding region are proposed. The possible roles of these structure motifs in the virus life cycle are discussed.

Base Sequence

A highly conserved RNA folding region coincident with the Rev response element of primate immunodeficiency viruses.

A series of unusual folding regions (UFR) immediately 3' to the cleavage site of the outer membrane protein (OMP) and transmembrane protein (TMP) were detected in the envelope gene RNA of the human immunodeficiency virus (HIV-1, HIV-2) and simian immunodeficiency virus (SIV) by an extensive Monte Carlo simulation. These RNA secondary structures were predicted to be both highly stable and statistically significant. In the calculation, twenty-five different sequence isolates of HIV-1, three isolates of HIV-2 and eight sequences of SIV were included. Although significant sequence divergence occurs in the env coding regions of these viruses, a distinct UFR of 234-nt is consistently located ten nucleotides 3' to the cleavage site of the OMP/TMP in HIV-1, and a 216-nt UFR occurs forty-six and forty-nine nucleotides downstream from the OMP/TMP cleavage site of HIV-2 and SIV, respectively. Compensatory base changes in the helical stem regions of these conserved RNA secondary structures are identified. These results support the hypothesis that these special RNA folding regions are functionally important and suggest that the role of this sequence as the Rev response element (RRE) is mediated by secondary structure as well as primary RNA sequence.

Base Sequence

Speeding up the dynamic algorithm for planar RNA folding.

The simplest dynamic algorithm for planar RNA folding searches for the maximum number of base pairs. The algorithm uses O(n3) steps. The more general case, where different weights (energies) are assigned to stacked base pairs and to the various types of single-stranded region topologies, requires a considerably longer computation time because of the partial backtracking involved. Limiting the loop size reduces the running time back to O(n3). Reduction in the number of steps in the calculations of the various RNA topologies has recently been suggested, thereby improving the time behavior. Here we show how a "jumping" procedure can be used to speed up the computation, not only for the maximal number of base pairs algorithm, but for the minimal energy algorithm as well.

Algorithms

Visna virus encodes a post-transcriptional regulator of viral structural gene expression.

Visna virus is an ungulate lentivirus that is distantly related to the primate lentiviruses, including human immunodeficiency virus type 1 (HIV-1). Replication of HIV-1 and of other complex primate retroviruses, including human T-cell leukemia virus type I (HTLV-I), requires the expression in trans of a virally encoded post-transcriptional activator of viral structural gene expression termed Rev (HIV-1) or Rex (HTLV-I). We demonstrate that the previously defined L open reading frame of visna virus encodes a protein, here termed Rev-V, that is required for the cytoplasmic expression of the incompletely spliced RNA that encodes the viral envelope protein. Transactivation by Rev-V was shown to require a cis-acting target sequence that coincides with a predicted RNA secondary structure located within the visna virus env gene. However, Rev-V was unable to function by using the structurally similar RNA target sequences previously defined for Rev or Rex and, therefore, displays a distinct sequence specificity. Remarkably, substitution of this visna virus target sequence in place of the HIV-1 Rev response element permitted the Rev-V protein to efficiently rescue the expression of HIV-1 structural proteins, including Gag, from a Rev- proviral clone. These results suggest that the post-transcriptional regulation of viral structural gene expression may be a characteristic feature of complex retroviruses.

Animals

A computational procedure for assessing the significance of RNA secondary structure.

In our recent series of papers, we have used the structures of statistical significance from Monte Carlo simulations to improve the predictions of secondary structure of RNA and to analyze the possible role of locally significant structures in the life cycle of human immunodeficiency virus. Because of intensive computational requirements for Monte Carlo simulation, it becomes impractical even using a supercomputer to assess the significance of a structure with a window size greater than 200 along an RNA sequence of 1000 bases or more. In this paper, we have developed a new procedure that drastically reduces the time needed to assess the significance of structures. In fact, the efficiency of this new method allows us to assess structures on the VAX as well as the CRAY.

Biometry

Thermodynamic stability and statistical significance of potential stem-loop structures situated at the frameshift sites of retroviruses.

RNA stem-loop structures situated just 3' to the frameshift sites of the retroviral gag-pol or gag-pro and pro-pol regions may make important contributions to frame-shifting in retroviruses. In this study, the thermodynamic stability and statistical significance of such secondary structural features relative to others in the sequence have been assessed using a newly developed method that combines calculations of the lowest free energy of formation of RNA secondary structures and the Monte Carlo simulations. Our results show that stem-loop structures situated just 3' to the frameshift sites are both highly stable and statistically significant relative to others in the gag-pol or gag-pro and pro-pol junction domains (both 300 nucleotides upstream and downstream from the possible frameshift sites are included) of Rous sarcoma virus (RSV), human immunodeficiency virus (HIV-1), bovine leukemia virus (BLV), human T-cell leukemia virus type II (HTLV-II), and mouse mammary tumor virus (MMTV). No other more stable, or significant folding regions are predicted in these domains.

Avian Sarcoma Viruses

A method for assessing the statistical significance of RNA folding.

We have developed a statistical method that is designed for analyzing potential RNA folded substructures. The statistical significance of RNA folding is assessed by the segment score. The segment score is defined as the difference between the lowest free energy calculated for the real biological sequence and the mean of the lowest free energies from random permutations of the real segment sequence, divided by the standard deviation of the random sample. This procedure was applied to the well-studied Escherichia coli 16S rRNA and potato spindle tuber viroid (PSTV) RNA. The results showed that the predictions of the locally significant secondary structures in these two molecules are in accord with the universally conserved local secondary structure elements (Gutell, Weiser & Noller, 1985, Prog. Nucl. Acid Res. molec. Biol. 32, 155-216; Riesner & Gross, 1985, A. Rev. Biochem. 54, 531-564). In addition, a statistical analysis indicated that the lowest free energies of a random sample set follow an approximately normal distribution. A reasonable size for the random sample set was determined statistically. Moreover, the statistical evaluation has been carried out using three different sets of energy rules--two sets (Salser, 1977, Cold Spring Harb. Symp. Quant Biol. 42, 985-1002; Freier, Kierzek, Jaeger, Sugimoto, Caruthers, Neilson & Turner, 1986, Proc. natn. Acad. Sci. U.S.A. 83, 9373-9377) take into account stacking energies and are based on experimental data and their computational extension (Salser, 1977)--the third set is a simplistic "unitary matrix" approach, where any base-pair is given a weight of "minus one" and an unpaired based is "zero". The Freier energy rules usually yield the strongest indication of significant folding region. However, the results derived from paired comparisons test don't provide sufficient evidence for concluding that a different set of energy rules is effective in changing the segment score level for local stem-loop structures in the 16S rRNA.

Analysis of Variance

Sequence divergence and open regions of RNA secondary structures in the envelope regions of the 17 human immunodeficiency virus isolates.

Genetic variation during the course of infection of an individual is a remarkable feature of the acquired immune deficiency syndrome (AIDS) disease. This variation has been studied for the envelope protein encoding regions of seventeen different sequences from various isolates of human immunodeficiency virus (HIV) using multiple sequence comparison and calculation of variability. The open regions with little intramolecular base pairing in these envelope sequences are predicted by a recently developed statistical method. The minimum length L for a run of hypervariable sites, conserved sites, or open regions that gives significance at the 1% (or 0.1%) level is then determined by a scan statistical method. The results show that significant clusters of open regions predicted at the RNA levels correlate with significant clusters of hypervariable sites in the HIV envelope gene. Those significant genomic variations in HIVs seem to be manifested mainly in the extracellular portion of the envelope protein. Twelve potential antigenic determinants are predicted using an antigenic index method. Interestingly, the majority of the significant hypervariable regions in the exterior envelope protein (gp120) were predicted potential epitopes.

Amino Acid Sequence

The HIV-1 rev trans-activator acts through a structured target sequence to activate nuclear export of unspliced viral mRNA.

Human immunodeficiency virus type 1 (HIV-1) replication requires the expression of two classes of viral mRNA. The early class of HIV-1 transcripts is fully spliced and encodes viral regulatory gene products. The functional expression of one of these nuclear regulatory proteins, termed Rev (formerly Art or Trs), induces the cytoplasmic expression of the incompletely spliced, late class of HIV-1 mRNAs that encode the viral structural proteins, including Gag and Env. Here, we provide evidence that this induction reflects the export from the cell nucleus to the cytoplasm of a pool of unspliced viral RNA constitutively expressed in the nucleus. The hypothesis that Rev acts on RNA transport, rather than splicing, is further supported by the observation that the cytoplasmic expression of a non-spliceable HIV-1 env gene sequence is also subject to Rev regulation. Here we show that this Rev response requires a specific target sequence which coincides with a complex RNA secondary structure present in the env gene. The response to Rev is fully maintained when this sequence is relocated to other exonic or intronic locations within env but is ablated by inversion. These results indicate that the HIV-1 rev gene product induces HIV-1 structural gene expression by activating the sequence-specific nuclear export of incompletely spliced HIV-1 RNA species.

Base Sequence

Tree graphs of RNA secondary structures and their comparisons.

To facilitate comparison of RNA secondary structures each structure is represented as an ordered labeled tree. Several alternate secondary structures yielding a set of trees can be computed for any given RNA molecule (sequence). Frequently recurring subtrees are searched in this set of trees. The consensus structure motifs are then selected and used to construct a secondary structure model of the RNA. Given the difficulties involved in RNA secondary structure calculations, this procedure may significantly improve our predictive capabilities. In addition, the change of secondary structures between two different RNA sequences is described as a transformation of ordered trees. The transferable ratio of tree A from tree B is defined as a proportion of the largest common subtrees in trees A and B occurring in tree A. The method is applied to the study of the mechanism of human alpha 1 globin pre-mRNA splicing. In the study, two tentative splicing mechanisms, A and B, with different orders of intron excision from alpha 1 globin pre-mRNA have been stimulated. A possible relationship between the structural features of the secondary structures and the order of intron excision in the pathway of precursor splicing of human alpha 1 globin is discussed.

Base Sequence

Functional comparison of the Rev trans-activators encoded by different primate immunodeficiency virus species.

The known primate lentiviruses can be divided into two subgroups consisting of the human immunodeficiency virus type 1 (HIV-1) isolates and the related HIV type 2 (HIV-2) and simian immunodeficiency virus (SIV) isolates. HIV-1 has been shown to encode a post-transcriptional trans-activator of viral structural gene expression, termed Rev, that is essential for viral replication in culture. Here, we demonstrate that HIV-2 and SIVmac also encode functional Rev proteins. As in the case of HIV-1, these Rev trans-activators are shown to induce the cytoplasmic expression of the unspliced viral transcripts that encode the viral structural proteins. Unexpectedly, the Rev proteins of HIV-2 and SIVmac proved incapable of activating the cytoplasmic expression of unspliced HIV-1 transcripts, whereas HIV-1 Rev was fully functional in the HIV-2/SIV system. This nonreciprocal complementation may imply a direct role for Rev in mediating the recognition of its viral RNA target sequence.

Animals