PubMed HealthSearch

Biomedical subjects

M Zuker

Publications and source records attributed to M Zuker.

18 recordsLinked to original sources

Suboptimal sequence alignment in molecular biology. Alignment with error analysis.

A molecular sequence alignment algorithm based on dynamic programming has been extended to allow the computation of all pairs of residues that can be part of optimal and suboptimal sequence alignments. The uncertainties inherent in sequence alignment can be displayed using a new form of dot plot. The method allows the qualitative assessment of whether or not two sequences are related, and can reveal what parts of the alignment are better determined than others. It also permits the computation of representative optimal and suboptimal alignments. The relation between alignment reliability and alignment parameters is discussed. Other applications are to cyclical permutations of sequences and the detection of self-similarity. An application to multiple sequence alignment is noted.

Algorithms

A comparison of optimal and suboptimal RNA secondary structures predicted by free energy minimization with structures determined by phylogenetic comparison.

This article describes the latest version of an RNA folding algorithm that predicts both optimal and suboptimal solutions based on free energy minimization. A number of RNA's with known structures deduced from comparative sequence analysis are folded to test program performance. The group of solutions obtained for each molecule is analysed to determine how many of the known helixes occur in the optimal solution and in the best suboptimal solution. In most cases, a structure about 80% correct is found with a free energy within 2% of the predicted lowest free energy structure.

Algorithms

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition

Predicting common foldings of homologous RNAs.

A new approach is proposed for determining common RNA secondary structures within a set of homologous RNAs. The approach is a combination of phylogenetic and thermodynamic methods which is based on the prediction of optimal and suboptimal secondary structures, topological similarity searches and phylogenetic comparative analysis. The optimal and suboptimal RNA secondary structures are predicted by energy minimization. Structural comparison of the predicted RNA secondary structures is used to find conserved structures that are topologically similar in all these homologous RNAs. The validity of the conserved structural elements found is then checked by phylogenetic comparison of the sequences. This procedure is used to predict common structures of ribonuclease P (RNAase P) RNAs.

Algorithms

Common structures of the 5' non-coding RNA in enteroviruses and rhinoviruses. Thermodynamical stability and statistical significance.

A total of 4051 suboptimal secondary structures are predicted by folding the 5' non-coding region of ten polioviruses, five human rhinoviruses and three coxsackieviruses using our new suboptimal folding algorithm for the prediction of both optimal and suboptimal RNA secondary structures. A comparative analysis of these RNA secondary structures reveals the conservation of common secondary structure that can be supported by phylogenetic data. The thermodynamic stability and statistical significance of these predicted, conserved helical elements are assessed and significant structure motifs in the 5' non-coding region are proposed. The possible roles of these structure motifs in the virus life cycle are discussed.

Base Sequence

Melting and chemical modification of a cyclized self-splicing group I intron: similarity of structures in 1 M Na+, in 10 mM Mg2+, and in the presence of substrate.

C IVS is the cyclized form of the intron from the RNA precursor of the Tetrahymena thermophila large subunit (LSU) ribosomal RNA. C IVS was mapped by chemical modification in 1 M Na+, 0.05 M Na+ and 10 mM Mg2+ (Na+/Mg2+), and Na+/Mg2+ with CUCU substrate. The results suggest the secondary structure is similar for all three conditions. Optical melting curves were also measured for C IVS in 1 M Na+ and Na+/Mg2+ and indicate the secondary structures have similar stabilities under both conditions. Computer predictions of secondary structure and stability are in good agreement with observations. The results suggest that many of the approximations used for computer prediction of secondary structure by free energy minimization are reasonable.

Aldehydes

Fluorescence decay kinetics of the tryptophyl residues of myoglobin: effect of heme ligation and evidence for discrete lifetime components.

The fluorescence decay kinetics of the tryptophyl residues of sperm whale and yellowfin tuna myoglobin have been determined by using time-correlated single photon counting, with picosecond resolution. Purification by HPLC techniques resulted in the isolation of samples that exclusively displayed picosecond decay kinetics. Lifetimes of 24.4 ps for Trp14 and 122.0 ps for Trp7 were found for oxy sperm whale myoglobin (pH 7), which agree with theoretical predictions [Hochstrasser, R. M., & Negus, D. K. (1984) Proc. Natl. Acad. Sci. U.S.A. 81, 4399-4403]. The effects of ligand binding and pH on the decay kinetics were investigated, and the results were shown to be consistent with the known crystal structures. Data for the met form of sperm whale myoglobin were analyzed both in terms of a sum of discrete exponential components and as a continuous gamma distribution of exponential decays. The results were not found to support the existence of multiple, structurally distinct conformation states in myoglobin.

Animals

Resolution of heterogeneous fluorescence into component decay-associated excitation spectra. Application to subtilisins.

Direct and indirect methods are described to combine steady-state and picosecond time-resolved fluorescence decay data to generate decay-associated excitation spectra. The heterogeneous fluorescence from a fluorophore mixture that models protein fluorescence was resolved into individual component excitation spectra. The two methods were also used to determine the excitation spectra associated with each of the decay time components for the proteins subtilisin Carlsberg and BPN'. On the basis of associated spectra, the decay components of both proteins were assigned to individual (or groups of) emitting species. The two approaches used to generate the decay-associated excitation spectra are compared and their general application to protein fluorescence studies is discussed.

Kinetics

On finding all suboptimal foldings of an RNA molecule.

An algorithm and a computer program have been prepared for determining RNA secondary structures within any prescribed increment of the computed global minimum free energy. The mathematical problem of determining how well defined a minimum energy folding is can now be solved. All predicted base pairs that can participate in suboptimal structures may be displayed and analyzed graphically. Representative suboptimal foldings are generated by selecting these base pairs one at a time and computing the best foldings that contain them. A distance criterion that ensures that no two structures are "too close" is used to avoid multiple generation of similar structures. Thermodynamic parameters, including free-energy increments for single-base stacking at the ends of helices and for terminal mismatched pairs in interior and hairpin loops, are incorporated into the underlying folding model of the above algorithm.

Animals

The alignment of protein structures in three dimensions.

This article extends the use of dynamic programming algorithms in molecular sequence comparison to the alignment of the alpha-carbon (C alpha-) coordinates of two protein structures in three dimensions. The algorithm is described in detail and is applied to the comparison of alpha-lactalbumin with both hen egg white lysozyme and T4 lysozyme. In the first case, the structures are similar, while the second comparison is between two distantly related molecules. References are made to the usual sequence alignments. A variety of complementary methods are introduced to display the results.

Algorithms

Improved predictions of secondary structures for RNA.

The accuracy of computer predictions of RNA secondary structure from sequence data and free energy parameters has been increased to roughly 70%. Performance is judged by comparison with structures known from phylogenetic analysis. The algorithm also generates suboptimal structures. On average, the best structure within 10% of the lowest free energy contains roughly 90% of phylogenetically known helixes. The algorithm does not include tertiary interactions or pseudoknots and employs a crude model for single-stranded regions. The only favorable interactions are base pairing and stacking of terminal unpaired nucleotides at the ends of helixes. The excellent performance is consistent with these interactions being the primary interactions determining RNA secondary structure.

Base Composition

The primary structure of the ribosomal A-protein (L12) from the halophilic eubacterium Haloanaerobium praevalens.

The ribosomal A-protein, equivalent to the ribosomal protein L12 from Escherichia coli, has been sequenced from the anaerobic halophilic eubacterium Haloanaerobium praevalens (DSM 2228). The protein contains 122 amino acids, has a composition of Asp6, Asn2, Thr2, Ser6, Glu22, Pro2, Gly13, Ala19, Val12, Met4, Ile5, Leu11, Phe3, Lys14, Arg1 and has a molecular weight of 12,691. The hydrophilicity profile was determined for this protein. A phylogenetic or cluster tree was calculated from computer analysis of the sequence data on eubacterial ribosomal A-proteins. H. praevalens clusters with a group that includes Bacillus subtilis, Micrococcus lysodeikticus, Bacillus stearothermophilus and Clostridium pasteurianum.

Amino Acid Sequence

Structure of a Bacillus subtilis endo-beta-1,4-glucanase gene.

The nucleotide sequence of the portion of a Bacillus subtilis (strain PAP115) 3 kb Pst I fragment which contains an endo-beta-1, 4-glucanase gene has been determined. This gene encodes a protein of 499 amino acid residues (Mr = 55,234) with a typical B. subtilis signal peptide. Escherichia coli which has been transformed with this gene produces an extracellular endoglucanase with an amino-terminus corresponding to the thirtieth encoded amino acid residue. The gene is preceded by a cryptic reading frame with a rho-independent terminator structure, and itself has such a structure in the immediate 3'-flanking region. We have also identified, in the 5'-flanking region, nucleotide sequences which resemble promoter elements recognized by Bacillus RNA polymerase E sigma 43. Comparison of the encoded amino acid sequence to other known beta-glucanases reveals a small region of similarity to the encoded protein of the Clostridium thermocellum celB gene. These similar regions may contain substrate-binding and/or catalytic sites.

Amino Acid Sequence

Effect of spermidine on the conformation of bacteriophage MS2 RNA. Electron microscopy and computer modeling.

The structure of single-stranded RNA from the bacteriophage MS2 has been examined by electron microscopy in the presence of the polyamine spermidine. The molecules are found in two alternate conformations. The first of these can be characterized as a cruciform structure composed of three large loops approximately 500 to 700 nucleotides in size. The interior of the molecule has extensive base-paired regions which connect distant regions of the molecule; the farthest being 2500 nucleotides apart. In the second conformation, the molecules appear rod-like. Two of the large loops disappear, and these regions form, instead, extensive long-range helices. Computer modeling has been employed to explore the base-pairing potential of the sequence of bacteriophage MS2 RNA. Double-stranded regions identified by electron microscopy are shown to occur in local G + C-rich stretches of the RNA. Detailed models have been calculated for two regions of long-range contact. One of these includes the ribosome-binding site for the viral coat protein gene. The results are discussed in the context of the known role of RNA structure in the regulation of viral gene expression.

Bacteriophages

Equivalency of linear least squares curve fitting and reciprocal functions in protein circular dichroic spectra analysis.

Two methods for the analysis of the circular dichroism spectra of proteins for determination of secondary structure have been examined. These are the linear curve fitting of the data, minimized in the least squares sense, and the method of reciprocal functions proposed by C.C. Baker and I. Isenberg, Biochemistry 15 (1976) 629. It is shown that the use of these two methods give results that are identical, providing the same set of reference spectra are used in each case, and, therefore, that no new information is obtained by the use of either one over the other.

Bacterial Proteins