PubMed Health⌕ Search

Biomedical subjects

S R Holbrook

Publications and source records attributed to S R Holbrook.

At least 19 recordsLinked to original sources

A computational approach to identify genes for functional RNAs in genomic sequences.

Currently there is no successful computational approach for identification of genes encoding novel functional RNAs (fRNAs) in genomic sequences. We have developed a machine learning approach using neural networks and support vector machines to extract common features among known RNAs for prediction of new RNA genes in the unannotated regions of prokaryotic and archaeal genomes. The Escherichia coli genome was used for development, but we have applied this method to several other bacterial and archaeal genomes. Networks based on nucleotide composition were 80-90% accurate in jackknife testing experiments for bacteria and 90-99% for hyperthermophilic archaea. We also achieved a significant improvement in accuracy by combining these predictions with those obtained using a second set of parameters consisting of known RNA sequence motifs and the calculated free energy of folding. Several known fRNAs not included in the training datasets were identified as well as several hundred predicted novel RNAs. These studies indicate that there are many unidentified RNAs in simple genomes that can be predicted computationally as a precursor to experimental study. Public access to our RNA gene predictions and an interface for user predictions is available via the web.

Computational Biology↗

Crystallization of RNA.

Even as the number of RNA structures determined and under study multiplies, the critical step in X-ray diffraction analysis, growth of single well-ordered crystals, remains at the boundary between art and science. Recent advances in methods of RNA synthesis, purification, and characterization, as well as empirical and technical improvements in crystallization techniques, the development of cryo-crystallography, and the wider availability of bright, tunable, X-rays from synchrotron sources are improving the chances of obtaining RNA crystals suitable for X-ray structural analysis. In this review, we summarize the current status of the design, preparation, purification, and analysis of RNA for crystallization and describe the latest approaches to obtaining diffraction-quality crystals.

Crystallization↗

The crystal structure of the Rev binding element of HIV-1 reveals novel base pairing and conformational variability.

The crystal and molecular structure of an RNA duplex corresponding to the high affinity Rev protein binding element (RBE) has been determined at 2.1-A resolution. Four unique duplexes are present in the crystal, comprising two structural variants. In each duplex, the RNA double helix consists of an annealed 12-mer and 14-mer that form an asymmetric internal loop consisting of G-G and G-A noncanonical base pairs and a flipped-out uridine. The 12-mer strand has an A-form conformation, whereas the 14-mer strand is distorted to accommodate the bulges and noncanonical base pairing. In contrast to the NMR model of the unbound RBE, an asymmetric G-G pair with N2-N7 and N1-O6 hydrogen bonding, is formed in each helix. The G-A base pairing agrees with the NMR structure in one structural variant, but forms a novel water-mediated pair in the other. A backbone flip and reorientation of the G-G base pair is required to assume the RBE conformation present in the NMR model of the complex between the RBE and the Rev peptide.

Base Sequence↗

Structure of an RNA internal loop consisting of tandem C-A+ base pairs.

The crystal structure of the RNA octamer 5'-CGC(CA)GCG-3' has been determined from X-ray diffraction data to 2.3 A resolution. In the crystal, this oligomer forms a self-complementary double helix in the asymmetric unit. Tandem non-Watson-Crick C-A and A-C base pairs comprise an internal loop in the middle of the duplex, which is incorporated with little distortion of the A-form double helix. From the geometry of the C-A base pairs, it is inferred that the adenosine imino group is protonated and donates a hydrogen bond to the carbonyl group of the cytosine. The wobble geometry of the C-A+ base pairs is very similar to that of the common U-G non-Watson-Crick pair.

Adenine↗

The crystal structure of an RNA oligomer incorporating tandem adenosine-inosine mismatches.

The X-ray crystallographic structure of the RNA duplex [r(CGCAIGCG)]2 has been refined to 2.5 A. It shows a symmetric internal loop of two non-Watson-Crick base pairs which form in the middle of the duplex. The tandem A-I/I-A pairs are related by a crystallographic two-fold axis. Both A(anti)-I(anti) mismatches are in a head-to-head conformation forming hydrogen bonds using the Watson-Crick positions. The octamer duplexes stack above one another in the cell forming a pseudo-infinite helix throughout the crystal. A hydrated calcium ion bridges between the 3'-terminal of one molecule and the backbone of another. The tandem A-I mismatches are incorporated with only minor distortion to the backbone. This is in contrast to the large helical perturbations often produced by sheared G-A pairs in RNA oligonucleotides.

Adenosine↗

RNA crystallography.

The current state of three-dimensional structure analysis of RNA by x-ray crystallography is summarized. The methods of sample preparation, crystallization, data collection, and structure solution are discussed, followed by a review of the RNA structures that have been determined and of common structural features, and finally, an appraisal of future prospects for x-ray crystal structure analysis of RNA.

Base Sequence↗

A curved RNA helix incorporating an internal loop with G.A and A.A non-Watson-Crick base pairing.

The crystal structure of the RNA dodecamer 5'-GGCC(GAAA)GGCC-3' has been determined from x-ray diffraction data to 2.3-A resolution. In the crystal, these oligomers form double helices around twofold symmetry axes. Four consecutive non-Watson-Crick base pairs make up an internal loop in the middle of the duplex, including sheared G.A pairs and novel asymmetric A.A pairs. This internal loop sequence produces a significant curvature and narrowing of the double helix. The helix is curved by 34 degrees from end to end and the diameter is narrowed by 24% in the internal loop. A Mn2+ ion is bound directly to the N7 of the first guanine in the Watson-Crick region following the internal loop and the phosphate of the preceding residue. This Mn2+ location corresponds to a metal binding site observed in the hammerhead catalytic RNA.

Adenine↗

Prediction of protein folding class using global description of amino acid sequence.

We present a method for predicting protein folding class based on global protein chain description and a voting process. Selection of the best descriptors was achieved by a computer-simulated neural network trained on a data base consisting of 83 folding classes. Protein-chain descriptors include overall composition, transition, and distribution of amino acid attributes, such as relative hydrophobicity, predicted secondary structure, and predicted solvent exposure. Cross-validation testing was performed on 15 of the largest classes. The test shows that proteins were assigned to the correct class (correct positive prediction) with an average accuracy of 71.7%, whereas the inverse prediction of proteins as not belonging to a particular class (correct negative prediction) was 90-95% accurate. When tested on 254 structures used in this study, the top two predictions contained the correct class in 91% of the cases.

Amino Acid Sequence↗

Structure of an RNA double helix including uracil-uracil base pairs in an internal loop.

The crystal structure of the RNA dodecamer 5'-GGACUUUGGUCC-3' has been determined from X-ray diffraction data to 2.6 A resolution. This oligomer forms an asymmetric double helix in the crystal. Four consecutive non-Watson-Crick base-pairs are formed in the middle of the duplex including the first intrahelical U-U (or T-T) pairs observed in an oligonucleotide crystal structure. Two different conformations of U-U pairs are observed in the context of the surrounding sequence. One of these pairs is highly twisted, allowing a bound water to bridge across strands in the major groove. The crystal packing illustrates a new form of RNA helix-helix interaction.

Base Composition↗

Use of low-molecular-weight polyethylene glycol in the crystallization of RNA oligomers.

We have crystallized a variety of RNA oligonucleotides in a form suitable for X-ray diffraction studies using polyethylene glycol with a low-molecular-weight distribution (PEG 400) as the precipitant. Crystallization experiments on a set of 26 RNA oligomers ranging from eight to 12 nucleotides in length resulted in eight diffraction-quality crystals. Of these eight RNA crystals, six utilized PEG 400 as the precipitating agent. We have also been able to obtain large single crystals of a DNA-RNA hybrid, transfer RNA (two different conditions) and a catalytic RNA from PEG 400 solutions. These results suggest that PEG 400 may be a generally useful alternative to 2-methyl-2,4-pentanediol (MPD) which has, thus far, been the most successful precipitant for DNA oligomers.

Journal Article↗

Prediction of protein folding class from amino acid composition.

An empirical relation between the amino acid composition and three-dimensional folding pattern of several classes of proteins has been determined. Computer simulated neural networks have been used to assign proteins to one of the following classes based on their amino acid composition and size: (1) 4 alpha-helical bundles, (2) parallel (alpha/beta)8 barrels, (3) nucleotide binding fold, (4) immunoglobulin fold, or (5) none of these. Networks trained on the known crystal structures as well as sequences of closely related proteins are shown to correctly predict folding classes of proteins not represented in the training set with an average accuracy of 87%. Other folding motifs can easily be added to the prediction scheme once larger databases become available. Analysis of the neural network weights reveals that amino acids favoring prediction of a folding class are usually over represented in that class and amino acids with unfavorable weights are underrepresented in composition. The neural networks utilize combinations of these multiple small variations in amino acid composition in order to make a prediction. The favorably weighted amino acids in a given class also form the most intramolecular interactions with other residues in proteins of that class. A detailed examination of the contacts of these amino acids reveals some general patterns that may help stabilize each folding class.

Amino Acid Sequence↗

PROBE: a computer program employing an integrated neural network approach to protein structure prediction.

A computer program, PROBE, has been designed for the prediction of protein structural features from amino acid sequence. This program integrates a variety of computer-simulated neural networks, each predicting an aspect of protein structure, into a single, easy-to-use package. The surface accessibility of each residue, the presence of disulfide bonds, the overall secondary structure composition and the residue secondary structures, including beta-turn type, are predicted. In addition, the overall amino acid composition and relative hydrophobicity are used to determine whether a protein belongs to one of four common folding motifs. PROBE is able to compare and synergistically improve the predictions by allowing communication between the different networks.

Amino Acid Sequence↗

Crystal structure of an RNA double helix incorporating a track of non-Watson-Crick base pairs.

The crystal structure of the RNA dodecamer duplex (r-GGACUUCGGUCC)2 has been determined. The dodecamers stack end-to-end in the crystal, simulating infinite A-form helices with only a break in the phosphodiester chain. These infinite helices are held together in the crystal by hydrogen bonding between ribose hydroxyl groups and a variety of donors and acceptors. The four noncomplementary nucleotides in the middle of the sequence did not form an internal loop, but rather a highly regular double-helix incorporating the non-Watson-Crick base pairs, G.U and U.C. This is the first direct observation of a U.C (or T.C) base pair in a crystal structure. The U.C pairs each form only a single base-base hydrogen bond, but are stabilized by a water molecule which bridges between the ring nitrogens and by four waters in the major groove which link the bases and phosphates. The lack of distortion introduced in the double helix by the U.C mismatch may explain its low efficiency of repair in DNA. The G.U wobble pair is also stabilized by a minor-groove water which bridges between the unpaired guanine amino and the ribose hydroxyl of the uracil. This structure emphasizes the importance of specific hydrogen bonding between not only the nucleotide bases, but also the ribose hydroxyls, phosphate oxygens and tightly bound waters in stabilization of the intramolecular and intermolecular structures of double helical RNA.

Base Sequence↗

Structural model of the nucleotide-binding conserved component of periplasmic permeases.

The amino acid sequences of 17 bacterial membrane proteins that are components of periplasmic permeases and function in the uptake of a variety of small molecules and ions are highly homologous to each other and contain sequence motifs characteristic of nucleotide-binding proteins. These proteins are known to bind ATP and are postulated to be the energy-coupling components of the permeases. Several medically important eukaryotic proteins, including the multidrug-resistance transporters and the protein encoded by the cystic fibrosis gene, are also homologous to this family. By multiple sequence alignment of these 17 proteins, the consensus sequence, secondary structure, and surface exposure were predicted. The secondary structural motifs that are conserved among nucleotide-binding proteins were identified in adenylate kinase, p21ras, and elongation factor Tu by superposition of their known tertiary structures. The equivalent secondary structural elements in the predicted conserved component were located. These, together with sequence information, served as guides for alignment with adenylate kinase. A model for the structure of the ATP-binding domain of the permease proteins is proposed by analogy to the adenylate kinase structure. The characteristics of several permease mutations and biochemical data lend support to the model.

Adenylate Kinase↗

A Macintosh computer program for designing DNA sequences that code for specific peptides and proteins.

A computer program (PINCERS) is described for use in the design of synthetic genes and mixed-probe DNA sequences. A protein sequence is reverse translated with generation of synonymous codons at each position producing a degenerate sequence. In order to locate potential restriction enzyme sites, the degenerate sequence is searched with a library of restriction enzymes for sites that utilize any combination of synonymous codons. These sites are indicated in a map so that they may be incorporated into the synthetic gene sequence. The program allows the user to select the appropriate codon usage table for the organism of interest and then to set a threshold usage frequency below which codons are not generated. PINCERS may also be used to assist in planning the synthesis of mixed-probe DNA sequences for cross-hybridization experiments. It can identify regions of specified length with the protein sequence that have the least overall degeneracy, thereby minimizing the number of probes to be synthesized and, therefore, maximizing the concentration of a given probe sequence.

DNA↗