PubMed Health⌕ Search

Biomedical subjects

Sean R Eddy

Publications and source records attributed to Sean R Eddy.

13 recordsLinked to original sources

Computational identification of non-coding RNAs in Saccharomyces cerevisiae by comparative genomics.

We screened for new structural non-coding RNAs (ncRNAs) in the genome sequence of the yeast Saccharomyces cerevisiae using computational comparative analysis of genome sequences from five related species of Saccharomyces. The screen identified 92 candidate ncRNA genes. Thirteen showed discrete transcripts when assayed by northern blot. Of these, eight appear to be novel ncRNAs ranging in size from 268 to 775 nt, including three new H/ACA box small nucleolar RNAs.

Base Sequence↗

An active DNA transposon family in rice.

The publication of draft sequences for the two subspecies of Oryza sativa (rice), japonica (cv. Nipponbare) and indica (cv. 93-11), provides a unique opportunity to study the dynamics of transposable elements in this important crop plant. Here we report the use of these sequences in a computational approach to identify the first active DNA transposons from rice and the first active miniature inverted-repeat transposable element (MITE) from any organism. A sequence classified as a Tourist-like MITE of 430 base pairs, called miniature Ping (mPing), was present in about 70 copies in Nipponbare and in about 14 copies in 93-11. These mPing elements, which are all nearly identical, transpose actively in an indica cell-culture line. Database searches identified a family of related transposase-encoding elements (called Pong), which also transpose actively in the same cells. Virtually all new insertions of mPing and Pong elements were into low-copy regions of the rice genome. Since the domestication of rice mPing MITEs have been amplified preferentially in cultivars adapted to environmental extremes-a situation that is reminiscent of the genomic shock theory for transposon activation.

Amino Acid Sequence↗

Rfam: an RNA family database.

Rfam is a collection of multiple sequence alignments and covariance models representing non-coding RNA families. Rfam is available on the web in the UK at http://www.sanger.ac.uk/Software/Rfam/ and in the US at http://rfam.wustl.edu/. These websites allow the user to search a query sequence against a library of covariance models, and view multiple sequence alignments and family annotation. The database can also be downloaded in flatfile form and searched locally using the INFERNAL package (http://infernal.wustl.edu/). The first release of Rfam (1.0) contains 25 families, which annotate over 50 000 non-coding RNA genes in the taxonomic divisions of the EMBL nucleotide database.

Animals↗

A uniform system for microRNA annotation.

MicroRNAs (miRNAs) are small noncoding RNA gene products about 22 nt long that are processed by Dicer from precursors with a characteristic hairpin secondary structure. Guidelines are presented for the identification and annotation of new miRNAs from diverse organisms, particularly so that miRNAs can be reliably distinguished from other RNAs such as small interfering RNAs. We describe specific criteria for the experimental verification of miRNAs, and conventions for naming miRNAs and miRNA genes. Finally, an online clearinghouse for miRNA gene name assignments is provided by the Rfam database of RNA families.

MicroRNAs↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗

A memory-efficient dynamic programming algorithm for optimal alignment of a sequence to an RNA secondary structure.

BACKGROUND: Covariance models (CMs) are probabilistic models of RNA secondary structure, analogous to profile hidden Markov models of linear sequence. The dynamic programming algorithm for aligning a CM to an RNA sequence of length N is O(N3) in memory. This is only practical for small RNAs. RESULTS: I describe a divide and conquer variant of the alignment algorithm that is analogous to memory-efficient Myers/Miller dynamic programming algorithms for linear sequence alignment. The new algorithm has an O(N2 log N) memory complexity, at the expense of a small constant factor in time. CONCLUSIONS: Optimal ribosomal RNA structural alignments that previously required up to 150 GB of memory now require less than 270 MB.

Algorithms↗

Noncoding RNA genes identified in AT-rich hyperthermophiles.

Noncoding RNA (ncRNA) genes that produce functional RNAs instead of encoding proteins seem to be somewhat more prevalent than previously thought. However, estimating their number and importance is difficult because systematic identification of ncRNA genes remains challenging. Here, we exploit a strong, surprising DNA composition bias in genomes of some hyperthermophilic organisms: simply screening for GC-rich regions in the AT-rich Methanococcus jannaschii and Pyrococcus furiosus genomes efficiently detects both known and new RNA genes with a high degree of secondary structure. A separate screen based on comparative analysis also successfully identifies noncoding RNA genes in P. furiosus. Nine of the 30 new candidate genes predicted by these screens have been verified to produce discrete, apparently noncoding transcripts with sizes ranging from 97 to 277 nucleotides.

Adenine↗

RIO: analyzing proteomes by automated phylogenomics using resampled inference of orthologs.

BACKGROUND: When analyzing protein sequences using sequence similarity searches, orthologous sequences (that diverged by speciation) are more reliable predictors of a new protein's function than paralogous sequences (that diverged by gene duplication). The utility of phylogenetic information in high-throughput genome annotation ("phylogenomics") is widely recognized, but existing approaches are either manual or not explicitly based on phylogenetic trees. RESULTS: Here we present RIO (Resampled Inference of Orthologs), a procedure for automated phylogenomics using explicit phylogenetic inference. RIO analyses are performed over bootstrap resampled phylogenetic trees to estimate the reliability of orthology assignments. We also introduce supplementary concepts that are helpful for functional inference. RIO has been implemented as Perl pipeline connecting several C and Java programs. It is available at http://www.genetics.wustl.edu/eddy/forester/. A web server is at http://www.rio.wustl.edu/. RIO was tested on the Arabidopsis thaliana and Caenorhabditis elegans proteomes. CONCLUSION: The RIO procedure is particularly useful for the automated detection of first representatives of novel protein subfamilies. We also describe how some orthologies can be misleading for functional inference.

Animals↗

Computational genomics of noncoding RNA genes.

The number of known noncoding RNA genes is expanding rapidly. Computational analysis of genome sequences, which has been revolutionary for protein gene analysis, should also be able to address questions of the number and diversity of noncoding RNA genes. However, noncoding RNAs present computational genomics with a new set of challenges.

Animals↗

Archaeal guide RNAs function in rRNA modification in the eukaryotic nucleus.

In eukaryotes, many Box C/D small nucleolar RNAs base pair with ribosomal RNA through short complementary guide sequences, thereby marking up to 100 individual nucleotides of ribosomal RNA for 2'-O-methylation. Function of the eukaryotic Box C/D RNAs depends upon interaction with at least six known proteins. Box C/D RNAs are not known to exist in Bacteria but were recently identified in Archaea by biochemical analysis and computational genomic screens and have likely evolved independently in Archaea and Eukarya for more than 2000 million years. We have microinjected Box C/D RNAs from Pyrococcus furiosus, a hyperthermophilic archaeon, into the nuclei of oocytes from the aquatic frog Xenopus laevis. Our results show that Box C/D RNAs derived from this prokaryote are retained in the nucleus, localize to nucleoli, and interact with the X. laevis Box C/D RNA binding proteins fibrillarin, Nop56, and Nop58. Furthermore, we have demonstrated the ability of archaeal Box C/D RNAs to direct site-specific 2'-O-methylation of ribosomal RNA. Our studies have revealed the remarkable ability of archaeal Box C/D RNAs to assemble into functional RNA-protein complexes in the eukaryotic nucleus.

Active Transport, Cell Nucleus↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.

Animals↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗