PubMed Health⌕ Search

Biomedical subjects

Shu Ouyang

Publications and source records attributed to Shu Ouyang.

17 recordsLinked to original sources

The rice kinase database. A phylogenomic database for the rice kinome.

The rice (Oryza sativa) genome contains 1,429 protein kinases, the vast majority of which have unknown functions. We created a phylogenomic database (http://rkd.ucdavis.edu) to facilitate functional analysis of this large gene family. Sequence and genomic data, including gene expression data and protein-protein interaction maps, can be displayed for each selected kinase in the context of a phylogenetic tree allowing for comparative analysis both within and between large kinase subfamilies. Interaction maps are easily accessed through links and displayed using Cytoscape, an open source software platform. Chromosomal distribution of all rice kinases can also be explored via an interactive interface.

Amino Acid Motifs↗

The TIGR Rice Genome Annotation Resource: improvements and new features.

In The Institute for Genomic Research Rice Genome Annotation project (http://rice.tigr.org), we have continued to update the rice genome sequence with new data and improve the quality of the annotation. In our current release of annotation (Release 4.0; January 12, 2006), we have identified 42,653 non-transposable element-related genes encoding 49,472 gene models as a result of the detection of alternative splicing. We have refined our identification methods for transposable element-related genes resulting in 13,237 genes that are related to transposable elements. Through incorporation of multiple transcript and proteomic expression data sets, we have been able to annotate 24 799 genes (31,739 gene models), representing approximately 50% of the total gene models, as expressed in the rice genome. All structural and functional annotation is viewable through our Rice Genome Browser which currently supports 59 tracks. Enhanced data access is available through web interfaces, FTP downloads and a Data Extractor tool developed in order to support discrete dataset downloads.

DNA Transposable Elements↗

Expressed sequence tags from loblolly pine embryos reveal similarities with angiosperm embryogenesis.

The process of embryogenesis in gymnosperms differs in significant ways from the more widely studied process in angiosperms. To further our understanding of embryogenesis in gymnosperms, we have generated Expressed Sequence Tags (ESTs) from four cDNA libraries constructed from un-normalized, normalized, and subtracted RNA populations of zygotic and somatic embryos of loblolly pine (Pinus taeda L.). A total of 68,721 ESTs were generated from 68,131 cDNA clones. Following clustering and assembly, these sequences collapsed into 5,274 contigs and 6,880 singleton sequences for a total of 12,154 non-redundant sequences. Searches of a non-identical amino acid database revealed a putative homolog for 9,189 sequences, leaving 2,965 sequences with no known function. More extensive searches of additional plant sequence data sets revealed a putative homolog for all but 1,388 (11.4%) of the sequences. Using gene ontologies, a known function could be assigned for 5,495 of the 12,154 total non-redundant sequences with 13,633 associations in total assigned. When compared to approximately 72,000 sequences in a collated P. taeda transcript assembly derived from >245,000 ESTs derived from root, xylem, stem, needles, pollen cone, and shoot ESTs, 3,458 (28.5%) of the non-redundant embryo sequences were unique and thereby provide a valuable addition to development of a complete loblolly pine transcriptome. To assess similarities between angiosperm and gymnosperm embryo development, we examined our EST collection for putative homologs of angiosperm genes implicated in embryogenesis. Out of 108 angiosperm embryogenesis-related genes, homologs were present for 83 of these genes suggesting that pine contains similar genes for embryogenesis and that our RNA sampling methods were successful. We also identified sequences from the pine embryo transcriptome that have no known function and may contribute to the programming of gene expression and embryo development.

Amino Acid Sequence↗

Genomic and genetic characterization of rice Cen3 reveals extensive transcription and evolutionary implications of a complex centromere.

The centromere is the chromosomal site for assembly of the kinetochore where spindle fibers attach during cell division. In most multicellular eukaryotes, centromeres are composed of long tracts of satellite repeats that are recalcitrant to sequencing and fine-scale genetic mapping. Here, we report the genomic and genetic characterization of the complete centromere of rice (Oryza sativa) chromosome 3. Using a DNA fiber-fluorescence in situ hybridization approach, we demonstrated that the centromere of chromosome 3 (Cen3) contains approximately 441 kb of the centromeric satellite repeat CentO. Cen3 includes an approximately 1,881-kb domain associated with the centromeric histone CENH3. This CENH3-associated chromatin domain is embedded within a 3,113-kb region that lacks genetic recombination. Extensive transcription was detected within the CENH3 binding domain based on comprehensive annotation of protein-coding genes coupled with empirical measurements of mRNA levels using RT-PCR and massively parallel signature sequencing. Genes <10 kb from the CentO satellite array were expressed in several rice tissues and displayed histone modification patterns consistent with euchromatin, suggesting that rice centromeric chromatin accommodates normal gene expression. These results support the hypothesis that centromeres can evolve from gene-containing genomic regions.

Centromere↗

Transcription and histone modifications in the recombination-free region spanning a rice centromere.

Centromeres are sites of spindle attachment for chromosome segregation. During meiosis, recombination is absent at centromeres and surrounding regions. To understand the molecular basis for recombination suppression, we have comprehensively annotated the 3.5-Mb region that spans a fully sequenced rice centromere. Although transcriptional analysis showed that the 750-kb CENH3-containing core is relatively deficient in genes, the recombination-free region differs little in gene density from flanking regions that recombine. Likewise, the density of transposable elements is similar between the recombination-free region and flanking regions. We also measured levels of histone H4 acetylation and histone H3 methylation at 176 genes within the 3.5-Mb span. Active genes showed enrichment of H4 acetylation and H3K4 dimethylation as expected, including genes within the core. Our inability to detect sequence or histone modification features that distinguish recombination-free regions from flanking regions that recombine suggest that recombination suppression is an epigenetic feature of centromeres maintained by the assembly of CENH3-containing nucleosomes within the core. CENH3-containing centrochromatin does not appear to be distinguished by a unique combination of H3 and H4 modifications. Rather, the varied distribution of histone modifications might reflect the composition and abundance of sequence elements that inhabit centromeric DNA.

Centromere↗

Comparative analyses of six solanaceous transcriptomes reveal a high degree of sequence conservation and species-specific transcripts.

BACKGROUND: The Solanaceae is a family of closely related species with diverse phenotypes that have been exploited for agronomic purposes. Previous studies involving a small number of genes suggested sequence conservation across the Solanaceae. The availability of large collections of Expressed Sequence Tags (ESTs) for the Solanaceae now provides the opportunity to assess sequence conservation and divergence on a genomic scale. RESULTS: All available ESTs and Expressed Transcripts (ETs), 449,224 sequences for six Solanaceae species (potato, tomato, pepper, petunia, tobacco and Nicotiana benthamiana), were clustered and assembled into gene indices. Examination of gene ontologies revealed that the transcripts within the gene indices encode a similar suite of biological processes. Although the ESTs and ETs were derived from a variety of tissues, 55-81% of the sequences had significant similarity at the nucleotide level with sequences among the six species. Putative orthologs could be identified for 28-58% of the sequences. This high degree of sequence conservation was supported by expression profiling using heterologous hybridizations to potato cDNA arrays that showed similar expression patterns in mature leaves for all six solanaceous species. 16-19% of the transcripts within the six Solanaceae gene indices did not have matches among Solanaceae, Arabidopsis, rice or 21 other plant gene indices. CONCLUSION: Results from this genome scale analysis confirmed a high level of sequence conservation at the nucleotide level of the coding sequence among Solanaceae. Additionally, the results indicated that part of the Solanaceae transcriptome is likely to be unique for each species.

Conserved Sequence↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

The institute for genomic research Osa1 rice genome annotation database.

We have developed a rice (Oryza sativa) genome annotation database (Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis (Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 non-transposable element-related genes, 18,545 (42.4%) were annotated with a putative function, 5,777 (13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 (44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms (5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser. All data can be obtained through the project Web pages at http://rice.tigr.org.

Computational Biology↗

Analyzing the potato abiotic stress transcriptome using expressed sequence tags.

To further increase our understanding of responses in potato to abiotic stress and the potato transcriptome in general, we generated 20 756 expressed sequence tags (ESTs) from a cDNA library constructed by pooling mRNA from heat-, cold-, salt-, and drought-stressed potato leaves and roots. These ESTs were clustered and assembled into a collection of 5240 unique sequences with 3344 contigs and 1896 singleton ESTs. Assignment of gene ontology terms (GOSlim/Plant) to the sequences revealed that 8101 assignments could be made with a total of 3863 molecular function assignments. Alignment to a set of 78 825 ESTs from other potato cDNA libraries derived from root, leaf, stolon, tuber, germinating eye, and callus tissues revealed 1476 sequences unique to abiotic stressed potato leaf and root tissue. Sequences present within the 5240 sequence set had similarity to genes known to be involved in abiotic stress responses in other plant species such as transcription factors, stress response genes, and signal transduction processes. In addition, we identified a number of genes unique to the abiotic stress library with unknown function, providing new candidate genes for investigation of abiotic stress responses in potato.

Arabidopsis↗

Structure, divergence, and distribution of the CRR centromeric retrotransposon family in rice.

The centromeric retrotransposon (CR) family in the grass species is one of few Ty3-gypsy groups of retroelements that preferentially transpose into highly specialized chromosomal domains. It has been demonstrated in both rice and maize that CRR (CR of rice) and CRM (CR of maize) elements are intermingled with centromeric satellite DNA and are highly concentrated within cytologically defined centromeres. We collected all of the CRR elements from rice chromosomes 1, 4, 8, and 10 that have been sequenced to high quality. Phylogenetic analysis revealed that the CRR elements are structurally diverged into four subfamilies, including two autonomous subfamilies (CRR1 and CRR2) and two nonautonomous subfamilies (noaCRR1 and noaCRR2). The CRR1/CRR2 elements contain all characteristic protein domains required for retrotransposition. In contrast, the noaCRR elements have different structures, containing only a gag or gag-pro domain or no open reading frames. The CRR and noaCRR elements share substantial sequence similarity in regions required for DNA replication and for recognition by integrase during retrotransposition. These data, coupled with the presence of young noaCRR elements in the rice genome and similar chromosomal distribution patterns between noaCRR1 and CRR1/CRR2 elements, suggest that the noaCRR elements were likely mobilized through the retrotransposition machinery from the autonomous CRR elements. Mechanisms of the targeting specificity of the CRR elements, as well as their role in centromere function, are discussed.

Base Sequence↗

Sequencing of a rice centromere uncovers active genes.

Centromeres are the last frontiers of complex eukaryotic genomes, consisting of highly repetitive sequences that resist mapping, cloning and sequencing. The centromere of rice Chromosome 8 (Cen8) has an unusually low abundance of highly repetitive satellite DNA, which allowed us to determine its sequence. A region of approximately 750 kb in Cen8 binds rice CENH3, the centromere-specific H3 histone. CENH3 binding is contained within a larger region that has abundant dimethylation of histone H3 at Lys9 (H3-Lys9), consistent with Cen8 being embedded in heterochromatin. Fourteen predicted and at least four active genes are interspersed in Cen8, along with CENH3 binding sites. The retrotransposons located in and outside of the CENH3 binding domain have similar ages and structural dynamics. These results suggest that Cen8 may represent an intermediate stage in the evolution of centromeres from genic regions, as in human neocentromeres, to fully mature centromeres that accumulate megabases of homogeneous satellite arrays.

Centromere↗

The TIGR Plant Repeat Databases: a collective resource for the identification of repetitive sequences in plants.

In a number of higher plants, a substantial portion of the genome is composed of repetitive sequences that can hinder genome annotation and sequencing efforts. To better understand the nature of repetitive sequences in plants and provide a resource for identifying such sequences, we constructed databases of repetitive sequences for 12 plant genera: Arabidopsis, Brassica, Glycine, Hordeum, Lotus, Lycopersicon, Medicago, Oryza, Solanum, Sorghum, Triticum and Zea (www.tigr.org/tdb/e2k1/plant. repeats/index.shtml). The repetitive sequences within each database have been coded into super-classes, classes and sub-classes based on sequence and structure similarity. These databases are available for sequence similarity searches as well as downloadable files either as entire databases or subsets of each database. To further the utility for comparative studies and to provide a resource for searching for repetitive sequences in other genera within these families, repetitive sequences have been combined into four databases to represent the Brassicaceae, Fabaceae, Gramineae and Solanaceae families. Collectively, these databases provide a resource for the identification, classification and analysis of repetitive sequences in plants.

Computational Biology↗

The TIGR rice genome annotation resource: annotating the rice genome and creating resources for plant biologists.

Rice is not only a major food staple for the world's population but it also is a model species for a major group of flowering plants, the monocotyledonous plants. Draft genomic sequence of two subspecies of rice, Oryza sativa spp. japonica and indica ssp. are publicly available. To provide the community with a resource to data-mine the rice genome, we have constructed an annotation resource for rice (http://www.tigr.org/tdb/e2k1/osa1/). In this resource, we have annotated the rice genome for gene content, identified motifs/domains within the predicted genes, constructed a rice repeat database, identified related sequences in other plant species, and identified syntenic sequences between rice and maize. All of the data is available through web-based interfaces, FTP downloads, and a Distributed Annotation System.

Chromosomes, Artificial↗

High-throughput fingerprinting of bacterial artificial chromosomes using the snapshot labeling kit and sizing of restriction fragments by capillary electrophoresis.

We have developed an automated, high-throughput fingerprinting technique for large genomic DNA fragments suitable for the construction of physical maps of large genomes. In the technique described here, BAC DNA is isolated in a 96-well plate format and simultaneously digested with four 6-bp-recognizing restriction endonucleases that generate 3' recessed ends and one 4-bp-recognizing restriction endonuclease that generates a blunt end. Each of the four recessed 3' ends is labeled with a different fluorescent dye, and restriction fragments are sized on a capillary DNA analyzer. The resulting fingerprints are edited with a fingerprint-editing computer program and contigs are assembled with the FPC computer program. The technique was evaluated by repeated fingerprinting of several BACs included as controls in plates during routine fingerprinting of a BAC library and by reconstruction of contigs of rice BAC clones with known positions on rice chromosome 10.

Chromosome Mapping↗

Molecular and cytological analyses of large tracks of centromeric DNA reveal the structure and evolutionary dynamics of maize centromeres.

We sequenced two maize bacterial artificial chromosome (BAC) clones anchored by the centromere-specific satellite repeat CentC. The two BACs, consisting of approximately 200 kb of cytologically defined centromeric DNA, are composed exclusively of satellite sequences and retrotransposons that can be classified as centromere specific or noncentromere specific on the basis of their distribution in the maize genome. Sequence analysis suggests that the original maize sequences were composed of CentC arrays that were expanded by retrotransposon invasions. Seven centromere-specific retrotransposons of maize (CRM) were found in BAC 16H10. The CRM elements inserted randomly into either CentC monomers or other retrotransposons. Sequence comparisons of the long terminal repeats (LTRs) of individual CRM elements indicated that these elements transposed within the last 1.22 million years. We observed that all of the previously reported centromere-specific retrotransposons in rice and barley, which belong to the same family as the CRM elements, also recently transposed with the oldest element having transposed approximately 3.8 million years ago. Highly conserved sequence motifs were found in the LTRs of the centromere-specific retrotransposons in the grass species, suggesting that the LTRs may be important for the centromere specificity of this retrotransposon family.

Base Sequence↗

Functional rice centromeres are marked by a satellite repeat and a centromere-specific retrotransposon.

The centromere of eukaryotic chromosomes is essential for the faithful segregation and inheritance of genetic information. In the majority of eukaryotic species, centromeres are associated with highly repetitive DNA, and as a consequence, the boundary for a functional centromere is difficult to define. In this study, we demonstrate that the centers of rice centromeres are occupied by a 155-bp satellite repeat, CentO, and a centromere-specific retrotransposon, CRR. The CentO satellite is located within the chromosomal regions to which the spindle fibers attach. CentO is quantitatively variable among the 12 rice centromeres, ranging from 65 kb to 2 Mb, and is interrupted irregularly by CRR elements. The break points of 14 rice centromere misdivision events were mapped to the middle of the CentO arrays, suggesting that the CentO satellite is located within the functional domain of rice centromeres. Our results demonstrate that the CentO satellite may be a key DNA element for rice centromere function.

Base Sequence↗

Type 1 capsule genes of Staphylococcus aureus are carried in a staphylococcal cassette chromosome genetic element.

The cap1 genes are required for the synthesis of type 1 capsular polysaccharide (CP1) in Staphylococcus aureus. We previously showed that the cap1 locus was associated with a discrete genetic element in S. aureus M. In this report, we defined the boundaries of the cap1 element by comparing its restriction pattern to that of a corresponding region from the CP1-negative strain Becker. The element was located in the SmaI-G chromosomal fragment of the standard mapping strain NCTC8325. The sequences of the entire cap1 element and the flanking regions were determined. We found that there were two additional cap1 genes not previously identified. The cap1 operon was located in a staphylococcal cassette chromosome (SCC) element similar to the resistance island SCCmec recently described for methicillin resistance in S. aureus. Notably, the SCCcap1 element was located at the same insertion site as all the SCCmec elements in the staphylococcal chromosome. The excision of SCCcap1 could be demonstrated only in the presence of the recombinase genes from an SCCmec element, verifying that SCCcap1 is a genuine SCC element but defective in mobilization. A novel enterotoxin gene, whose transcript was detected by Northern blotting, was found next to the SCCcap1 locus. We propose that the enterotoxin gene and SCCcap1 were inserted into this locus at the juxtaposition by independent events. Sequence comparison revealed numerous DNA rearrangements and mutations in SCCcap1 and the left flanking region, suggesting that the SCCcap1 had been inserted at the SCC attC site a long time ago. In addition, most genes in this region were incomplete, with the exception of the 15 cap1 genes, implying that the cap1 genes confer a survival advantage on strain M.

3' Flanking Region↗