PubMed Health⌕ Search

Biomedical subjects

Yeisoo Yu

Publications and source records attributed to Yeisoo Yu.

15 recordsLinked to original sources

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

A global assembly of cotton ESTs.

Approximately 185,000 Gossypium EST sequences comprising >94,800,000 nucleotides were amassed from 30 cDNA libraries constructed from a variety of tissues and organs under a range of conditions, including drought stress and pathogen challenges. These libraries were derived from allopolyploid cotton (Gossypium hirsutum; A(T) and D(T) genomes) as well as its two diploid progenitors, Gossypium arboreum (A genome) and Gossypium raimondii (D genome). ESTs were assembled using the Program for Assembling and Viewing ESTs (PAVE), resulting in 22,030 contigs and 29,077 singletons (51,107 unigenes). Further comparisons among the singletons and contigs led to recognition of 33,665 exemplar sequences that represent a nonredundant set of putative Gossypium genes containing partial or full-length coding regions and usually one or two UTRs. The assembly, along with their UniProt BLASTX hits, GO annotation, and Pfam analysis results, are freely accessible as a public resource for cotton genomics. Because ESTs from diploid and allotetraploid Gossypium were combined in a single assembly, we were in many cases able to bioinformatically distinguish duplicated genes in allotetraploid cotton and assign them to either the A or D genome. The assembly and associated information provide a framework for future investigation of cotton functional and evolutionary genomics.

DNA, Complementary↗

Utilization of a zebra finch BAC library to determine the structure of an avian androgen receptor genomic region.

The zebra finch (Taeniopygia guttata) is an important model organism for studying behavior, neuroscience, avian biology, and evolution. To support the study of its genome, we constructed a BAC library (TG__Ba) using DNA from livers of females. The BAC library consists of 147,456 clones with 98% containing inserts of an average size of 134 kb and represents 15.5 haploid genome equivalents. By sequencing a whole BAC, a full-length androgen receptor open reading frame was identified, the first in an avian species. Comparison of BAC end sequences and the whole BAC sequence with the chicken genome draft sequence showed a high degree of conserved synteny between the zebra finch and the chicken genome.

Animals↗

The Oryza bacterial artificial chromosome library resource: construction and analysis of 12 deep-coverage large-insert BAC libraries that represent the 10 genome types of the genus Oryza.

Rice (Oryza sativa L.) is the most important food crop in the world and a model system for plant biology. With the completion of a finished genome sequence we must now functionally characterize the rice genome by a variety of methods, including comparative genomic analysis between cereal species and within the genus Oryza. Oryza contains two cultivated and 22 wild species that represent 10 distinct genome types. The wild species contain an essentially untapped reservoir of agriculturally important genes that must be harnessed if we are to maintain a safe and secure food supply for the 21st century. As a first step to functionally characterize the rice genome from a comparative standpoint, we report the construction and analysis of a comprehensive set of 12 BAC libraries that represent the 10 genome types of Oryza. To estimate the number of clones required to generate 10 genome equivalent BAC libraries we determined the genome sizes of nine of the 12 species using flow cytometry. Each library represents a minimum of 10 genome equivalents, has an average insert size range between 123 and 161 kb, an average organellar content of 0.4%-4.1% and nonrecombinant content between 0% and 5%. Genome coverage was estimated mathematically and empirically by hybridization and extensive contig and BAC end sequence analysis. A preliminary analysis of BAC end sequences of clones from these libraries indicated that LTR retrotransposons are the predominant class of repeat elements in Oryza and a roughly linear relationship of these elements with genome size was observed.

Base Sequence↗

Random sheared fosmid library as a new genomic tool to accelerate complete finishing of rice (Oryza sativa spp. Nipponbare) genome sequence: sequencing of gap-specific fosmid clones uncovers new euchromatic portions of the genome.

The International Rice Genome Sequencing Project has recently announced the high-quality finished sequence that covers nearly 95% of the japonica rice genome representing 370 Mbp. Nevertheless, the current physical map of japonica rice contains 62 physical gaps corresponding to approximately 5% of the genome, that have not been identified/represented in the comprehensive array of publicly available BAC, PAC and other genomic library resources. Without finishing these gaps, it is impossible to identify the complete complement of genes encoded by rice genome and will also leave us ignorant of some 5% of the genome and its unknown functions. In this article, we report the construction and characterization of a tenfold redundant, 40 kbp insert fosmid library generated by random mechanical shearing. We demonstrated its utility in refining the physical map of rice by identifying and in silico mapping 22 gap-specific fosmid clones with particular emphasis on chromosomes 1, 2, 6, 7, 8, 9 and 10. Further sequencing of 12 of the gap-specific fosmid clones uncovered unique rice genome sequence that was not previously reported in the finished IRGSP sequence and emphasizes the need to complete finishing of the rice genome.

Base Sequence↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

Toward closing rice telomere gaps: mapping and sequence characterization of rice subtelomere regions.

Despite the collective efforts of the international community to sequence the complete rice genome, telomeric regions of most chromosome arms remain uncharacterized. In this report we present sequence data from subtelomere regions obtained by analyzing telomeric clones from two 8.8 x genome equivalent 10-kb libraries derived from partial restriction digestion with HaeIII or Sau3AI (OSJNPb HaeIII and OSJNPc Sau3AI). Seven telomere clones were identified and contain 25-100 copies of the telomere repeat (CCCTAAA)(n) on one end and unique sequences on the opposite end. Polymorphic sequence-tagged site markers from five clones and one additional PCR product were genetically mapped on the ends of chromosome arms 2S, 5L, 10S, 10L, 7L, and 7S. We found distinct chromosome-specific telomere-associated tandem repeats (TATR) on chromosome 7 (TATR7) and on the short arm of chromosome 10 (TATR10s) that showed no significant homology to any International Rice Genome Sequencing Project (IRGSP) genomic sequence. The TATR7, a degenerate tandem repeat which is interrupted by transposable elements, appeared on both ends of chromosome 7. The TATR10s was found to contain an inverted array of three tandem repeats displaying an interesting secondary folding pattern that resembles a telomere loop (t-loop) and which may be involved in a protective function against chromosomal end degradation.

Base Sequence↗

In-depth sequence analysis of the tomato chromosome 12 centromeric region: identification of a large CAA block and characterization of pericentromere retrotranposons.

We sequenced a continuous 326-kb DNA stretch of a microscopically defined centromeric region of tomato chromosome 12. A total of 84% of the sequence (270 kb) was composed of a nested complex of repeat sequences including 27 retrotransposons, two transposable elements, three MITEs, two terminal repeat retrotransposons in miniature (TRIMs), ten unclassified repeats and three chloroplast DNA insertions. The retrotransposons were grouped into three families of Ty3-Gypsy type long terminal repeat (LTR) retrotransposons (PCRT1-PCRT3) and one LINE-like retrotransposon (PCRT4). High-resolution fluorescence in situ hybridization analyses on pachytene complements revealed that PCRT1a occurs on the pericentromere heterochromatin blocks. PCRT1 was the prevalent retrotransposon family occupying more than 60% of the 326-kb sequence with 19 members grouped into eight subfamilies (PCRT1a-PCRT1h) based on LTR sequence. The PCRT1a subfamily is a rapidly amplified element occupying tens of megabases. The other PCRT1 subfamilies (PCRT1b-PCRT1h) were highly degenerated and interrupted by insertions of other elements. The PCRT1 family shows identity with a previously identified tomato-specific repeat TGR2 and a CENP-B like sequence. A second previously described genomic repeat, TGR3, was identified as a part of the LTR sequence of an Athila-like PCRT2 element of which four copies were found in the 326-kb stretch. A large block of trinucleotide microsatellite (CAA)n occupies the centromere and large portions of the flanking pericentromere heterochromatin blocks of chromosome 12 and most of the other chromosomes. Five putative genes in the remaining 14% of the centromere region were identified, of which one is similar to a transcription regulator (ToCPL1) and a candidate jointless-2 gene.

Base Sequence↗

Candidate gene database and transcript map for peach, a model species for fruit trees.

Peach (Prunus persica) is a model species for the Rosaceae, which includes a number of economically important fruit tree species. To develop an extensive Prunus expressed sequence tag (EST) database for identifying and cloning the genes important to fruit and tree development, we generated 9,984 high-quality ESTs from a peach cDNA library of developing fruit mesocarp. After assembly and annotation, a putative peach unigene set consisting of 3,842 ESTs was defined. Gene ontology (GO) classification was assigned based on the annotation of the single "best hit" match against the Swiss-Prot database. No significant homology could be found in the GenBank nr databases for 24.3% of the sequences. Using core markers from the general Prunus genetic map, we anchored bacterial artificial chromosome (BAC) clones on the genetic map, thereby providing a framework for the construction of a physical and transcript map. A transcript map was developed by hybridizing 1,236 ESTs from the putative peach unigene set and an additional 68 peach cDNA clones against the peach BAC library. Hybridizing ESTs to genetically anchored BACs immediately localized 11.2% of the ESTs on the genetic map. ESTs showed a clustering of expressed genes in defined regions of the linkage groups. [The data were built into a regularly updated Genome Database for Rosaceae (GDR), available at (http://www.genome.clemson.edu/gdr/).].

Breeding↗

The oryza map alignment project: the golden path to unlocking the genetic potential of wild rice species.

The wild species of the genus Oryza offer enormous potential to make a significant impact on agricultural productivity of the cultivated rice species Oryza sativa and Oryza glaberrima. To unlock the genetic potential of wild rice we have initiated a project entitled the 'Oryza Map Alignment Project' (OMAP) with the ultimate goal of constructing and aligning BAC/STC based physical maps of 11 wild and one cultivated rice species to the International Rice Genome Sequencing Project's finished reference genome--O. sativa ssp. japonica c. v. Nipponbare. The 11 wild rice species comprise nine different genome types and include six diploid genomes (AA, BB, CC, EE, FF and GG) and four tetrapliod genomes (BBCC, CCDD, HHKK and HHJJ) with broad geographical distribution and ecological adaptation. In this paper we describe our strategy to construct robust physical maps of all 12 rice species with an emphasis on the AA diploid O. nivara--thought to be the progenitor of modern cultivated rice.

Chromosome Mapping↗

Large-scale identification of expressed sequence tags involved in rice and rice blast fungus interaction.

To better understand the molecular basis of the defense response against the rice blast fungus (Magnaporthe grisea), a large-scale expressed sequence tag (EST) sequencing approach was used to identify genes involved in the early infection stages in rice (Oryza sativa). Six cDNA libraries were constructed using infected leaf tissues harvested from 6 conditions: resistant, partially resistant, and susceptible reactions at both 6 and 24 h after inoculation. Two additional libraries were constructed using uninoculated leaves and leaves from the lesion mimic mutant spl11. A total of 68,920 ESTs were generated from 8 libraries. Clustering and assembly analyses resulted in 13,570 unique sequences from 10,934 contigs and 2,636 singletons. Gene function classification showed that 42% of the ESTs were predicted to have putative gene function. Comparison of the pathogen-challenged libraries with the uninoculated control library revealed an increase in the percentage of genes in the functional categories of defense and signal transduction mechanisms and cell cycle control, cell division, and chromosome partitioning. In addition, hierarchical clustering analysis grouped the eight libraries based on their disease reactions. A total of 7,748 new and unique ESTs were identified from our collection compared with the KOME full-length cDNA collection. Interestingly, we found that rice ESTs are more closely related to sorghum (Sorghum bicolor) ESTs than to barley (Hordeum vulgare), wheat (Triticum aestivum), and maize (Zea mays) ESTs. The large cataloged collection of rice ESTs in this study provides a solid foundation for further characterization of the rice defense response and is a useful public genomic resource for rice functional genomics studies.

Disease Susceptibility↗

Sequence composition and genome organization of maize.

Zea mays L. ssp. mays, or corn, one of the most important crops and a model for plant genetics, has a genome approximately 80% the size of the human genome. To gain global insight into the organization of its genome, we have sequenced the ends of large insert clones, yielding a cumulative length of one-eighth of the genome with a DNA sequence read every 6.2 kb, thereby describing a large percentage of the genes and transposable elements of maize in an unbiased approach. Based on the accumulative 307 Mb of sequence, repeat sequences occupy 58% and genic regions occupy 7.5%. A conservative estimate predicts approximately 59,000 genes, which is higher than in any other organism sequenced so far. Because the sequences are derived from bacterial artificial chromosome clones, which are ordered in overlapping bins, tagged genes are also ordered along continuous chromosomal segments. Based on this positional information, roughly one-third of the genes appear to consist of tandemly arrayed gene families. Although the ancestor of maize arose by tetraploidization, fewer than half of the genes appear to be present in two orthologous copies, indicating that the maize genome has undergone significant gene loss since the duplication event.

Chromosomes, Artificial, Bacterial↗

Construction and utility of 10-kb libraries for efficient clone-gap closure for rice genome sequencing.

Rice is an important crop and a model system for monocot genomics, and is a target for whole genome sequencing by the International Rice Genome Sequencing Project (IRGSP). The IRGSP is using a clone by clone approach to sequence rice based on minimum tiles of BAC or PAC clones. For chromosomes 10 and 3 we are using an integrated physical map based on two fingerprinted and end-sequenced BAC libraries to identifying a minimum tiling path of clones. In this study we constructed and tested two rice genomic libraries with an average insert size of 10 kb (10-kb library) to support the gap closure and finishing phases of the rice genome sequencing project. The HaeIII library contains 166,752 clones covering approximately 4.6x rice genome equivalents with an average insert size of 10.5 kb. The Sau3AI library contains 138,960 clones covering 4.2x genome equivalents with an average insert size of 11.6 kb. Both libraries were gridded in duplicate onto 11 high-density filters in a 5 x 5 pattern to facilitate screening by hybridization. The libraries contain an unbiased coverage of the rice genome with less than 5% contamination by clones containing organelle DNA or no insert. An efficient method was developed, consisting of pooled overgo hybridization, the selection of 10-kb gap spanning clones using end sequences, transposon sequencing and utilization of in silico draft sequence, to close relatively small gaps between sequenced BAC clones. Using this method we were able to close a majority of the gaps (up to approximately 50 kb) identified during the finishing phase of chromosome-10 sequencing. This method represents a useful way to close clone gaps and thus to complete the entire rice genome.

Chromosomes, Artificial, Bacterial↗

A draft sequence of the rice genome (Oryza sativa L. ssp. japonica).

The genome of the japonica subspecies of rice, an important cereal and model monocot, was sequenced and assembled by whole-genome shotgun sequencing. The assembled sequence covers 93% of the 420-megabase genome. Gene predictions on the assembled sequence suggest that the genome contains 32,000 to 50,000 genes. Homologs of 98% of the known maize, wheat, and barley proteins are found in rice. Synteny and gene homology between rice and the other cereal genomes are extensive, whereas synteny with Arabidopsis is limited. Assignment of candidate rice orthologs to Arabidopsis genes is possible in many cases. The rice genome sequence provides a foundation for the improvement of cereals, our most important crops.

Arabidopsis↗

An integrated physical and genetic map of the rice genome.

Rice was chosen as a model organism for genome sequencing because of its economic importance, small genome size, and syntenic relationship with other cereal species. We have constructed a bacterial artificial chromosome fingerprint-based physical map of the rice genome to facilitate the whole-genome sequencing of rice. Most of the rice genome ( approximately 90.6%) was anchored genetically by overgo hybridization, DNA gel blot hybridization, and in silico anchoring. Genome sequencing data also were integrated into the rice physical map. Comparison of the genetic and physical maps reveals that recombination is suppressed severely in centromeric regions as well as on the short arms of chromosomes 4 and 10. This integrated high-resolution physical map of the rice genome will greatly facilitate whole-genome sequencing by helping to identify a minimum tiling path of clones to sequence. Furthermore, the physical map will aid map-based cloning of agronomically important genes and will provide an important tool for the comparative analysis of grass genomes.

Chromosomes, Artificial, Bacterial↗