PubMed Health⌕ Search

Biomedical subjects

Zhirong Bao

Publications and source records attributed to Zhirong Bao.

7 recordsLinked to original sources

Automated cell lineage tracing in Caenorhabditis elegans.

The invariant cell lineage and cell fate of Caenorhabditis elegans provide a unique opportunity to decode the molecular mechanisms of animal development. To exploit this opportunity, we have developed a system for automated cell lineage tracing during C. elegans embryogenesis, based on 3D, time-lapse imaging and automated image analysis. Using ubiquitously expressed histone-GFP fusion protein to label cells/nuclei and a confocal microscope, the imaging protocol captures embryogenesis at high spatial (31 planes at 1 microm apart) and temporal (every minute) resolution without apparent effects on development. A set of image analysis algorithms then automatically recognizes cells at each time point, tracks cell movements, divisions and deaths over time and assigns cell identities based on the canonical naming scheme. Starting from the four-cell stage (or earlier), our software, named starrynite, can trace the lineage up to the 350-cell stage in 25 min on a desktop computer. The few errors of automated lineaging can then be corrected in a few hours with a graphic interface that allows easy navigation of the images and the reported lineage tree. The system can be used to characterize lineage phenotypes of genes and/or extended to determine gene expression patterns in a living embryo at the single-cell level. We envision that this automation will make it practical to systematically decipher the developmental genes and pathways encoded in the genome of C. elegans.

Animals↗

Genomics in C. elegans: so many genes, such a little worm.

The Caenorhabditis elegans genome sequence is now complete, fully contiguous telomere to telomere and totaling 100,291,840 bp. The sequence has catalyzed the collection of systematic data sets and analyses, including a curated set of 19,735 protein-coding genes--with >90% directly supported by experimental evidence--and >1300 noncoding RNA genes. High-throughput efforts are under way to complete the gene sets, along with studies to characterize gene expression, function, and regulation on a genome-wide scale. The success of the worm project has had a profound effect on genome sequencing and on genomics more broadly. We now have a solid platform on which to build toward the lofty goal of a true molecular understanding of worm biology with all its implications including those for human health.

Animals↗

Pack-MULE transposable elements mediate gene evolution in plants.

Mutator-like transposable elements (MULEs) are found in many eukaryotic genomes and are especially prevalent in higher plants. In maize, rice and Arabidopsis a few MULEs were shown to carry fragments of cellular genes. These chimaeric elements are called Pack-MULEs in this study. The abundance of MULEs in rice and the availability of most of the genome sequence permitted a systematic analysis of the prevalence and nature of Pack-MULEs in an entire genome. Here we report that there are over 3,000 Pack-MULEs in rice containing fragments derived from more than 1,000 cellular genes. Pack-MULEs frequently contain fragments from multiple chromosomal loci that are fused to form new open reading frames, some of which are expressed as chimaeric transcripts. About 5% of the Pack-MULEs are represented in collections of complementary DNA. Functional analysis of amino acid sequences and proteomic data indicate that some captured gene fragments might be functional. Comparison of the cellular genes and Pack-MULE counterparts indicates that fragments of genomic DNA have been captured, rearranged and amplified over millions of years. Given the abundance of Pack-MULEs in rice and the widespread occurrence of MULEs in all characterized plant genomes, gene fragment acquisition by Pack-MULEs might represent an important new mechanism for the evolution of genes in higher plants.

Base Sequence↗

The genome sequence of Caenorhabditis briggsae: a platform for comparative genomics.

The soil nematodes Caenorhabditis briggsae and Caenorhabditis elegans diverged from a common ancestor roughly 100 million years ago and yet are almost indistinguishable by eye. They have the same chromosome number and genome sizes, and they occupy the same ecological niche. To explore the basis for this striking conservation of structure and function, we have sequenced the C. briggsae genome to a high-quality draft stage and compared it to the finished C. elegans sequence. We predict approximately 19,500 protein-coding genes in the C. briggsae genome, roughly the same as in C. elegans. Of these, 12,200 have clear C. elegans orthologs, a further 6,500 have one or more clearly detectable C. elegans homologs, and approximately 800 C. briggsae genes have no detectable matches in C. elegans. Almost all of the noncoding RNAs (ncRNAs) known are shared between the two species. The two genomes exhibit extensive colinearity, and the rate of divergence appears to be higher in the chromosomal arms than in the centers. Operons, a distinctive feature of C. elegans, are highly conserved in C. briggsae, with the arrangement of genes being preserved in 96% of cases. The difference in size between the C. briggsae (estimated at approximately 104 Mbp) and C. elegans (100.3 Mbp) genomes is almost entirely due to repetitive sequence, which accounts for 22.4% of the C. briggsae genome in contrast to 16.5% of the C. elegans genome. Few, if any, repeat families are shared, suggesting that most were acquired after the two species diverged or are undergoing rapid evolution. Coclustering the C. elegans and C. briggsae proteins reveals 2,169 protein families of two or more members. Most of these are shared between the two species, but some appear to be expanding or contracting, and there seem to be as many as several hundred novel C. briggsae gene families. The C. briggsae draft sequence will greatly improve the annotation of the C. elegans genome. Based on similarity to C. briggsae, we found strong evidence for 1,300 new C. elegans genes. In addition, comparisons of the two genomes will help to understand the evolutionary forces that mold nematode genomes.

Animals↗

An active DNA transposon family in rice.

The publication of draft sequences for the two subspecies of Oryza sativa (rice), japonica (cv. Nipponbare) and indica (cv. 93-11), provides a unique opportunity to study the dynamics of transposable elements in this important crop plant. Here we report the use of these sequences in a computational approach to identify the first active DNA transposons from rice and the first active miniature inverted-repeat transposable element (MITE) from any organism. A sequence classified as a Tourist-like MITE of 430 base pairs, called miniature Ping (mPing), was present in about 70 copies in Nipponbare and in about 14 copies in 93-11. These mPing elements, which are all nearly identical, transpose actively in an indica cell-culture line. Database searches identified a family of related transposase-encoding elements (called Pong), which also transpose actively in the same cells. Virtually all new insertions of mPing and Pong elements were into low-copy regions of the rice genome. Since the domestication of rice mPing MITEs have been amplified preferentially in cultivars adapted to environmental extremes-a situation that is reminiscent of the genomic shock theory for transposon activation.

Amino Acid Sequence↗

Dasheng: a recently amplified nonautonomous long terminal repeat element that is a major component of pericentromeric regions in rice.

A new and unusual family of LTR elements, Dasheng, has been discovered in the genome of Oryza sativa following database searches of approximately 100 Mb of rice genomic sequence and 78 Mb of BAC-end sequence information. With all of the cis-elements but none of the coding domains normally associated with retrotransposons (e.g., gag, pol), Dasheng is a novel nonautonomous LTR element with high copy number. Over half of the approximately 1000 Dasheng elements in the rice genome are full length (5.6-8.6 kb), and 60% are estimated to have amplified in the past 500,000 years. Using a modified AFLP technique called transposon display, 215 elements were mapped to all 12 rice chromosomes. Interestingly, more than half of the mapped elements are clustered in the heterochromatic regions around centromeres. The distribution pattern was further confirmed by FISH analysis. Despite clustering in heterochromatin, Dasheng elements are not nested, suggesting their potential value as molecular markers for these marker-poor regions. Taken together, Dasheng is one of the highest-copy-number LTR elements and one of the most recent elements to amplify in the rice genome.

Centromere↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗