PubMed Health⌕ Search

Biomedical subjects

George S Yang

Publications and source records attributed to George S Yang.

5 recordsLinked to original sources

Sequencing and analysis of 10,967 full-length cDNA clones from Xenopus laevis and Xenopus tropicalis reveals post-tetraploidization transcriptome remodeling.

Sequencing of full-insert clones from full-length cDNA libraries from both Xenopus laevis and Xenopus tropicalis has been ongoing as part of the Xenopus Gene Collection Initiative. Here we present 10,967 full ORF verified cDNA clones (8049 from X. laevis and 2918 from X. tropicalis) as a community resource. Because the genome of X. laevis, but not X. tropicalis, has undergone allotetraploidization, comparison of coding sequences from these two clawed (pipid) frogs provides a unique angle for exploring the molecular evolution of duplicate genes. Within our clone set, we have identified 445 gene trios, each comprised of an allotetraploidization-derived X. laevis gene pair and their shared X. tropicalis ortholog. Pairwise dN/dS, comparisons within trios show strong evidence for purifying selection acting on all three members. However, dN/dS ratios between X. laevis gene pairs are elevated relative to their X. tropicalis ortholog. This difference is highly significant and indicates an overall relaxation of selective pressures on duplicated gene pairs. We have found that the paralogs that have been lost since the tetraploidization event are enriched for several molecular functions, but have found no such enrichment in the extant paralogs. Approximately 14% of the paralogous pairs analyzed here also show differential expression indicative of subfunctionalization.

Animals↗

Analysis of long-lived C. elegans daf-2 mutants using serial analysis of gene expression.

We have identified longevity-associated genes in a long-lived Caenorhabditis elegans daf-2 (insulin/IGF receptor) mutant using serial analysis of gene expression (SAGE), a method that efficiently quantifies large numbers of mRNA transcripts by sequencing short tags. Reduction of daf-2 signaling in these mutant worms leads to a doubling in mean lifespan. We prepared C. elegans SAGE libraries from 1, 6, and 10-d-old adult daf-2 and from 1 and 6-d-old control adults. Differences in gene expression between daf-2 libraries representing different ages and between daf-2 versus control libraries identified not only single genes, but whole gene families that were differentially regulated. These gene families are part of major metabolic pathways including lipid, protein, and energy metabolism, stress response, and cell structure. Similar expression patterns of closely related family members emphasize the importance of these genes in aging-related processes. Global analysis of metabolism-associated genes showed hypometabolic features in mid-life daf-2 mutants that diminish with advanced age. Comparison of our results to recent microarray studies highlights sets of overlapping genes that are highly conserved throughout evolution and thus represent strong candidate genes that control aging and longevity.

Age Factors↗

High-throughput sequencing: a failure mode analysis.

BACKGROUND: Basic manufacturing principles are becoming increasingly important in high-throughput sequencing facilities where there is a constant drive to increase quality, increase efficiency, and decrease operating costs. While high-throughput centres report failure rates typically on the order of 10%, the causes of sporadic sequencing failures are seldom analyzed in detail and have not, in the past, been formally reported. RESULTS: Here we report the results of a failure mode analysis of our production sequencing facility based on detailed evaluation of 9,216 ESTs generated from two cDNA libraries. Two categories of failures are described; process-related failures (failures due to equipment or sample handling) and template-related failures (failures that are revealed by close inspection of electropherograms and are likely due to properties of the template DNA sequence itself). CONCLUSIONS: Preventative action based on a detailed understanding of failure modes is likely to improve the performance of other production sequencing pipelines.

Automation↗

The Genome sequence of the SARS-associated coronavirus.

We sequenced the 29,751-base genome of the severe acute respiratory syndrome (SARS)-associated coronavirus known as the Tor2 isolate. The genome sequence reveals that this coronavirus is only moderately related to other known coronaviruses, including two human coronaviruses, HCoV-OC43 and HCoV-229E. Phylogenetic analysis of the predicted viral proteins indicates that the virus does not closely resemble any of the three previously known groups of coronaviruses. The genome sequence will aid in the diagnosis of SARS virus infection in humans and potential animal hosts (using polymerase chain reaction and immunological tests), in the development of antivirals (including neutralizing antibodies), and in the identification of putative epitopes for vaccine development.

3' Untranslated Regions↗

An efficient strategy for large-scale high-throughput transposon-mediated sequencing of cDNA clones.

We describe an efficient high-throughput method for accurate DNA sequencing of entire cDNA clones. Developed as part of our involvement in the Mammalian Gene Collection full-length cDNA sequencing initiative, the method has been used and refined in our laboratory since September 2000. Amenable to large scale projects, we have used the method to generate >7 Mb of accurate sequence from 3695 candidate full-length cDNAs. Sequencing is accomplished through the insertion of Mu transposon into cDNAs, followed by sequencing reactions primed with Mu-specific sequencing primers. Transposon insertion reactions are not performed with individual cDNAs but rather on pools of up to 96 clones. This pooling strategy reduces the number of transposon insertion sequencing libraries that would otherwise be required, reducing the costs and enhancing the efficiency of the transposon library construction procedure. Sequences generated using transposon-specific sequencing primers are assembled to yield the full-length cDNA sequence, with sequence editing and other sequence finishing activities performed as required to resolve sequence ambiguities. Although analysis of the many thousands (22 785) of sequenced Mu transposon insertion events revealed a weak sequence preference for Mu insertion, we observed insertion of the Mu transposon into 1015 of the possible 1024 5mer candidate insertion sites.

Bacteriophage mu↗