PubMed Health⌕ Search

Biomedical subjects

Terry Clark

Publications and source records attributed to Terry Clark.

7 recordsLinked to original sources

Serial analysis of gene expression study of a hybrid rice strain (LYP9) and its parental cultivars.

Using the serial analysis of gene expression technique, we surveyed transcriptomes of three major tissues (panicles, leaves, and roots) of a super-hybrid rice (Oryza sativa) strain, LYP9, in comparison to its parental cultivars, 93-11 (indica) and PA64s (japonica). We acquired 465,679 tags from the serial analysis of gene expression libraries, which were consolidated into 68,483 unique tags. Focusing our initial functional analyses on a subset of the data that are supported by full-length cDNAs and the tags (genes) differentially expressed in the hybrid at a significant level (P<0.01), we identified 595 up-regulated (22 tags in panicles, 228 in leaves, and 345 in roots) and 25 down-regulated (seven tags in panicles, 15 in leaves, and three in roots) in LYP9. Most of the tag-identified and up-regulated genes were found related to enhancing carbon- and nitrogen-assimilation, including photosynthesis in leaves, nitrogen uptake in roots, and rapid growth in both roots and panicles. Among the down-regulated genes in LYP9, there is an essential enzyme in photorespiration, alanine:glyoxylate aminotransferase 1. Our study adds a new set of data crucial for the understanding of molecular mechanisms of heterosis and gene regulation networks of the cultivated rice.

Chromosome Mapping↗

Detecting novel low-abundant transcripts in Drosophila.

Increasing evidence suggests that low-abundant transcripts may play fundamental roles in biological processes. In an attempt to estimate the prevalence of low-abundant transcripts in eukaryotic genomes, we performed a transcriptome analysis in Drosophila using the SAGE technique. We collected 244,313 SAGE tags from transcripts expressed in Drosophila embryonic, larval, pupae, adult, and testicular tissue. From these SAGE tags, we identified 40,823 unique SAGE tags. Our analysis showed that 55% of the 40,823 unique SAGE tags are novel without matches in currently known Drosophila transcripts, and most of the novel SAGE tags have low copy numbers. Further analysis indicated that these novel SAGE tags represent novel low-abundant transcripts expressed from loci outside of currently annotated exons including the intergenic and intronic regions, and antisense of the currently annotated exons in the Drosophila genome. Our study reveals the presence of a significant number of novel low-abundant transcripts in Drosophila, and highlights the need to isolate these novel low-abundant transcripts for further biological studies.

Animals↗

A structured interface to the object-oriented genomics unified schema for XML-formatted data.

Data management systems are fast becoming required components in many biology laboratories as the role of computer-based information grows. Although the need for data management systems is on the rise, their inherent complexities can deter the full and routine use of their computational capabilities. The significant undertaking to implement a capable production system can be reduced in part by adapting an established data management system. In such a way, we are leveraging the Genomics Unified Schema (GUS) developed at the Computational Biology and Informatics Laboratory at the University of Pennsylvania as a foundation for managing and analysing DNA sequence data in centromere research projects around Arabidopsis thaliana and related species. Because GUS provides a core schema that includes support for genome sequences, mRNA and its expression, and annotated chromosomes, it is ideal for synthesising a variety of parameters to analyse these repetitive and highly dynamic portions of the genome. Despite this, production-strength data management frameworks are complex, requiring dedicated efforts to adapt and maintain. The work reported in this article addresses one component of such an effort, namely the pivotal task of marshalling data from various sources into GUS. In order to harness GUS for our project, and motivated by efficiency needs, we developed a structured framework for transferring data into GUS from outside sources. This technology is embodied in a GUS object-layer processor, XMLGUS. XMLGUS facilitates incorporating data into GUS by (i) formulating an XML interface that includes relational database key constraint definitions, (ii) regularising traversal through that XML, (iii) realising automatic processing of the XML with database key constraints and (iv) allowing for special processing of input data within the framework for automated processing. The application of XMLGUS to production pipeline processing for a sequencing project and inputting the Arabidopsis genome into GUS is discussed. XMLGUS is available from the Flora website (http://flora.ittc.ku.edu/).

Chromosome Mapping↗

Oligo(dT) primer generates a high frequency of truncated cDNAs through internal poly(A) priming during reverse transcription.

We have analyzed a systematic flaw in the current system of gene identification: the oligo(dT) primer widely used for cDNA synthesis generates a high frequency of truncated cDNAs through internal poly(A) priming. Such truncated cDNAs may contribute to 12% of the expressed sequence tags in the current dbEST database. By using a synthetic transcript and real mRNA templates as models, we characterized the patterns of internal poly(A) priming by oligo(dT) primer. We further demonstrated that the internal poly(A) priming can be effectively diminished by replacing the oligo(dT) primer with a set of anchored oligo(dT) primers for reverse transcription. Our study indicates that cDNAs designed for genomewide gene identification should be synthesized by use of the anchored oligo(dT) primers, rather than the oligo(dT) primers, to diminish the generation of truncated cDNAs caused by internal poly(A) priming.

Animals↗

Correct identification of genes from serial analysis of gene expression tag sequences.

SAGE (serial analysis of gene expression) is a remarkable technique for genome-wide analysis of gene expression. It is crucial to understand the extent to which SAGE can accurately indicate a gene or expressed sequence tag (EST) with a single tag. We analyzed the effect of the size of SAGE tag on gene identification. Our observation indicates that SAGE tags are in general not long enough to achieve the degree of uniqueness of identification originally envisaged. Our observations also indicate that the limitation of using SAGE tag to identify a gene can be overcome by converting SAGE tags into longer 3' EST sequences with the generation of longer cDNA fragments from SAGE tages for gene identification (GLGI) method.

Expressed Sequence Tags↗

Computational Analysis of Gene Identification with SAGE.

SAGE is one of the few techniques capable of uniformly probing gene expression at a genome level irrespective of mRNA abundance and without a priori knowledge of the transcripts present. However, individual SAGE tags can match many sequences in the reference database, complicating gene identification. We perform a baseline evaluation of gene identification with SAGE using UniGene Human as the reference database by analyzing 1) the distributions of tags for various length tag sets formed for UniGene Human and 2) the tag-to-sequence mapping using a SAGE tag set consisting of 37,522 tags derived from human myeloid cells. The extensive multiplicity of the dbEST component of UniGene significantly detracts from gains that might be expected by extending tags within the scope of the SAGE protocol. In order to achieve reasonable sequence specificity for gene identification with the content of the commonly used UniGene sequence collection, tags on the order of hundreds of bases in length are required. One way to produce tags of such lengths is with GLGI, which extends SAGE tags to the 3' end of cDNA. We show that the longer sequences produced by GLGI relieve significantly the multiple match condition. In the myeloid sample, we also found a correlation between multiple match severity and high copy number. We extrapolate these findings, providing insights into the use of UniGene Human as a reference for gene identification.

Animals↗