PubMed Health⌕ Search

Biomedical subjects

Jonathan Schug

Publications and source records attributed to Jonathan Schug.

10 recordsLinked to original sources

EPConDB: a web resource for gene expression related to pancreatic development, beta-cell function and diabetes.

EPConDB (http://www.cbil.upenn.edu/EPConDB) is a public web site that supports research in diabetes, pancreatic development and beta-cell function by providing information about genes expressed in cells of the pancreas. EPConDB displays expression profiles for individual genes and information about transcripts, promoter elements and transcription factor binding sites. Gene expression results are obtained from studies examining tissue expression, pancreatic development and growth, differentiation of insulin-producing cells, islet or beta-cell injury, and genetic models of impaired beta-cell function. The expression datasets are derived using different microarray platforms, including the BCBC PancChips and Affymetrix gene expression arrays. Other datasets include semi-quantitative RT-PCR and MPSS expression studies. For selected microarray studies, lists of differentially expressed genes, derived from PaGE analysis, are displayed on the site. EPConDB provides database queries and tools to examine the relationship between a gene, its transcriptional regulation, protein function and expression in pancreatic tissues.

Animals↗

Glucocorticoid receptor-dependent gene regulatory networks.

While the molecular mechanisms of glucocorticoid regulation of transcription have been studied in detail, the global networks regulated by the glucocorticoid receptor (GR) remain unknown. To address this question, we performed an orthogonal analysis to identify direct targets of the GR. First, we analyzed the expression profile of mouse livers in the presence or absence of exogenous glucocorticoid, resulting in over 1,300 differentially expressed genes. We then executed genome-wide location analysis on chromatin from the same livers, identifying more than 300 promoters that are bound by the GR. Intersecting the two lists yielded 53 genes whose expression is functionally dependent upon the ligand-bound GR. Further network and sequence analysis of the functional targets enabled us to suggest interactions between the GR and other transcription factors at specific target genes. Together, our results further our understanding of the GR and its targets, and provide the basis for more targeted glucocorticoid therapies.

Animals↗

Promoter features related to tissue specificity as measured by Shannon entropy.

BACKGROUND: The regulatory mechanisms underlying tissue specificity are a crucial part of the development and maintenance of multicellular organisms. A genome-wide analysis of promoters in the context of gene-expression patterns in tissue surveys provides a means of identifying the general principles for these mechanisms. RESULTS: We introduce a definition of tissue specificity based on Shannon entropy to rank human genes according to their overall tissue specificity and by their specificity to particular tissues. We apply our definition to microarray-based and expressed sequence tag (EST)-based expression data for human genes and use similar data for mouse genes to validate our results. We show that most genes show statistically significant tissue-dependent variations in expression level. We find that the most tissue-specific genes typically have a TATA box, no CpG island, and often code for extracellular proteins. As expected, CpG islands are found in most of the least tissue-specific genes, which often code for proteins located in the nucleus or mitochondrion. The class of genes with no CpG island or TATA box are the most common mid-specificity genes and commonly code for proteins located in a membrane. Sp1 was found to be a weak indicator of less-specific expression. YY1 binding sites, either as initiators or as downstream sites, were strongly associated with the least-specific genes. CONCLUSIONS: We have begun to understand the components of promoters that distinguish tissue-specific from ubiquitous genes, to identify associations that can predict the broad class of gene expression from sequence data alone.

Animals↗

Orthogonal analysis of C/EBPbeta targets in vivo during liver proliferation.

CCAAT enhancer-binding protein beta (C/EBPbeta), a basic-leucine zipper transcription factor, is an important effector of signals in physiologic growth and cancer. The identification of direct C/EBPbeta targets in vivo has been limited by functional compensation by other C/EBP family proteins and the low stringency of the consensus sequence. Here we use the combined power of expression profiling and high-throughput chromatin immunoprecipitation to identify direct and biologically relevant targets of C/EBPbeta. We identified 25 potential C/EBPbeta targets, of which 88% of those tested were confirmed as in vivo C/EBPbeta-binding sites. Six of these genes also displayed differential expression in C/EBPbeta-/- livers. Computational analysis revealed that bona fide C/EBPbeta target genes can be distinguished by the presence of binding motifs for specific additional transcription factors in the vicinity of the C/EBPbeta site. This approach is generally applicable to the discovery of direct, biologically relevant targets of mammalian transcription factors.

Animals↗

Widespread distribution of antisense transcripts in the Plasmodium falciparum genome.

The availability of the complete genome sequence of Plasmodium falciparum has facilitated high-throughput profiling of its complex life cycle, following the application of micro-array, proteomic, and serial analysis of gene expression (SAGE) technologies in this system. These, in turn, have yielded unprecedented insight into global gene expression, including the foremost demonstration of antisense transcription in the parasite. For example, owing to its inherent ability to sample novel ORFs and to predict transcript orientation, SAGE analysis in asexual forms led to the initial discovery of highly abundant antisense RNAs. To determine the extent of this phenomenon in P. falciparum, we have surveyed the distribution of both sense and antisense transcripts across the asexual transcriptome for the first time. To this end, a relational database integrating SAGE expression data with genome annotation information was constructed. This allowed the comprehensive annotation of a total of 17245 SAGE tags, extending over a 350-fold expression range. Transcripts from approximately 30% of the estimated 3D7 gene loci were present at detectable levels in mixed asexual stages, where loci involved in invasion and immune evasion; and carbohydrate metabolism were highly represented in the sense transcriptome. Approximately 12% of SAGE tags, however, were derived from the non-coding strand of nuclear-encoded ORFs, indicating that endogenous antisense RNAs are widespread in this system. Notably, these antisense transcripts were absent from the mitochondrial genome. Interestingly, we note that sense and antisense tag counts from single loci across the transcriptome were inversely related. Taken together, this data may provide first hints as to the possible function of antisense transcription in this system.

Animals↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

Drug-induced alterations in gene expression of the asexual blood forms of Plasmodium falciparum.

The complete genome sequence of the malarial parasite, Plasmodium falciparum [Gardner, M.J., Hall, N., Fung, E., White, O., Berriman, M., and Hyman, R.W. (2002) Nature 419: 498-510; Hyman, R.W., Fung, E., Conway, A., Kurdi, O., Mao, J., Miranda, M. et al. (2002) Nature 419: 534-537], has provided researchers with the informational base for establishing genomic [Volkman, S.K., Hartl, D.L., Wirth, D.F., Nielsen, K.M., Choi, M., Batalov, S., et al. (2002) Science 298: 216-218], proteomic [Florens, L., Washburn, M.P.J.D.R., Anthony, R.M., Grainger, M., Haynes, J.D., et al. (2002) Nature 419: 520-526; Lasonder, E., Ishihama, Y., Anderson, J.S., Vermunt, A.M.W., Pain, A., Sauerwein, R.W., et al. (2002) Nature 419: 537-542] and genome-wide RNA expression [Ben Mamoun, C., Gluzman, I.Y., Hott, C., MacMillan, S.K., Amarakone, A.S., Anderson, D.L., et al. (2001) Mol Microbiol 39: 26-36; Hayward, R.E., Derisi, J.L., Alfadhli, S., Kaslow, D.C., Brown, P.O., and Rathod, P.K. (2000) Mol Microbiol 35: 6-14] analyses in this system. In fact, we have previously utilized SAGE (serial analysis of gene expression) to identify abundant loci that probably constitute part of the active metabolome [Patankar, S., Munasinghe, A., Shoaibi, A., Cummings, L.M., and Wirth, D.F. (2001) Mol Biol Cell 12: 3114-3125], as well as to characterize antisense transcription on a global scale in Plasmodium. In the present study, the comprehensive annotation of SAGE libraries derived from an asexual stage population exposed to drug and its matched control was used to assess the modulation of gene expression by chloroquine. Here, we observed a constellation of changes, with the differential regulation of over 100 transcripts, and have confirmed the data by alternate methods. A few responsive loci, including PfMDR1, have previously been implicated in the mechanism of chloroquine action/resistance. Several others, however, were derived from unexpected categories, including a large number of unknown open reading frames (ORFs), whose induction after drug exposure may provide first hints to their possible function.

Animals↗

PlasmoDB: the Plasmodium genome resource. An integrated database providing tools for accessing, analyzing and mapping expression and sequence data (both finished and unfinished).

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates finished and draft genome sequence data and annotation emerging from Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for cross-species comparisons. Sequence information is also integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects. The relational schemas used to build PlasmoDB [Genomics Unified Schema (GUS) and RNA Abundance Database (RAD)] employ a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically based queries of the database. A version of the database is also available on CD-ROM (Plasmodium GenePlot), facilitating access to the data in situations where Internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to enhance utilization of the vast quantities of data emerging from genome-scale projects by the global malaria research community.

Animals↗

Predicting gene ontology functions from ProDom and CDD protein domains.

A heuristic algorithm for associating Gene Ontology (GO) defined molecular functions to protein domains as listed in the ProDom and CDD databases is described. The algorithm generates rules for function-domain associations based on the intersection of functions assigned to gene products by the GO consortium that contain ProDom and/or CDD domains at varying levels of sequence similarity. The hierarchical nature of GO molecular functions is incorporated into rule generation. Manual review of a subset of the rules generated indicates an accuracy rate of 87% for ProDom rules and 84% for CDD rules. The utility of these associations is that novel sequences can be assigned a putative function if sufficient similarity exists to a ProDom or CDD domain for which one or more GO functions has been associated. Although functional assignments are increasingly being made for gene products from model organisms, it is likely that the needs of investigators will continue to outpace the efforts of curators, particularly for nonmodel organisms. A comparison with other methods in terms of coverage and agreement was performed, indicating the utility of the approach. The domain-function associations and function assignments are available from our website http://www.cbil.upenn.edu/GO.

Algorithms↗