PubMed Health⌕ Search

Biomedical subjects

David Botstein

Publications and source records attributed to David Botstein.

At least 19 recordsLinked to original sources

Endothelial cell diversity revealed by global expression profiling.

The vascular system is locally specialized to accommodate widely varying blood flow and pressure and the distinct needs of individual tissues. The endothelial cells (ECs) that line the lumens of blood and lymphatic vessels play an integral role in the regional specialization of vascular structure and physiology. However, our understanding of EC diversity is limited. To explore EC specialization on a global scale, we used DNA microarrays to determine the expression profile of 53 cultured ECs. We found that ECs from different blood vessels and microvascular ECs from different tissues have distinct and characteristic gene expression profiles. Pervasive differences in gene expression patterns distinguish the ECs of large vessels from microvascular ECs. We identified groups of genes characteristic of arterial and venous endothelium. Hey2, the human homologue of the zebrafish gene gridlock, was selectively expressed in arterial ECs and induced the expression of several arterial-specific genes. Several genes critical in the establishment of left/right asymmetry were expressed preferentially in venous ECs, suggesting coordination between vascular differentiation and body plan development. Tissue-specific expression patterns in different tissue microvascular ECs suggest they are distinct differentiated cell types that play roles in the local physiology of their respective organs and tissues.

Cells, Cultured↗

Gene expression patterns in ovarian carcinomas.

We used DNA microarrays to characterize the global gene expression patterns in surface epithelial cancers of the ovary. We identified groups of genes that distinguished the clear cell subtype from other ovarian carcinomas, grade I and II from grade III serous papillary carcinomas, and ovarian from breast carcinomas. Six clear cell carcinomas were distinguished from 36 other ovarian carcinomas (predominantly serous papillary) based on their gene expression patterns. The differences may yield insights into the worse prognosis and therapeutic resistance associated with clear cell carcinomas. A comparison of the gene expression patterns in the ovarian cancers to published data of gene expression in breast cancers revealed a large number of differentially expressed genes. We identified a group of 62 genes that correctly classified all 125 breast and ovarian cancer specimens. Among the best discriminators more highly expressed in the ovarian carcinomas were PAX8 (paired box gene 8), mesothelin, and ephrin-B1 (EFNB1). Although estrogen receptor was expressed in both the ovarian and breast cancers, genes that are coregulated with the estrogen receptor in breast cancers, including GATA-3, LIV-1, and X-box binding protein 1, did not show a similar pattern of coexpression in the ovarian cancers.

Adenocarcinoma↗

Repeated observation of breast tumor subtypes in independent gene expression data sets.

Characteristic patterns of gene expression measured by DNA microarrays have been used to classify tumors into clinically relevant subgroups. In this study, we have refined the previously defined subtypes of breast tumors that could be distinguished by their distinct patterns of gene expression. A total of 115 malignant breast tumors were analyzed by hierarchical clustering based on patterns of expression of 534 "intrinsic" genes and shown to subdivide into one basal-like, one ERBB2-overexpressing, two luminal-like, and one normal breast tissue-like subgroup. The genes used for classification were selected based on their similar expression levels between pairs of consecutive samples taken from the same tumor separated by 15 weeks of neoadjuvant treatment. Similar cluster analyses of two published, independent data sets representing different patient cohorts from different laboratories, uncovered some of the same breast cancer subtypes. In the one data set that included information on time to development of distant metastasis, subtypes were associated with significant differences in this clinical feature. By including a group of tumors from BRCA1 carriers in the analysis, we found that this genotype predisposes to the basal tumor subtype. Our results strongly support the idea that many of these breast tumor subtypes represent biologically distinct disease entities.

Breast Neoplasms↗

A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).

Genomic sequencing is no longer a novelty, but gene function annotation remains a key challenge in modern biology. A variety of functional genomics experimental techniques are available, from classic methods such as affinity precipitation to advanced high-throughput techniques such as gene expression microarrays. In the future, more disparate methods will be developed, further increasing the need for integrated computational analysis of data generated by these studies. We address this problem with MAGIC (Multisource Association of Genes by Integration of Clusters), a general framework that uses formal Bayesian reasoning to integrate heterogeneous types of high-throughput biological data (such as large-scale two-hybrid screens and multiple microarray analyses) for accurate gene function prediction. The system formally incorporates expert knowledge about relative accuracies of data sources to combine them within a normative framework. MAGIC provides a belief level with its output that allows the user to vary the stringency of predictions. We applied MAGIC to Saccharomyces cerevisiae genetic and physical interactions, microarray, and transcription factor binding sites data and assessed the biological relevance of gene groupings using Gene Ontology annotations produced by the Saccharomyces Genome Database. We found that by creating functional groupings based on heterogeneous data types, MAGIC improved accuracy of the groupings compared with microarray analysis alone. We describe several of the biological gene groupings identified.

Algorithms↗

Variation in gene expression patterns in human gastric cancers.

Gastric cancer is the world's second most common cause of cancer death. We analyzed gene expression patterns in 90 primary gastric cancers, 14 metastatic gastric cancers, and 22 nonneoplastic gastric tissues, using cDNA microarrays representing approximately 30,300 genes. Gastric cancers were distinguished from nonneoplastic gastric tissues by characteristic differences in their gene expression patterns. We found a diversity of gene expression patterns in gastric cancer, reflecting variation in intrinsic properties of tumor and normal cells and variation in the cellular composition of these complex tissues. We identified several genes whose expression levels were significantly correlated with patient survival. The variations in gene expression patterns among cancers in different patients suggest differences in pathogenetic pathways and potential therapeutic strategies.

Adenocarcinoma↗

Generalized singular value decomposition for comparative analysis of genome-scale expression data sets of two different organisms.

We describe a comparative mathematical framework for two genome-scale expression data sets. This framework formulates expression as superposition of the effects of regulatory programs, biological processes, and experimental artifacts common to both data sets, as well as those that are exclusive to one data set or the other, by using generalized singular value decomposition. This framework enables comparative reconstruction and classification of the genes and arrays of both data sets. We illustrate this framework with a comparison of yeast and human cell-cycle expression data sets.

Cell Cycle↗

Variation in gene expression patterns in follicular lymphoma and the response to rituximab.

Analysis of the patterns of gene expression in follicular lymphomas from 24 patients suggested that two groups of tumors might be distinguished. All patients, whose biopsies were obtained before any treatment, were treated with rituximab, a monoclonal antibody directed against the B cell antigen, CD20. Gene expression patterns in the tumors that subsequently failed to respond to rituximab appeared more similar to those of normal lymphoid tissues than to gene expression patterns of tumors from rituximab responders. These findings suggest the possibility that the response of follicular lymphoma to rituximab treatment may be predicted from the gene expression pattern of tumors.

Adult↗

SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.

The explosion in the number of functional genomic datasets generated with tools such as DNA microarrays has created a critical need for resources that facilitate the interpretation of large-scale biological data. SOURCE is a web-based database that brings together information from a broad range of resources, and provides it in manner particularly useful for genome-scale analyses. SOURCE's GeneReports include aliases, chromosomal location, functional descriptions, GeneOntology annotations, gene expression data, and links to external databases. We curate published microarray gene expression datasets and allow users to rapidly identify sets of co-regulated genes across a variety of tissues and a large number of conditions using a simple and intuitive interface. SOURCE provides content both in gene and cDNA clone-centric pages, and thus simplifies analysis of datasets generated using cDNA microarrays. SOURCE is continuously updated and contains the most recent and accurate information available for human, mouse, and rat genes. By allowing dynamic linking to individual gene or clone reports, SOURCE facilitates browsing of large genomic datasets. Finally, SOURCEs batch interface allows rapid extraction of data for thousands of genes or clones at once and thus facilitates statistical analyses such as assessing the enrichment of functional attributes within clusters of genes. SOURCE is available at http://source.stanford.edu.

Animals↗

Saccharomyces Genome Database (SGD) provides biochemical and structural information for budding yeast proteins.

The Saccharomyces Genome Database (SGD: http://genome-www.stanford.edu/Saccharomyces/) has recently developed new resources to provide more complete information about proteins from the budding yeast Saccharomyces cerevisiae. The PDB Homologs page provides structural information from the Protein Data Bank (PDB) about yeast proteins and/or their homologs. SGD has also created a resource that utilizes the eMOTIF database for motif information about a given protein. A third new resource is the Protein Information page, which contains protein physical and chemical properties, such as molecular weight and hydropathicity scores, predicted from the translated ORF sequence.

Amino Acid Motifs↗

The Stanford Microarray Database: data access and quality assessment tools.

The Stanford Microarray Database (SMD; http://genome-www.stanford.edu/microarray/) serves as a microarray research database for Stanford investigators and their collaborators. In addition, SMD functions as a resource for the entire scientific community, by making freely available all of its source code and providing full public access to data published by SMD users, along with many tools to explore and analyze those data. SMD currently provides public access to data from 3500 microarrays, including data from 85 publications, and this total is increasing rapidly. In this article, we describe some of SMD's newer tools for accessing public data, assessing data quality and for data analysis.

Animals↗

A genome scan for hypertension susceptibility loci in populations of Chinese and Japanese origins.

BACKGROUND: Our understanding of genes that predispose to essential hypertension is poor. METHODS: A genome-wide scan for linkage at approximately 10 cM resolution was done on 1425 sibpairs of Chinese and Japanese origins that were concordant for hypertension (N = 661), low-normal blood pressure (BP) (N = 184), or discordant for BP (N = 580). RESULTS: There was no significant evidence of linkage to a single locus in the genome. There was suggestive evidence of linkage to chromosome 10p, with a LOD score of 2.5. CONCLUSIONS: We can exclude the possibility that a single gene accounts for at least 15% of the variance in hypertension in this population.

Adult↗

Discovering genotypes underlying human phenotypes: past successes for mendelian disease, future approaches for complex disease.

The past two decades have witnessed an explosion in the identification, largely by positional cloning, of genes associated with mendelian diseases. The roughly 1,200 genes that have been characterized have clarified our understanding of the molecular basis of human genetic disease. The principles derived from these successes should be applied now to strategies aimed at finding the considerably more elusive genes that underlie complex disease phenotypes. The distribution of types of mutation in mendelian disease genes argues for serious consideration of the early application of a genomic-scale sequence-based approach to association studies and against complete reliance on a positional cloning approach based on a map of anonymous single nucleotide polymorphism haplotypes.

Alleles↗

Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data.

Much of a cell's activity is organized as a network of interacting modules: sets of genes coregulated to respond to different conditions. We present a probabilistic method for identifying regulatory modules from gene expression data. Our procedure identifies modules of coregulated genes, their regulators and the conditions under which regulation occurs, generating testable hypotheses in the form 'regulator X regulates module Y under conditions W'. We applied the method to a Saccharomyces cerevisiae expression data set, showing its ability to identify functionally coherent modules and their correct regulators. We present microarray experiments supporting three novel predictions, suggesting regulatory roles for previously uncharacterized proteins.

Algorithms↗

A systematic approach to reconstructing transcription networks in Saccharomycescerevisiae.

Decomposing regulatory networks into functional modules is a first step toward deciphering the logical structure of complex networks. We propose a systematic approach to reconstructing transcription modules (defined by a transcription factor and its target genes) and identifying conditionsperturbations under which a particular transcription module is activateddeactivated. Our approach integrates information from regulatory sequences, genome-wide mRNA expression data, and functional annotation. We systematically analyzed gene expression profiling experiments in which the yeast cell was subjected to various environmental or genetic perturbations. We were able to construct transcription modules with high specificity and sensitivity for many transcription factors, and predict the activation of these modules under anticipated as well as unexpected conditions. These findings generate testable hypotheses when combined with existing knowledge on signaling pathways and protein-protein interactions. Correlating the activation of a module to a specific perturbation predicts links in the cell's regulatory networks, and examining coactivated modules suggests specific instances of crosstalk between regulatory pathways.

Gene Expression Profiling↗

Overview of the Alliance for Cellular Signaling.

The Alliance for Cellular Signaling is a large-scale collaboration designed to answer global questions about signalling networks. Pathways will be studied intensively in two cells--B lymphocytes (the cells of the immune system) and cardiac myocytes--to facilitate quantitative modelling. One goal is to catalyse complementary research in individual laboratories; to facilitate this, all alliance data are freely available for use by the entire research community.

B-Lymphocytes↗

Phospholipase A2 group IIA expression in gastric adenocarcinoma is associated with prolonged survival and less frequent metastasis.

We analyzed gene expression patterns in human gastric cancers by using cDNA microarrays representing approximately equal 30,300 genes. Expression of PLA2G2A, a gene previously implicated as a modifier of the Apc(Min/+) (multiple intestinal neoplasia 1) mutant phenotype in the mouse, was significantly correlated with patient survival. We confirmed this observation in an independent set of patient samples by using quantitative RT-PCR. Beyond its potential diagnostic and prognostic significance, this result suggests the intriguing possibility that the activity of PLA2G2A may suppress progression or metastasis of human gastric cancer.

Adenocarcinoma↗

Characteristic genome rearrangements in experimental evolution of Saccharomyces cerevisiae.

Genome rearrangements, especially amplifications and deletions, have regularly been observed as responses to sustained application of the same strong selective pressure in microbial populations growing in continuous culture. We studied eight strains of budding yeast (Saccharomyces cerevisiae) isolated after 100-500 generations of growth in glucose-limited chemostats. Changes in DNA copy number were assessed at single-gene resolution by using DNA microarray-based comparative genomic hybridization. Six of these evolved strains were aneuploid as the result of gross chromosomal rearrangements. Most of the aneuploid regions were the result of translocations, including three instances of a shared breakpoint on chromosome 14 immediately adjacent to CIT1, which encodes the citrate synthase that performs a key regulated step in the tricarboxylic acid cycle. Three strains had amplifications in a region of chromosome 4 that includes the high-affinity hexose transporters; one of these also had the aforementioned chromosome 14 break. Three strains had extensive overlapping deletions of the right arm of chromosome 15. Further analysis showed that each of these genome rearrangements was bounded by transposon-related sequences at the breakpoints. The observation of repeated, independent, but nevertheless very similar, chromosomal rearrangements in response to persistent selection of growing cells parallels the genome rearrangements that characteristically accompany tumor progression.

Aneuploidy↗