PubMed Health⌕ Search

Biomedical subjects

Robert L Strausberg

Publications and source records attributed to Robert L Strausberg.

At least 19 recordsLinked to original sources

Somatic sequence alterations in twenty-one genes selected by expression profile analysis of breast carcinomas.

INTRODUCTION: Genomic alterations have been observed in breast carcinomas that affect the capacity of cells to regulate proliferation, signaling, and metastasis. Re-sequence studies have investigated candidate genes based on prior genetic observations (changes in copy number or regions of genetic instability) or other laboratory observations and have defined critical somatic mutations in genes such as TP53 and PIK3CA. METHODS: We have extended the paradigm and analyzed 21 genes primarily identified by expression profiling studies, which are useful for breast cancer subtyping and prognosis. This study conducted a bidirectional re-sequence analysis of all exons and 5', 3', and evolutionarily conserved regions (spanning more than 16 megabases) in 91 breast tumor samples. RESULTS: Eighty-seven unique somatic alterations were identified in 16 genes. Seventy-eight were single base pair alterations, of which 23 were missense mutations; 55 were distributed across conserved intronic regions or the 5' and 3' regions. There were nine insertion/deletions. Because there is no a priori way to predict whether any one of the identified synonymous and noncoding somatic alterations disrupt function, analysis unique to each gene will be required to establish whether it is a tumor suppressor gene or whether there is no effect. In five genes, no somatic alterations were observed. CONCLUSION: The study confirms the value of re-sequence analysis in cancer gene discovery and underscores the importance of characterizing somatic alterations across genes that are related not only by function, or functional pathways, but also based upon expression patterns.

Breast Neoplasms↗

Ancient noncoding elements conserved in the human genome.

Cartilaginous fishes represent the living group of jawed vertebrates that diverged from the common ancestor of human and teleost fish lineages about 530 million years ago. We generated approximately 1.4x genome sequence coverage for a cartilaginous fish, the elephant shark (Callorhinchus milii), and compared this genome with the human genome to identify conserved noncoding elements (CNEs). The elephant shark sequence revealed twice as many CNEs as were identified by whole-genome comparisons between teleost fishes and human. The ancient vertebrate-specific CNEs in the elephant shark and human genomes are likely to play key regulatory roles in vertebrate gene expression.

Animals↗

High expression of a cytokeratin-associated protein in many cancers.

We have described previously a cDNA library made from membrane-bound polysomal mRNA prepared from breast and prostate cancer cell lines. The library is highly enriched for cDNAs encoding membrane proteins, secreted proteins, and cytokeratins. To characterize this library, 25,277 cDNA clones were sequenced and aligned with various databases; 1,439 clones did not align with known genes. From this set of clones we identified a previously uncharacterized gene encoding a 334-aa protein. Although protein structural motif prediction programs indicate that the gene encodes a membrane protein comprising a signal sequence, a series of leucine-rich repeats, and a single transmembrane domain with a cytoplasmic tail, confocal microscopy of MCF7 breast cancer cells demonstrates that the protein is not directly associated with the plasma membrane or intracellular membranes but instead colocalizes with intermediate filaments and cytokeratins within the cell. Immunofluorescence studies also show that protein expression is increased greatly in mitotic MCF7 cells, and immunohistochemistry demonstrates its expression in human breast cancer cells. Analysis of mRNA levels in 25 different normal tissues by RT-PCR shows that this gene is expressed highly in normal prostate and salivary gland, very weakly in colon, pancreas, and intestine, and not at all in other tissues. RT-PCR studies on human cancer samples show that the RNA is expressed highly in many cancer cell lines and cancer specimens, including 26 of 33 human breast cancers, 3 of 3 prostate cancers, 3 of 3 colon cancers, and 3 of 3 pancreatic cancers. We name the protein CAPC, cytokeratin-associated protein in cancer.

Amino Acid Sequence↗

Establishment of the epithelial-specific transcriptome of normal and malignant human breast cells based on MPSS and array expression data.

INTRODUCTION: Diverse microarray and sequencing technologies have been widely used to characterise the molecular changes in malignant epithelial cells in breast cancers. Such gene expression studies to identify markers and targets in tumour cells are, however, compromised by the cellular heterogeneity of solid breast tumours and by the lack of appropriate counterparts representing normal breast epithelial cells. METHODS: Malignant neoplastic epithelial cells from primary breast cancers and luminal and myoepithelial cells isolated from normal human breast tissue were isolated by immunomagnetic separation methods. Pools of RNA from highly enriched preparations of these cell types were subjected to expression profiling using massively parallel signature sequencing (MPSS) and four different genome wide microarray platforms. Functional related transcripts of the differential tumour epithelial transcriptome were used for gene set enrichment analysis to identify enrichment of luminal and myoepithelial type genes. Clinical pathological validation of a small number of genes was performed on tissue microarrays. RESULTS: MPSS identified 6,553 differentially expressed genes between the pool of normal luminal cells and that of primary tumours substantially enriched for epithelial cells, of which 98% were represented and 60% were confirmed by microarray profiling. Significant expression level changes between these two samples detected only by microarray technology were shown by 4,149 transcripts, resulting in a combined differential tumour epithelial transcriptome of 8,051 genes. Microarray gene signatures identified a comprehensive list of 907 and 955 transcripts whose expression differed between luminal epithelial cells and myoepithelial cells, respectively. Functional annotation and gene set enrichment analysis highlighted a group of genes related to skeletal development that were associated with the myoepithelial/basal cells and upregulated in the tumour sample. One of the most highly overexpressed genes in this category, that encoding periostin, was analysed immunohistochemically on breast cancer tissue microarrays and its expression in neoplastic cells correlated with poor outcome in a cohort of poor prognosis estrogen receptor-positive tumours. CONCLUSION: Using highly enriched cell populations in combination with multiplatform gene expression profiling studies, a comprehensive analysis of molecular changes between the normal and malignant breast tissue was established. This study provides a basis for the identification of novel and potentially important targets for diagnosis, prognosis and therapy in breast cancer.

Biomarkers, Tumor↗

Sequence survey of receptor tyrosine kinases reveals mutations in glioblastomas.

It is now clear that tyrosine kinases represent attractive targets for therapeutic intervention in cancer. Recent advances in DNA sequencing technology now provide the opportunity to survey mutational changes in cancer in a high-throughput and comprehensive manner. Here we report on the sequence analysis of members of the receptor tyrosine kinase (RTK) gene family in the genomes of glioblastoma brain tumors. Previous studies have identified a number of molecular alterations in glioblastoma, including amplification of the RTK epidermal growth factor receptor. We have identified mutations in two other RTKs: (i) fibroblast growth receptor 1, including the first mutations in the kinase domain in this gene observed in any cancer, and (ii) a frameshift mutation in the platelet-derived growth factor receptor-alpha gene. Fibroblast growth receptor 1, platelet-derived growth factor receptor-alpha, and epidermal growth factor receptor are all potential entry points to the phosphatidylinositol 3-kinase and mitogen-activated protein kinase intracellular signaling pathways already known to be important for neoplasia. Our results demonstrate the utility of applying DNA sequencing technology to systematically assess the coding sequence of genes within cancer genomes.

Adult↗

Identification of cancer/testis-antigen genes by massively parallel signature sequencing.

Massively parallel signature sequencing (MPSS) generates millions of short sequence tags corresponding to transcripts from a single RNA preparation. Most MPSS tags can be unambiguously assigned to genes, thereby generating a comprehensive expression profile of the tissue of origin. From the comparison of MPSS data from 32 normal human tissues, we identified 1,056 genes that are predominantly expressed in the testis. Further evaluation by using MPSS tags from cancer cell lines and EST data from a wide variety of tumors identified 202 of these genes as candidates for encoding cancer/testis (CT) antigens. Of these genes, the expression in normal tissues was assessed by RT-PCR in a subset of 166 intron-containing genes, and those with confirmed testis-predominant expression were further evaluated for their expression in 21 cancer cell lines. Thus, 20 CT or CT-like genes were identified, with several exhibiting expression in five or more of the cancer cell lines examined. One of these genes is a member of a CT gene family that we designated as CT45. The CT45 family comprises six highly similar (>98% cDNA identity) genes that are clustered in tandem within a 125-kb region on Xq26.3. CT45 was found to be frequently expressed in both cancer cell lines and lung cancer specimens. Thus, MPSS analysis has resulted in a significant extension of our knowledge of CT antigens, leading to the discovery of a distinctive X-linked CT-antigen gene family.

Antigens, Neoplasm↗

Tumor microenvironments, the immune system and cancer survival.

The study of cancer immunology has recently been reinvigorated by the application of new research tools and technologies, as well as by refined bioinformatics methods for interpretation of complex datasets. Recent microarray analyses of lymphomas suggest that the prognosis of cancer patients is related to an interplay between cancer cells and their microenvironment, including the immune response.

Cell Transformation, Neoplastic↗

The interactive online SKY/M-FISH & CGH database and the Entrez cancer chromosomes search database: linkage of chromosomal aberrations with the genome sequence.

To catalog data on chromosomal aberrations in cancer derived from emerging molecular cytogenetic techniques and to integrate these data with genome maps, we have established two resources, the NCI and NCBI SKY/M-FISH & CGH Database and the Cancer Chromosomes database. The goal of the former is to allow investigators to submit and analyze clinical and research cytogenetic data. It contains a karyotype parser tool, which automatically converts the ISCN short-form karyotype into an internal representation displayed in detailed form and as a colored ideogram with band overlay, and also has a tool to compare CGH profiles from multiple cases. The Cancer Chromosomes database integrates the SKY/M-FISH & CGH Database with the Mitelman Database of Chromosome Aberrations in Cancer and the Recurrent Chromosome Aberrations in Cancer database. These three datasets can now be searched seamlessly by use of the Entrez search and retrieval system for chromosome aberrations, clinical data, and reference citations. Common diagnoses, anatomic sites, chromosome breakpoints, junctions, numerical and structural abnormalities, and bands gained and lost among selected cases can be compared by use of the "similarity" report. Because the model used for CGH data is a subset of the karyotype data, it is now possible to examine the similarities between CGH results and karyotypes directly. All chromosomal bands are directly linked to the Entrez Map Viewer database, providing integration of cytogenetic data with the sequence assembly. These resources, developed as a part of the Cancer Chromosome Aberration Project (CCAP) initiative, aid the search for new cancer-associated genes and foster insights into the causes and consequences of genetic alterations in cancer.

Base Sequence↗

An atlas of human gene expression from massively parallel signature sequencing (MPSS).

We have used massively parallel signature sequencing (MPSS) to sample the transcriptomes of 32 normal human tissues to an unprecedented depth, thus documenting the patterns of expression of almost 20,000 genes with high sensitivity and specificity. The data confirm the widely held belief that differences in gene expression between cell and tissue types are largely determined by transcripts derived from a limited number of tissue-specific genes, rather than by combinations of more promiscuously expressed genes. Expression of a little more than half of all known human genes seems to account for both the common requirements and the specific functions of the tissues sampled. A classification of tissues based on patterns of gene expression largely reproduces classifications based on anatomical and biochemical properties. The unbiased sampling of the human transcriptome achieved by MPSS supports the idea that most human genes have been mapped, if not functionally characterized. This data set should prove useful for the identification of tissue-specific genes, for the study of global changes induced by pathological conditions, and for the definition of a minimal set of genes necessary for basic cell maintenance. The data are available on the Web at http://mpss.licr.org and http://sgb.lynxgen.com.

Algorithms↗

Mutation of GATA3 in human breast tumors.

GATA3 is an essential transcription factor that was first identified as a regulator of immune cell function. In recent microarray analyses of human breast tumors, both normal breast luminal epithelium and estrogen receptor (ESR1)-positive tumors showed high expression of GATA3. We sequenced genomic DNA from 111 breast tumors and three breast-tumor-derived cell lines and identified somatic mutations of GATA3 in five tumors and the MCF-7 cell line. These mutations cluster in the vicinity of the highly conserved second zinc-finger that is required for DNA binding. In addition to these five, we identified using cDNA sequencing a unique mis-splicing variant that caused a frameshift mutation. One of the somatic mutations we identified was identical to a germline GATA3 mutation reported in two kindreds with HDR syndrome/OMIM #146255, which is an autosomal dominant syndrome caused by the haplo-insufficiency of GATA3. The ectopic expression of GATA3 in human 293T cells caused the induction of 73 genes including six cytokeratins, and inhibited cell line doubling times. These data suggest that GATA3 is involved in growth control and the maintenance of the differentiated state in epithelial cells, and that GATA3 variants may contribute to tumorigenesis in ESR1-positive breast tumors.

Amino Acid Sequence↗

Human, mouse, and rat genome large-scale rearrangements: stability versus speciation.

Using paired-end sequences from bacterial artificial chromosomes, we have constructed high-resolution synteny and rearrangement breakpoint maps among human, mouse, and rat genomes. Among the >300 syntenic blocks identified are segments of over 40 Mb without any detected interspecies rearrangements, as well as regions with frequently broken synteny and extensive rearrangements. As closely related species, mouse and rat share the majority of the breakpoints and often have the same types of rearrangements when compared with the human genome. However, the breakpoints not shared between them indicate that mouse rearrangements are more often interchromosomal, whereas intrachromosomal rearrangements are more prominent in rat. Centromeres may have played a significant role in reorganizing a number of chromosomes in all three species. The comparison of the three species indicates that genome rearrangements follow a path that accommodates a delicate balance between maintaining a basic structure underlying all mammalian species and permitting variations that are necessary for speciation.

Animals↗

Oncogenomics and the development of new cancer therapies.

Scientists have sequenced the human genome and identified most of its genes. Now it is time to use these genomic data, and the high-throughput technology developed to generate them, to tackle major health problems such as cancer. To accelerate our understanding of this disease and to produce targeted therapies, further basic mutational and functional genomic information is required. A systematic and coordinated approach, with the results freely available, should speed up progress. This will best be accomplished through an international academic and pharmaceutical oncogenomics initiative.

Biomedical Research↗

Generation and analysis of melanoma SAGE libraries: SAGE advice on the melanoma transcriptome.

In this study, we generated three SAGE libraries from melanoma tissues. Using bioinformatics tools usually applied to microarray data, we identified several genes, including novel transcripts, which are preferentially expressed in melanoma. SAGE results converged with previous microarray analysis on the importance of intracellular calcium and G-protein signaling, and the Wnt/Frizzled family. We also examined the expression of CD74, which was specifically, albeit not abundantly, expressed in the melanoma libraries using a melanoma progression tissue microarray, and demonstrate that this protein is expressed by melanoma cells but not by benign melanocytes. Many genes involved in intracellular calcium and G-protein signaling were highly expressed in melanoma, results we had observed earlier from microarray studies (Bittner et al., 2000). One of the genes most highly expressed in our melanoma SAGE libraries was a calcium-regulated gene, calpain 3 (p94). Immunohistochemical analysis demonstrated that calpain 3 moves from the nuclei of non-neoplastic cells to the cytoplasm of malignant cells, suggesting activation of this intracellular proteinase. Our SAGE results and the clinical validation data demonstrate how SAGE profiles can highlight specific links between signaling pathways as well as associations with tumor progression. This may provide insights into new genes that may be useful for the diagnosis and therapy of melanoma.

Aged↗

The generation and utilization of a cancer-oriented representation of the human transcriptome by using expressed sequence tags.

Whereas genome sequencing defines the genetic potential of an organism, transcript sequencing defines the utilization of this potential and links the genome with most areas of biology. To exploit the information within the human genome in the fight against cancer, we have deposited some two million expressed sequence tags (ESTs) from human tumors and their corresponding normal tissues in the public databases. The data currently define approximately 23,500 genes, of which only approximately 1,250 are still represented only by ESTs. Examination of the EST coverage of known cancer-related (CR) genes reveals that <1% do not have corresponding ESTs, indicating that the representation of genes associated with commonly studied tumors is high. The careful recording of the origin of all ESTs we have produced has enabled detailed definition of where the genes they represent are expressed in the human body. More than 100,000 ESTs are available for seven tissues, indicating a surprising variability of gene usage that has led to the discovery of a significant number of genes with restricted expression, and that may thus be therapeutically useful. The ESTs also reveal novel nonsynonymous germline variants (although the one-pass nature of the data necessitates careful validation) and many alternatively spliced transcripts. Although widely exploited by the scientific community, vindicating our totally open source policy, the EST data generated still provide extensive information that remains to be systematically explored, and that may further facilitate progress toward both the understanding and treatment of human cancers.

Chromosome Mapping↗

Comparison of medulloblastoma and normal neural transcriptomes identifies a restricted set of activated genes.

Over 1.4 million transcript tags expressed in 20 different human medulloblastomas were counted using serial analysis of gene expression. Digital gene expression profiles in the medulloblastoma were compared to multiple regions of the normal human brain, revealing 30 transcripts with high expression in multiple tumors and little or no expression in the normal cerebellum and other adult and pediatric brain regions. Using independent medulloblastoma samples and normal tissue, real-time PCR verified eight of nine selected genes as candidate tumor-associated antigens. Differential protein expression for CD24, prolactin and Topo2A was further confirmed by immunohistochemical analysis using medulloblastoma and normal brain sections and a tissue microarray. The genes highly expressed in the medulloblastoma include PRAME, a cancer-testis antigen and potential targets for immunotherapy.

Brain↗

From knowing to controlling: a path from genomics to drugs using small molecule probes.

The National Cancer Institute Initiative in Chemical Genetics is designed to encourage the development of small molecular probes. The probes are useful for activating or inactivating protein functions, thereby providing resources that help discern the functions of gene products in normal and disease cells, as well as in tissues. This initiative includes "ChemBank," a suite of informatics tools and databases aimed at promoting the development and use of chemical genetics by scientists worldwide. The information generated with such tools should provide a critical link from genomic discovery to drug development.

Animals↗

Comprehensive sampling of gene expression in human cell lines with massively parallel signature sequencing.

Whereas information is rapidly accumulating about the structure and position of genes encoded in the human genome, less is known about the complexity and relative abundance of their expression in individual human cells and tissues. Here, we describe the characteristics of the transcriptomes of two cultured cell lines, HB4a (normal breast epithelium) and HCT-116 (colon adenocarcinoma), using massively parallel signature sequencing (MPSS). We generated in excess of 10(7) short signature sequences per cell line, thus providing a comprehensive snapshot of gene expression, within the technical limitations of the method. The number of genes expressed at one copy per cell or more in either of the lines was estimated to be between 10,000 and 15,000. The vast majority of the transcripts found in these cells can be mapped to known genes and their polyadenylation variants. Among the genes that could be identified from their signature sequences, approximately 8,500 were expressed by both cell lines, whereas 6,000 showed cellular specificity. Taking into account sequence tags that map uniquely to the genome but not to known transcripts, overall the data are consistent with an upper limit of 17,000 for the total number of genes expressed at more than one copy per cell in one or both of the two cell lines examined.

Adenocarcinoma↗

Sequence-based cancer genomics: progress, lessons and opportunities.

Technologies that provide a genome-wide view offer an unprecedented opportunity to scrutinize the molecular biology of the cancer cell. The information that is derived from these technologies is well suited to the development of public databases of alterations in the cancer genome and its expression. Here, we describe the synergistic efforts of research programmes in Brazil, the United Kingdom and the United States towards building integrated databases that are widely accessible to the research community, to enable basic and applied applications in cancer research.

Brazil↗