PubMed Health⌕ Search

PubMed · 11108479

Tissue classification with gene expression profiles.

Abstract

Constantly improving gene expression profiling technologies are expected to provide understanding and insight into cancer-related cellular processes. Gene expression data is also expected to significantly aid in the development of efficient cancer diagnosis and classification platforms. In this work we examine three sets of gene expression data measured across sets of tumor(s) and normal clinical samples: The first set consists of 2,000 genes, measured in 62 epithelial colon samples (Alon et al., 1999). The second consists of approximately equal to 100,000 clones, measured in 32 ovarian samples (unpublished extension of data set described in Schummer et al. (1999)). The third set consists of approximately equal to 7,100 genes, measured in 72 bone marrow and peripheral blood samples (Golub et al, 1999). We examine the use of scoring methods, measuring separation of tissue type (e.g., tumors from normals) using individual gene expression levels. These are then coupled with high-dimensional classification methods to assess the classification power of complete expression profiles. We present results of performing leave-one-out cross validation (LOOCV) experiments on the three data sets, employing nearest neighbor classifier, SVM (Cortes and Vapnik, 1995), AdaBoost (Freund and Schapire, 1997) and a novel clustering-based classification technique. As tumor samples can differ from normal samples in their cell-type composition, we also perform LOOCV experiments using appropriately modified sets of genes, attempting to eliminate the resulting bias. We demonstrate success rate of at least 90% in tumor versus normal classification, using sets of selected genes, with, as well as without, cellular-contamination-related members. These results are insensitive to the exact selection mechanism, over a certain range.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Ben-Dor, L Bruhn, N Friedman, I Nachman, M Schummer, Z Yakhini. 2000. Tissue classification with gene expression profiles.. https://doi.org/10.1089/106652700750050943

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Fuzzy species among recombinogenic bacteria.

BACKGROUND: It is a matter of ongoing debate whether a universal species concept is possible for bacteria. Indeed, it is not clear whether closely related isolates of bacteria typically form discrete genotypic clusters that can be assigned as species. The most challenging test of whether species can be clearly delineated is provided by analysis of large populations of closely-related, highly recombinogenic, bacteria that colonise the same body site. We have used concatenated sequences of seven house-keeping loci from 770 strains of 11 named Neisseria species, and phylogenetic trees, to investigate whether genotypic clusters can be resolved among these recombinogenic bacteria and, if so, the extent to which they correspond to named species. RESULTS: Alleles at individual loci were widely distributed among the named species but this distorting effect of recombination was largely buffered by using concatenated sequences, which resolved clusters corresponding to the three species most numerous in the sample, N. meningitidis, N. lactamica and N. gonorrhoeae. A few isolates arose from the branch that separated N. meningitidis from N. lactamica leading us to describe these species as 'fuzzy'. CONCLUSION: A multilocus approach using large samples of closely related isolates delineates species even in the highly recombinogenic human Neisseria where individual loci are inadequate for the task. This approach should be applied by taxonomists to large samples of other groups of closely-related bacteria, and especially to those where species delineation has historically been difficult, to determine whether genotypic clusters can be delineated, and to guide the definition of species.

Cluster Analysis↗

Representation is faithfully preserved in global cDNA amplified exponentially from sub-picogram quantities of mRNA.

Analysis of transcript representation on gene microarrays requires microgram amounts of total RNA or DNA. Without amplification, such amounts are obtainable only from millions of cells. However, it may be desirable to determine transcript representation in few or even single cells in aspiration biopsies, rare population subsets isolated by cell sorting or laser capture, or micromanipulated single cells. Nucleic-acid amplification methods could be used in these cases, but it is difficult to amplify different transcripts in a sample without distorting quantitative relationships between them. Linear isothermal RNA amplification has been used to amplify as little as 10 ng of total cellular RNA, corresponding to the amount obtainable from thousands of cells, while still preserving the original abundance relationships. However, the available procedures require multiple steps, are labor intensive and time consuming, and have not been shown to preserve abundance information from smaller starting amounts. Exponential amplification, on the other hand, is a relatively simple technology, but is generally considered to bias abundance relationships unacceptably. These constraints have placed beyond current reach the secure and routine application of microarray analysis to single or small numbers of cells. Here we describe results obtained with a rapid and highly optimized global reverse transcription#150;PCR (RT-PCR) procedure. Contrary to prevalent expectations, the exponential approach preserves abundance relationships through amplification as high as 3 x 10(11)-fold. Further, it reduces by a million-fold the input amount of RNA needed for microarray analysis, and yields reproducible results from the picogram range of total RNA obtainable from single cells.

Cluster Analysis↗