PubMed Health⌕ Search

PubMed · 9697170

Cluster analysis and data visualization of large-scale gene expression data.

Abstract

The discovery of any new gene requires an analysis of the expression context for that gene. Now that the cDNA and genomic sequencing projects are progressing at such a rapid rate, high throughput gene expression screening approaches are beginning to appear to take advantage of that data. We present a strategy for the analysis for large-scale quantitative gene expression measurement data from time course experiments. Our approach takes advantage of cluster analysis and graphical visualization methods to reveal correlated patterns of gene expression from time series data. The coherence of these patterns suggests an order that conforms to a notion of shared pathways and control processes that can be experimentally verified.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

G S Michaels, D B Carr, M Askenazi, S Fuhrman, X Wen, R Somogyi. 1998. Cluster analysis and data visualization of large-scale gene expression data.. https://pubmed.ncbi.nlm.nih.gov/9697170/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

miR-503-3p promotes epithelial-mesenchymal transition in breast cancer by directly targeting SMAD2 and E-cadherin.

Although progress in clinical and basic research has significantly increased our understanding of breast cancer, little is known about the molecular mechanism underlying breast cancer metastasis. Identification of effective therapeutic targets to prevent breast cancer metastasis is urgently needed. The function of miR-503-3p has been investigated in other cancers, but its role in breast cancer remains undefined. Here, we found that miR-503-3p was overexpressed in breast cancer tissue and plasma compared with adjacent normal breast tissue and with plasma from healthy individuals. Moreover, we identified miR-503-3p to be an oncogene of breast cancer cell proliferation, migration and invasion. Upregulation of miR-503-3p in breast cancer cells inhibited expression of epithelial-mesenchymal transition (EMT)-related protein SMAD2 and the epithelial marker protein E-cadherin by directly binding to their mRNA 3' untranslated region, whereas increased expression of mesenchymal marker proteins, including vimentin and N-cadherin. Taken together, our findings support a critical role for miR-503-3p in induction of breast cancer EMT and suggest that plasma miR-503-3p may be a useful diagnostic biomarker for breast cancer.

Base Sequence↗

Identification and characterization of Prp45p and Prp46p, essential pre-mRNA splicing factors.

Through exhaustive two-hybrid screens using a budding yeast genomic library, and starting with the splicing factor and DEAH-box RNA helicase Prp22p as bait, we identified yeast Prp45p and Prp46p. We show that as well as interacting in two-hybrid screens, Prp45p and Prp46p interact with each other in vitro. We demonstrate that Prp45p and Prp46p are spliceosome associated throughout the splicing process and both are essential for pre-mRNA splicing. Under nonsplicing conditions they also associate in coprecipitation assays with low levels of the U2, U5, and U6 snRNAs that may indicate their presence in endogenous activated spliceosomes or in a postsplicing snRNP complex.

Base Sequence↗

Genetic engineering of the Trichoderma reesei endoglucanase I (Cel7B) for enhanced partitioning in aqueous two-phase systems containing thermoseparating ethylene oxide--propylene oxide copolymers.

Endoglucanases (endo-1,4-beta-D-glucan-4-glucanohydrolase, EC 3.2.1.4) are industrially important enzymes. In this study endoglucanase I (EGI or Cel7B) of the filamentous fungi Trichoderma reesei has been genetically engineered to investigate the influence of tryptophan rich peptide extensions (tags) on partitioning in an aqueous two-phase model system. EGI is a two-domain enzyme and is composed of a N-terminal catalytic domain and a C-terminal cellulose binding domain, separated by a linker. The aim was to find an optimal tag and fusion position, which further could be utilised for large scale extractions. Peptide tags of different length and composition were attached at various localisations of EGI. The fusion proteins were expressed from T. reesei with the use of the gpdA promoter from Aspergillus nidulans. Variations in secreted levels between the engineered proteins were obtained. The partitioning of EGI in an aqueous two-phase system composed of a thermoseparating ethylene oxide-propylene oxide random copolymer (EO(50)PO(50)) and dextran, could be significantly improved by relatively minor genetic engineering. The (Trp-Pro)(4) tag added after a short stretch of the linker, containing five proline residues, gave in the highest partition coefficient of 12.8. The yield in the top phase was 94%. The specific activity was 83% of the specific activity of unmodified EGI on soluble substrate. The efficiency of a tag fused to a protein is shown by the tag efficiency factor (TEF). A hypothetical TEF of 1.0 would indicate full tag exposure and optimal contribution to the protein partitioning by the fused tag. The location of the fusion point after the sequence of five proline residues in the linker of EGI is the most beneficial in two-phase separation. The highest TEF (0.97) was obtained with the (Trp-Pro)(2) tag at this position, indicating full exposure and intactness of the tag. However, the peptide tag composed of (Trp-Pro)(4) improved the partition properties the most but had lower TEF in comparison to (Trp-Pro)(2).

Base Sequence↗