PubMed Health⌕ Search

PubMed · 15623938

Microarray data analysis: from hypotheses to conclusions using gene expression data.

Abstract

We review several commonly used methods for the design and analysis of microarray data. To begin with, some experimental design issues are addressed. Several approaches for pre-processing the data (filtering and normalization) before the statistical analysis stage are then discussed. A common first step in this type of analysis is gene selection based on statistical testing. Two approaches, permutation and model-based methods are explained and we emphasize the need to correct for multiple testing. Moreover, powerful approaches based on gene sets are mentioned. Clustering of either genes or samples is frequently performed when analyzing microarray data. We summarize the basics of both supervised and unsupervised clustering (classification). The latter may be of use for creating diagnostic arrays, for example. Construction of biological networks, such as pathways, is a statistically challenging but complex task that is a relatively new development and hence mentioned only briefly. We finish with some remarks on literature and software. The emphasis in this paper is on the philosophy behind several statistical issues and on a critical interpretation of microarray related analysis methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nicola J Armstrong, Mark A van de Wiel. 2004. Microarray data analysis: from hypotheses to conclusions using gene expression data.. https://doi.org/10.1155/2004%2F943940

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

fMRI temporal clustering analysis in patients with frequent interictal epileptiform discharges: comparison with EEG-driven analysis.

Temporal clustering analysis (TCA) is an exploratory data-driven technique that has been proposed for the analysis of resting fMRI to localise epileptiform activity without need for simultaneous EEG. Conventionally, fMRI of epileptic activity has been limited to those patients with subtle clinical events or frequent interictal epileptiform EEG discharges, requiring simultaneous EEG recording, from which a linear model is derived to make valid statistical inferences from the fMRI data. We sought to evaluate TCA by comparing the results with those of EEG correlated fMRI in eight selected cases. Cases were selected with clear epileptogenic localisation or lateralisation on the basis of concordant EEG and structural MRI findings, in addition to concordant activations seen on EEG-derived fMRI analyses. In three, areas of activation were seen with TCA but none corresponding to the electro-clinical localisation or activations obtained with EEG driven analysis. Temporal clusters were closely coincident with times of maximal head motion. We feel this is a serious confound to this approach and recommend that interpretation of TCA that does not address motion and physiological noise be treated with caution. New techniques to localise epileptogenic activity with fMRI alone require validation with an appropriate independent measure. In the investigation of interictal epileptiform activity, this is best done with simultaneous EEG recording.

Cluster Analysis↗

Confirmation of human protein interaction data by human expression data.

BACKGROUND: With microarray technology the expression of thousands of genes can be measured simultaneously. It is well known that the expression levels of genes of interacting proteins are correlated significantly more strongly in Saccharomyces cerevisiae than those of proteins that are not interacting. The objective of this work is to investigate whether this observation extends to the human genome. RESULTS: We investigated the quantitative relationship between expression levels of genes encoding interacting proteins and genes encoding random protein pairs. Therefore we studied 1369 interacting human protein pairs and human gene expression levels of 155 arrays. We were able to establish a statistically significantly higher correlation between the expression levels of genes whose proteins interact compared to random protein pairs. Additionally we were able to provide evidence that genes encoding proteins belonging to the same GO-class show correlated expression levels. CONCLUSION: This finding is concurrent with the naive hypothesis that the scales of production of interacting proteins are linked because an efficient interaction demands that involved proteins are available to some degree. The goal of further research in this field will be to understand the biological mechanisms behind this observation.

Cluster Analysis↗