PubMed Health⌕ Search

PubMed · 14581611

Robust singular value decomposition analysis of microarray data.

Abstract

In microarray data there are a number of biological samples, each assessed for the level of gene expression for a typically large number of genes. There is a need to examine these data with statistical techniques to help discern possible patterns in the data. Our technique applies a combination of mathematical and statistical methods to progressively take the data set apart so that different aspects can be examined for both general patterns and very specific effects. Unfortunately, these data tables are often corrupted with extreme values (outliers), missing values, and non-normal distributions that preclude standard analysis. We develop a robust analysis method to address these problems. The benefits of this robust analysis will be both the understanding of large-scale shifts in gene effects and the isolation of particular sample-by-gene effects that might be either unusual interactions or the result of experimental flaws. Our method requires a single pass and does not resort to complex "cleaning" or imputation of the data table before analysis. We illustrate the method with a commercial data set.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Li Liu, Douglas M Hawkins, Sujoy Ghosh, S Stanley Young. 2003-10-27. Robust singular value decomposition analysis of microarray data.. https://doi.org/10.1073/pnas.1733249100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Thermal unfolding simulations of a multimeric protein--transition state and unfolding pathways.

The folding of an oligomeric protein poses an extra challenge to the folding problem because the protein not only has to fold correctly; it has to avoid nonproductive aggregation. We have carried out over 100 molecular dynamics simulations using an implicit solvation model at different temperatures to study the unfolding of one of the smallest known tetramers, p53 tetramerization domain (p53tet). We found that unfolding started with disruption of the native tetrameric hydrophobic core. The transition state for the tetramer to dimer transition was characterized as a diverse ensemble of different structures using Phi value analysis in quantitative agreement with experimental data. Despite the diversity, the ensemble was still native-like with common features such as partially exposed tetramer hydrophobic core and shifts in the dimer-dimer arrangements. After passing the transition state, the secondary and tertiary structures continued to unfold until the primary dimers broke free. The free dimer had little secondary structure left and the final free monomers were random-coil like. Both the transition states and the unfolding pathways from these trajectories were very diverse, in agreement with the new view of protein folding. The multiple simulations showed that the folding of p53tet is a mixture of the framework and nucleation-condensation mechanisms and the folding is coupled to the complex formation. We have also calculated the entropy and effective energy for the different states along the unfolding pathway and found that the tetramerization is stabilized by hydrophobic interactions.

Cluster Analysis↗

A new strategy of cooperativity of biclustering and hierarchical clustering: a case of analyzing yeast genomic microarray datasets.

Hierarchical clustering is difficult to be deployed effectively in finding meaningful subtrees since genes rarely exhibit similar expression pattern across a wide range of conditions. It is also difficult to find a suitable level in cleaving a big hierarchy tree. Biclustering is a promising methodology in the field of the analysis of gene expression data of genechip. Generally it can be employed in identification of gene groups, which show a coherent expression profile across a subset of conditions. But in some cases of biclustering analysis of gene expressions, the genes in one bicluster are involved in more than one functional group, or all genes in one bicluster are involved in unknown functional groups (e.g. pattern VI and VIII in our studies). Then, how to predict the function of genes in these patterns? In the present research, we developed a new strategy of combining both of the clustering methods, hierarchical clustering and biclustering. The reserved conditions in datasets for hierarchical clustering were elicited according to the conditions in biclusters, and after hierarchical clustering, more detailed results in predicting unknown genes in certain patterns were obtained. This strategy of cooperating both of the methods during clustering procedure should be an effective guideline for functional predictions.

Cluster Analysis↗