PubMed Health⌕ Search

Biomedical subjects

Lev Klebanov

Publications and source records attributed to Lev Klebanov.

6 recordsLinked to original sources

Treating expression levels of different genes as a sample in microarray data analysis: is it worth a risk?

One of the prevailing ideas in the literature on microarray data analysis is to pool the expression measures across genes and treat them as a sample drawn from some distribution. Several universal laws were proposed to analytically describe this distribution. This idea raises a number of concerns. The expression levels of genes are not identically distributed random variables so that treating them as a sample amounts to sampling from a mixture of equally weighted distributions, each being associated with a different gene. The expression levels of different genes are heavily dependent random variables so that the law of large numbers and statistical goodness-of-fit tests are normally inapplicable to this kind of data. This dependence represents a very serious pitfall in microarray data analysis.

Gene Expression Profiling↗

A new type of stochastic dependence revealed in gene expression data.

Modern methods of microarray data analysis are biased towards selecting those genes that display the most pronounced differential expression. The magnitude of differential expression does not necessarily indicate biological significance and other criteria are needed to supplement the information on differential expression. Three large sets of microarray data on childhood leukemia were analyzed by an original method introduced in this paper. A new type of stochastic dependence between expression levels in gene pairs was deciphered by our analysis. This modulation-like unidirectional dependence between expression signals arises when the expression of a "gene-modulator'' is stochastically proportional to that of a "gene-driver''. A total of more than 35% of all pairs formed from 12550 genes were conservatively estimated to belong to this type. There are genes that tend to form Type A relationships with the overwhelming majority of genes. However, this picture is not static: the composition of Type A gene pairs may undergo dramatic changes when comparing two phenotypes. The ability to identify genes that act as ;;modulators'' provides a potential strategy of prioritizing candidate genes.

Child↗

Correlation between gene expression levels and limitations of the empirical bayes methodology for finding differentially expressed genes.

Stochastic dependence between gene expression levels in microarray data is of critical importance for the methods of statistical inference that resort to pooling test statistics across genes. The empirical Bayes methodology in the nonparametric and parametric formulations, as well as closely related methods employing a two-component mixture model, represent typical examples. It is frequently assumed that dependence between gene expressions (or associated test statistics) is sufficiently weak to justify the application of such methods for selecting differentially expressed genes. By applying resampling techniques to simulated and real biological data sets, we have studied a potential impact of the correlation between gene expression levels on the statistical inference based on the empirical Bayes methodology. We report evidence from these analyses that this impact may be quite strong, leading to a high variance of the number of differentially expressed genes. This study also pinpoints specific components of the empirical Bayes method where the reported effect manifests itself.

Journal Article↗

The effects of normalization on the correlation structure of microarray data.

BACKGROUND: Stochastic dependence between gene expression levels in microarray data is of critical importance for the methods of statistical inference that resort to pooling test-statistics across genes. It is frequently assumed that dependence between genes (or tests) is sufficiently weak to justify the proposed methods of testing for differentially expressed genes. A potential impact of between-gene correlations on the performance of such methods has yet to be explored. RESULTS: The paper presents a systematic study of correlation between the t-statistics associated with different genes. We report the effects of four different normalization methods using a large set of microarray data on childhood leukemia in addition to several sets of simulated data. Our findings help decipher the correlation structure of microarray data before and after the application of normalization procedures. CONCLUSION: A long-range correlation in microarray data manifests itself in thousands of genes that are heavily correlated with a given gene in terms of the associated t-statistics. By using normalization methods it is possible to significantly reduce correlation between the t-statistics computed for different genes. Normalization procedures affect both the true correlation, stemming from gene interactions, and the spurious correlation induced by random noise. When analyzing real world biological data sets, normalization procedures are unable to completely remove correlation between the test statistics. The long-range correlation structure also persists in normalized data.

Algorithms↗

Multivariate search for differentially expressed gene combinations.

BACKGROUND: To identify differentially expressed genes, it is standard practice to test a two-sample hypothesis for each gene with a proper adjustment for multiple testing. Such tests are essentially univariate and disregard the multidimensional structure of microarray data. A more general two-sample hypothesis is formulated in terms of the joint distribution of any sub-vector of expression signals. RESULTS: By building on an earlier proposed multivariate test statistic, we propose a new algorithm for identifying differentially expressed gene combinations. The algorithm includes an improved random search procedure designed to generate candidate gene combinations of a given size. Cross-validation is used to provide replication stability of the search procedure. A permutation two-sample test is used for significance testing. We design a multiple testing procedure to control the family-wise error rate (FWER) when selecting significant combinations of genes that result from a successive selection procedure. A target set of genes is composed of all significant combinations selected via random search. CONCLUSIONS: A new algorithm has been developed to identify differentially expressed gene combinations. The performance of the proposed search-and-testing procedure has been evaluated by computer simulations and analysis of replicated Affymetrix gene array data on age-related changes in gene expression in the inner ear of CBA mice.

Aging↗