PubMed Health⌕ Search

PubMed · 11867085

Variable selection and pattern recognition with gene expression data generated by the microarray technology.

Abstract

Lack of adequate statistical methods for the analysis of microarray data remains the most critical deterrent to uncovering the true potential of these promising techniques in basic and translational biological studies. The popular practice of drawing important biological conclusions from just one replicate (slide) should be discouraged. In this paper, we discuss some modern trends in statistical analysis of microarray data with a special focus on statistical classification (pattern recognition) and variable selection. In addressing these issues we consider the utility of some distances between random vectors and their nonparametric estimates obtained from gene expression data. Performance of the proposed distances is tested by computer simulations and analysis of gene expression data on two different types of human leukemia. In experimental settings, the error rate is estimated by cross-validation, while a control sample is generated in computer simulation experiments aimed at testing the proposed gene selection procedures and associated classification rules.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Szabo, K Boucher, W L Carroll, L B Klebanov, A D Tsodikov, A Y Yakovlev. 2002. Variable selection and pattern recognition with gene expression data generated by the microarray technology.. https://doi.org/10.1016/s0025-5564(01)00103-1

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Non-parametric estimators of a monotonic dose-response curve and bootstrap confidence intervals.

In this paper we consider study designs which include a placebo and an active control group as well as several dose groups of a new drug. A monotonically increasing dose-response function is assumed, and the objective is to estimate a dose with equivalent response to the active control group, including a confidence interval for this dose. We present different non-parametric methods to estimate the monotonic dose-response curve. These are derived from the isotonic regression estimator, a non-negative least squares estimator, and a bias adjusted non-negative least squares estimator using linear interpolation. The different confidence intervals are based upon an approach described by Korn, and upon two different bootstrap approaches. One of these bootstrap approaches is standard, and the second ensures that resampling is done from empiric distributions which comply with the order restrictions imposed. In our simulations we did not find any differences between the two bootstrap methods, and both clearly outperform Korn's confidence intervals. The non-negative least squares estimator yields biased results for moderate sample sizes. The bias adjustment for this estimator works well, even for small and moderate sample sizes, and surprisingly outperforms the isotonic regression method in certain situations.

Computer Simulation↗