PubMed Health⌕ Search

Biomedical subjects

John D Storey

Publications and source records attributed to John D Storey.

14 recordsLinked to original sources

The optimal discovery procedure for large-scale significance testing, with applications to comparative microarray experiments.

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the optimal discovery procedure (ODP), which has recently been introduced and theoretically shown to optimally perform multiple significance tests. Whereas existing procedures essentially use data from only one feature at a time, the ODP approach uses the relevant information from the entire data set when testing each feature. In particular, we propose a generally applicable estimate of the ODP for identifying differentially expressed genes in microarray experiments. This microarray method consistently shows favorable performance over five highly used existing methods. For example, in testing for differential expression between two breast cancer tumor types, the ODP provides increases from 72% to 185% in the number of genes called significant at a false discovery rate of 3%. Our proposed microarray method is freely available to academic users in the open-source, point-and-click EDGE software package.

Apoptosis Regulatory Proteins↗

Relaxed significance criteria for linkage analysis.

Linkage analysis involves performing significance tests at many loci located throughout the genome. Traditional criteria for declaring a linkage statistically significant have been formulated with the goal of controlling the rate at which any single false positive occurs, called the genomewise error rate (GWER). As complex traits have become the focus of linkage analysis, it is increasingly common to expect that a number of loci are truly linked to the trait. This is especially true in mapping quantitative trait loci (QTL), where sometimes dozens of QTL may exist. Therefore, alternatives to the strict goal of preventing any single false positive have recently been explored, such as the false discovery rate (FDR) criterion. Here, we characterize some of the challenges that arise when defining relaxed significance criteria that allow for at least one false positive linkage to occur. In particular, we show that the FDR suffers from several problems when applied to linkage analysis of a single trait. We therefore conclude that the general applicability of FDR for declaring significant linkages in the analysis of a single trait is dubious. Instead, we propose a significance criterion that is more relaxed than the traditional GWER, but does not appear to suffer from the problems of the FDR. A generalized version of the GWER is proposed, called GWERk, that allows one to provide a more liberal balance between true positives and false positives at no additional cost in computation or assumptions.

Algorithms↗

A new approach to intensity-dependent normalization of two-channel microarrays.

A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions, which limit their general applicability. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show that a dye-swap is accurate under general assumptions, even under intensity-dependent dye bias, and that a dye-swap removes dye bias from a single pair of samples in general. Based on a flexible model of the relationship between mRNA amount and single-channel fluorescence intensity, we demonstrate the general applicability of a dye-swap approach. We then propose a common array dye-swap (CADS) method for the normalization of two-channel microarrays. We show that CADS removes both dye bias and array-specific effects, and preserves the true differential expression signal for every gene under the assumptions of the model.

Computer Simulation↗

EDGE: extraction and analysis of differential gene expression.

EDGE (Extraction of Differential Gene Expression) is an open source, point-and-click software program for the significance analysis of DNA microarray experiments. EDGE can perform both standard and time course differential expression analysis. The functions are based on newly developed statistical theory and methods. This document introduces the EDGE software package.

Algorithms↗

Significance analysis of time course microarray experiments.

Characterizing the genome-wide dynamic regulation of gene expression is important and will be of much interest in the future. However, there is currently no established method for identifying differentially expressed genes in a time course study. Here we propose a significance method for analyzing time course microarray studies that can be applied to the typical types of comparisons and sampling schemes. This method is applied to two studies on humans. In one study, genes are identified that show differential expression over time in response to in vivo endotoxin administration. By using our method, 7,409 genes are called significant at a 1% false-discovery rate level, whereas several existing approaches fail to identify any genes. In another study, 417 genes are identified at a 10% false-discovery rate level that show expression changing with age in the kidney cortex. Here it is also shown that as many as 47% of the genes change with age in a manner more complex than simple exponential growth or decay. The methodology proposed here has been implemented in the freely distributed and open-source edge software package.

Adult↗

Genetic interactions between polymorphisms that affect gene expression in yeast.

Interactions between polymorphisms at different quantitative trait loci (QTLs) are thought to contribute to the genetics of many traits, and can markedly affect the power of genetic studies to detect QTLs. Interacting loci have been identified in many organisms. However, the prevalence of interactions, and the nucleotide changes underlying them, are largely unknown. Here we search for naturally occurring genetic interactions in a large set of quantitative phenotypes--the levels of all transcripts in a cross between two strains of Saccharomyces cerevisiae. For each transcript, we searched for secondary loci interacting with primary QTLs detected by their individual effects. Such locus pairs were estimated to be involved in the inheritance of 57% of transcripts; statistically significant pairs were identified for 225 transcripts. Among these, 67% of secondary loci had individual effects too small to be significant in a genome-wide scan. Engineered polymorphisms in isogenic strains confirmed an interaction between the mating-type locus MAT and the pheromone response gene GPA1. Our results indicate that genetic interactions are widespread in the genetics of transcript levels, and that many QTLs will be missed by single-locus tests but can be detected by two-stage tests that allow for interactions.

Crosses, Genetic↗

Multiple locus linkage analysis of genomewide expression in yeast.

With the ability to measure thousands of related phenotypes from a single biological sample, it is now feasible to genetically dissect systems-level biological phenomena. The genetics of transcriptional regulation and protein abundance are likely to be complex, meaning that genetic variation at multiple loci will influence these phenotypes. Several recent studies have investigated the role of genetic variation in transcription by applying traditional linkage analysis methods to genomewide expression data, where each gene expression level was treated as a quantitative trait and analyzed separately from one another. Here, we develop a new, computationally efficient method for simultaneously mapping multiple gene expression quantitative trait loci that directly uses all of the available data. Information shared across gene expression traits is captured in a way that makes minimal assumptions about the statistical properties of the data. The method produces easy-to-interpret measures of statistical significance for both individual loci and the overall joint significance of multiple loci selected for a given expression trait. We apply the new method to a cross between two strains of the budding yeast Saccharomyces cerevisiae, and estimate that at least 37% of all gene expression traits show two simultaneous linkages, where we have allowed for epistatic interactions. Pairs of jointly linking quantitative trait loci are identified with high confidence for 170 gene expression traits, where it is expected that both loci are true positives for at least 153 traits. In addition, we are able to show that epistatic interactions contribute to gene expression variation for at least 14% of all traits. We compare the proposed approach to an exhaustive two-dimensional scan over all pairs of loci. Surprisingly, we demonstrate that an exhaustive two-dimensional scan is less powerful than the sequential search used here. In addition, we show that a two-dimensional scan does not truly allow one to test for simultaneous linkage, and the statistical significance measured from this existing method cannot be interpreted among many traits.

Chromosome Mapping↗

Longitudinal transcriptional analysis of developing neointimal vascular occlusion and pulmonary hypertension in rats.

Pneumonectomized rats injected with the alkaloid toxin, monocrotaline, develop progressive neointimal pulmonary vascular obliteration and pulmonary hypertension resulting in right ventricular failure and death. The antiproliferative immunosuppressant, triptolide, attenuates neointimal formation and pulmonary hypertension in this disease model (Faul JL, Nishimura T, Berry GJ, Benson GV, Pearl RG, and Kao PN. Am J Respir Crit Care Med 162: 2252-2258, 2000). Pneumonectomized rats, injected with monocrotaline on day 7, were killed at days 14, 21, 28, and 35 for measurements of physiology and gene expression patterns. These data were compared with pneumonectomized, monocrotaline-injected animals that received triptolide from day 5 to day 35. The hypothesis was tested that a group of functionally related genes would be significantly coexpressed during the development of disease and downregulated in response to treatment. Transcriptional analysis using total lung RNA was performed on replicate animals for each experimental time point with exploratory data analysis followed by statistical significance analysis. Marked, statistically significant increases in proteases (particularly derived from mast cells) were noted that parallel the development of vascular obliteration and pulmonary hypertension. Mast-cell-derived proteases may play a role in regulating the development of neointimal pulmonary vascular occlusion and pulmonary hypertension in response to injury.

Animals↗

Statistical significance for genomewide studies.

With the increase in genomewide experiments and the sequencing of multiple genomes, the analysis of large data sets has become commonplace in biology. It is often the case that thousands of features in a genomewide data set are tested against some null hypothesis, where a number of features are expected to be significant. Here we propose an approach to measuring statistical significance in these genomewide studies based on the concept of the false discovery rate. This approach offers a sensible balance between the number of true and false positives that is automatically calibrated and easily interpreted. In doing so, a measure of statistical significance called the q value is associated with each tested feature. The q value is similar to the well known p value, except it is a measure of significance in terms of the false discovery rate rather than the false positive rate. Our approach avoids a flood of false positive results, while offering a more liberal criterion than what has been used in genome scans for linkage.

Algorithms↗

Genome-wide analysis of mRNA translation profiles in Saccharomyces cerevisiae.

We have analyzed the translational status of each mRNA in rapidly growing Saccharomyces cerevisiae. mRNAs were separated by velocity sedimentation on a sucrose gradient, and 14 fractions across the gradient were analyzed by quantitative microarray analysis, providing a profile of ribosome association with mRNAs for thousands of genes. For most genes, the majority of mRNA molecules were associated with ribosomes and presumably engaged in translation. This systematic approach enabled us to recognize genes with unusual behavior. For 43 genes, most mRNA molecules were not associated with ribosomes, suggesting that they may be translationally controlled. For 53 genes, including GCN4, CPA1, and ICY2, three genes for which translational control is known to play a key role in regulation, most mRNA molecules were associated with a single ribosome. The number of ribosomes associated with mRNAs increased with increasing length of the putative protein-coding sequence, consistent with longer transit times for ribosomes translating longer coding sequences. The density at which ribosomes were distributed on each mRNA (i.e., the number of ribosomes per unit ORF length) was well below the maximum packing density for nearly all mRNAs, consistent with initiation as the rate-limiting step in translation. Global analysis revealed an unexpected correlation: Ribosome density decreases with increasing ORF length. Models to account for this surprising observation are discussed.

Genome, Fungal↗

Precision and functional specificity in mRNA decay.

Posttranscriptional processing of mRNA is an integral component of the gene expression program. By using DNA microarrays, we precisely measured the decay of each yeast mRNA, after thermal inactivation of a temperature-sensitive RNA polymerase II. The half-lives varied widely, ranging from approximately 3 min to more than 90 min. We found no simple correlation between mRNA half-lives and ORF size, codon bias, ribosome density, or abundance. However, the decay rates of mRNAs encoding groups of proteins that act together in stoichiometric complexes were generally closely matched, and other evidence pointed to a more general relationship between physiological function and mRNA turnover rates. The results provide strong evidence that precise control of the decay of each mRNA is a fundamental feature of the gene expression program in yeast.

Glycolysis↗

In vivo regulation of human skeletal muscle gene expression by thyroid hormone.

Thyroid hormones are key regulators of metabolism that modulate transcription via nuclear receptors. Hyperthyroidism is associated with increased metabolic rate, protein breakdown, and weight loss. Although the molecular actions of thyroid hormones have been studied thoroughly, their pleiotropic effects are mediated by complex changes in expression of an unknown number of target genes. Here, we measured patterns of skeletal muscle gene expression in five healthy men treated for 14 days with 75 microg of triiodothyronine, using 24,000 cDNA element microarrays. To analyze the data, we used a new statistical method that identifies significant changes in expression and estimates the false discovery rate. The 381 up-regulated genes were involved in a wide range of cellular functions including transcriptional control, mRNA maturation, protein turnover, signal transduction, cellular trafficking, and energy metabolism. Only two genes were down-regulated. Most of the genes are novel targets of thyroid hormone. Cluster analysis of triiodothyronine-regulated gene expression among 19 different human tissues or cell lines revealed sets of coregulated genes that serve similar biologic functions. These results define molecular signatures that help to understand the physiology and pathophysiology of thyroid hormone action.

Administration, Oral↗