Sensitive and quantitative universal Pyrosequencing methylation analysis of CpG sites.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to K A Baggerly.
Explore the source record for details and available documents.
BACKGROUND: A key problem in immunohistochemistry is assessing when two sample histograms are significantly different. One test that is commonly used for this purpose in the univariate case is the chi-squared test. Comparing multivariate distributions is qualitatively harder, as the "curse of dimensionality" means that the number of bins can grow exponentially. For the chi-squared test to be useful, data-dependent binning methods must be employed. An example of how this can be done is provided by the "probability binning" method of Roederer et al. (1,2,3). METHODS: We derive the theoretical distribution of the probability binning statistic, giving it a more rigorous foundation. We show that the null distribution is a scaled chi-square, and show how it can be related to the standard chi-squared statistic. RESULTS: A small simulation shows how the theoretical results can be used to (a) modify the probability binning statistic to make it more sensitive and (b) suggest variant statistics which, while still exploiting the data-dependent strengths of the probability binning procedure, may be easier to work with. CONCLUSIONS: The probability binning procedure effectively uses adaptive binning to locate structure in high-dimensional data. The derivation of a theoretical basis provides a more detailed interpretation of its behavior and renders the probability binning method more flexible.
Application of powerful, high-throughput genomics technologies is becoming more common and these technologies are evolving at a rapid pace. Genomics facilities are being established in major research institutions to produce inexpensive, customized cDNA microarrays that are accessible to researchers in a broad range of fields. These high-throughput platforms have generated a massive onslaught of data, which threatens to overwhelm researchers. Although microarrays show great promise, the technology has not matured to the point of consistently generating robust and reliable data when used in the average laboratory. This article addresses several aspects related to the handling of the deluge of microarray data and extracting reliable information from these data. We review the essential elements of data acquisition, data processing and data analysis, and briefly discuss issues related to the quality, validation and storage of data. Our goal is to point out some of the problems that must be overcome before this promising technology can achieve its full potential.
A major goal of microarray experiments is to determine which genes are differentially expressed between samples. Differential expression has been assessed by taking ratios of expression levels of different samples at a spot on the array and flagging spots (genes) where the magnitude of the fold difference exceeds some threshold. More recent work has attempted to incorporate the fact that the variability of these ratios is not constant. Most methods are variants of Student's t-test. These variants standardize the ratios by dividing by an estimate of the standard deviation of that ratio; spots with large standardized values are flagged. Estimating these standard deviations requires replication of the measurements, either within a slide or between slides, or the use of a model describing what the standard deviation should be. Starting from considerations of the kinetics driving microarray hybridization, we derive models for the intensity of a replicated spot, when replication is performed within and between arrays. Replication within slides leads to a beta-binomial model, and replication between slides leads to a gamma-Poisson model. These models predict how the variance of a log ratio changes with the total intensity of the signal at the spot, independent of the identity of the gene. Ratios for genes with a small amount of total signal are highly variable, whereas ratios for genes with a large amount of total signal are fairly stable. Log ratios are scaled by the standard deviations given by these functions, giving model-based versions of Studentization. An example is given.
Unequal sister chromatid exchange has been proposed as one of several possible mechanisms for gene amplification resulting in tandemly repeated sequences on chromosomes. Two requirements for testing this hypothesis are analytical observations and a mathematical model. Recently observations were reported for the number of tandemly repeated sequences on chromosomes of cells growing in the presence of a toxic drug and the mechanism was proposed to be unequal sister chromatid exchange. We now develop a mathematical model of this process based on the following hypotheses, (i) the extent of slippage between paired sister chromatids is a random variable with geometric distribution, (ii) the number of crossover sites is a random variable with a Poisson distribution, and (iii) cells with less than a threshold number of copies of an essential gene are eliminated when grown in selective conditions. Iterating the model at successive cell divisions results in a Markov chain with a denumerable infinity of states. The resulting distributions of gene copy number per cell at a particular population size are compared to published data on the CAD gene in BHK cells growing in the presence of the drug PALA (Smith et al., 1990, Cell, 63, 1219). The mathematical model can reproduce the observed means and standard deviations of gene copy number per cell and allows construction of confidence region estimates of parameters describing the extent of slippage, density of crossover sites, and strength of selection. An important prediction of the model is that in non-selective conditions the cells with amplified sequences gradually disappear from the population even if they are not at a growth disadvantage, though rare cells with a very large number of amplified sequences might continue to exist. The success of modeling suggests that the proposed mechanism of gene amplification by unequal sister chromatid exchange is consistent with the number of tandemly repeated sequences on chromosomes observed in some circumstances.
This paper presents two new ways of analysing data which may be obtained from pulse labelling a population of cells with bromodeoxyuridine and analysing that population as a function of time with bivariate flow cytometry. The progression of cells is measured by the change in position in the cell cycle, as shown by a change in the mean DNA content of the labelled and unlabelled cells. The particular measures of the mean DNA content used are extensions of the relative movement of the labelled undivided cells, RMlu(t), which was introduced by Begg and co-workers to measure the DNA synthesis time, TS. In general, the relative movement is defined as the mean DNA fluorescence of a population of cells less the DNA fluorescence of the cells in G1 and divided by the difference in DNA fluorescence of the cells in G2 + M and G1. In this paper we examine the relative movements of all the labelled cells and all of the unlabelled cells, denoted RML(t) and RMU(t) respectively. It is found that RML(t) and RMU(t) exhibit clear cyclic behaviour and distinguishable characteristics which depend directly on the transit times (T) of the cell cycle phases, i.e. TG1, TS and TG2 + M. Furthermore, the peak heights of the RMU(t) curve are shown to depend strongly on the growth fraction of the population under consideration. A theoretical treatment of the curves so obtained is presented, and is shown to yield values in close agreement with those from other methods for measuring these transit times and a lower limit to values for the growth fraction of Chinese hamster ovary cells grown in vitro.