PubMed Health⌕ Search

Biomedical subjects

Jaroslav P Novak

Publications and source records attributed to Jaroslav P Novak.

4 recordsLinked to original sources

Generalization of DNA microarray dispersion properties: microarray equivalent of t-distribution.

BACKGROUND: DNA microarrays are a powerful technology that can provide a wealth of gene expression data for disease studies, drug development, and a wide scope of other investigations. Because of the large volume and inherent variability of DNA microarray data, many new statistical methods have been developed for evaluating the significance of the observed differences in gene expression. However, until now little attention has been given to the characterization of dispersion of DNA microarray data. RESULTS: Here we examine the expression data obtained from 682 Affymetrix GeneChips with 22 different types and we demonstrate that the Gaussian (normal) frequency distribution is characteristic for the variability of gene expression values. However, typically 5 to 15% of the samples deviate from normality. Furthermore, it is shown that the frequency distributions of the difference of expression in subsets of ordered, consecutive pairs of genes (consecutive samples) in pair-wise comparisons of replicate experiments are also normal. We describe a consecutive sampling method, which is employed to calculate the characteristic function approximating standard deviation and show that the standard deviation derived from the consecutive samples is equivalent to the standard deviation obtained from individual genes. Finally, we determine the boundaries of probability intervals and demonstrate that the coefficients defining the intervals are independent of sample characteristics, variability of data, laboratory conditions and type of chips. These coefficients are very closely correlated with Student's t-distribution. CONCLUSION: In this study we ascertained that the non-systematic variations possess Gaussian distribution, determined the probability intervals and demonstrated that the K(alpha) coefficients defining these intervals are invariant; these coefficients offer a convenient universal measure of dispersion of data. The fact that the K(alpha) distributions are so close to t-distribution and independent of conditions and type of arrays suggests that the quantitative data provided by Affymetrix technology give "true" representation of physical processes, involved in measurement of RNA abundance. REVIEWERS: This article was reviewed by Yoav Gilad (nominated by Doron Lancet), Sach Mukherjee (nominated by Sandrine Dudoit) and Amir Niknejad and Shmuel Friedland (nominated by Neil Smalheiser).

Journal Article↗

Variation in fiberoptic bead-based oligonucleotide microarrays: dispersion characteristics among hybridization and biological replicate samples.

BACKGROUND: Gene expression microarray technology continues to evolve and its use has expanded into all areas of biology. However, the high dimensionality of the data makes analysis a difficult challenge. Evaluating measurements and estimating the significance of the observed differences among samples remain important issues that must be addressed for each technology platform. In this work we use a consecutive sampling method to characterize the dispersion patterns of data generated from Illumina fiberoptic bead-based oligonucleotide arrays. RESULTS: To describe general properties of the dispersion we used a linear function SD = a + bY(mean), approximating the standard deviation across arrays (Y(mean) is the mean expression of a given consecutive sample). First we examined three levels of variability: 1) same cell culture, same reverse transcription, duplicate hybridizations; 2) same cell culture, reverse transcription replicates; 3) parallel cultures. Each higher level is expected to introduce a new source of variability. We observed minor differences in the constant term: the mean values are 3.5, 3.1 and 3.5, respectively. However, the mean coefficient b increased from 0.045 to 0.147 and 0.133. We compared the coefficients derived from the consecutive sampling to those obtained from the standard deviation of individual gene expressions and found them in good agreement. In the second experiment samples we detected 11 genes with systematically different expressions between the experiment samples treated with glucose oxidase and controls and corroborated the selection using the Mann-Whitney and other tests. We also compared the consecutive sampling and coincidence method to t-test: the average percentage of consistency was above 80 for the former and below 50 for the latter. CONCLUSION: Our results indicate that the consecutive sampling method and standard deviation function provide a convenient description of the overall dispersion of Illumina arrays. We observed that the constant term of the standard deviation function is at average approximately the same for duplicate hybridization as for the assays with additional sources of variability. Furthermore, among the genes affected by glucose oxidase treatment we identified 6 genes in oxidative stress pathways and 5 genes involved in DNA repair. Finally, we noted that the consecutive sampling and coincidence test provide, under given conditions, more consistent results than the t-test. REVIEWERS: This article was reviewed by Alexander Karpikov (nominated by Mark Gerstein), Jordan King and Eugene V. Koonin.

Journal Article↗

Distinct pattern of lung gene expression in the Cftr-KO mice developing spontaneous lung disease compared with their littermate controls.

Cystic fibrosis (CF) is caused by a defect in the CF transmembrane conductance regulator (CFTR) protein that functions as a chloride channel. Dysfunction of the CFTR protein results in salty sweat, pancreatic insufficiency, intestinal obstruction, male infertility, and severe pulmonary disease. Most of the morbidity and mortality of CF patients results from pulmonary complications. Differences in susceptibility to bacterial infection and variable degree of CF lung disease among CF patients remain unexplained. Many phenotypic expressions of the disease do not directly correlate with the type of mutation in the Cftr gene. Using a unique CF mouse model that mimics aspects of human CF lung disease, we analyzed the differential gene expression pattern between the normal lungs of wild-type mice (WT) and the affected lungs of CFTR knockout mice (KO). Using microarray analysis followed by quantitation of candidate gene mRNA and protein expression, we identified many interesting genes involved in the development of CF lung disease in mice. These findings point to distinct mechanisms of gene expression regulation between mice with CF and control mice.

Animals↗

Characterization of variability in large-scale gene expression data: implications for study design.

Large-scale gene expression measurement techniques provide a unique opportunity to gain insight into biological processes under normal and pathological conditions. To interpret the changes in expression profiles for thousands of genes, we face the nontrivial problem of understanding the significance of these changes. In practice, the sources of background variability in expression data can be divided into three categories: technical, physiological, and sampling. To assess the relative importance of these sources of background variation, we generated replicate gene expression profiles on high-density Affymetrix GeneChip oligonucleotide arrays, using either identical RNA samples or RNA samples obtained under similar biological states. We derived a novel measure of dispersion in two-way comparisons, using a linear characteristic function. When comparing expression profiles from replicate tests using the same RNA sample (a test for technical variability), we observed a level of dispersion similar to the pattern obtained with RNA samples from replicate cultures of the same cell line (a test for physiological variability). On the other hand, a higher level of dispersion was observed when tissue samples of different animals were compared (an example of sampling variability). This implies that, in experiments in which samples from different subjects are used, the variation induced by the stimulus may be masked by non-stimuli-related differences in the subjects' biological state. These analyses underscore the need for replica experiments to reliably interpret large-scale expression data sets, even with simple microarray experiments.

Animals↗