PubMed Health⌕ Search

Biomedical subjects

R Sásik

Publications and source records attributed to R Sásik.

13 recordsLinked to original sources

Microarray truths and consequences.

For many, analysis of a microarray experiment starts with a spreadsheet of expression levels. While great attention is duly paid to RNA extraction, preparation and hybridization, relatively little care is devoted to extraction of expression levels from the fluorescent image. By delegating this step to a click of the mouse the exact extraction process is masked and researchers may be unwittingly compromising their data. In this review, we describe the most common mistakes committed on the path from the image to the spreadsheet and their impact on data quality. Remedies are further proposed for most of the popular microarray platforms in use today.

Oligonucleotide Array Sequence Analysis↗

Extracting transcriptional events from temporal gene expression patterns during Dictyostelium development.

MOTIVATION: The DNA microarray technology can generate a large amount of data describing the time-course of gene expression. These data, when properly interpreted, can yield a great deal of information concerning differential gene expression during development. Much current effort in bioinformatics has been devoted to the analysis of gene expression data, usually via some 'clustering analysis' on the raw data in some abstract high dimensional space. Here, we describe a method where we first 'process' the raw time-course data using a simple biologically based kinetic model of gene expression. This allows us to reduce the vast data to a few vital attributes characterizing each expression profile, e.g. the times of the onset and cessation of the expression of the developmentally regulated genes. These vital attributes can then be trivially clustered by visual inspection to reveal biologically significant effects. RESULTS: We have applied this approach to microarray expression data from samples isolated every 2 h throughout the 24 h developmental program of Dictyostelium discoideum. mRNA accumulation patterns for 50 developmental genes were found to fit the kinetic model with a p-value of 0.05 or better. Transcription of these genes appears to be initiated in bursts at well-defined periods during development, in a manner suggestive of a dependent sequence. This approach can be applied to analyses of other temporal gene expression patterns, including those of the cell cycle.

Algorithms↗

Statistical analysis of high-density oligonucleotide arrays: a multiplicative noise model.

MOTIVATION: High-density oligonucleotide arrays (GeneChip, Affymetrix, Santa Clara, CA) have become a standard research tool in many areas of biomedical research. They quantitatively monitor the expression of thousands of genes simultaneously by measuring fluorescence from gene-specific targets or probes. The relationship between signal intensities and transcript abundance as well as normalization issues have been the focus of much recent attention (Hill et al., 2001; Chudin et al., 2002; Naef et al., 2002a). It is desirable that a researcher has the best possible analytical tools to make the most of the information that this powerful technology has to offer. At present there are three analytical methods available: the newly released Affymetrix Microarray Suite 5.0 (AMS) software that accompanies the GeneChip product, the method of Li and Wong (LW; Li and Wong, 2001), and the method of Naef et al. (FN; Naef et al., 2001). The AMS method is tailored for analysis of a single microarray, and can therefore be used with any experimental design. The LW method on the other hand depends on a large number of microarrays in an experiment and cannot be used for an isolated microarray, and the FN method is particular to paired microarrays, such as resulting from an experiment in which each 'treatment' sample has a corresponding 'control' sample. Our focus is on analysis of experiments in which there is a series of samples. In this case only the AMS, LW, and the method described in this paper can be used. The present method is model-based, like the LW method, but assumes multiplicative not additive noise, and employs elimination of statistically significant outliers for improved results. Unlike LW and AMS, we do not assume probe-specific background (measured by the so-called mismatch probes). Rather, we assume uniform background, whose level is estimated using both the mismatch and perfect match probe intensities. RESULTS: We present a new method for GeneChip analysis, based on a statistical model with multiplicative noise. We demonstrated that this method yields results superior to those obtained by the Affymetrix Microarray Suite 5.0 software and to those obtained by the model-based method of Li and Wong (Li and Wong, 2001). The present method eliminates the hard-to-interpret negative expression indices, and the binary 'presence' calls (present or absent) are replaced by the statistical significance (p-value) of gene expression. We have found that thresholding the p-values at the (0.1)(16)-level produces about the same number of 'present' calls as the AMS software. By testing our method on a pair of replicate GeneChips (hybridized with the same cRNA), we found that 95.6% of data points lie within the 1.25-fold interval. In other words, our method had a 4.4% type I error rate at the 1.25-fold level. The error rate of the LW method was 15%, and that of the AMS method was 29%. There were no points outside the 2-fold interval with the present method. Analysis of variance (ANOVA) of another experiment with multiple replicates shows that this reduction of variance is not accompanied by a corresponding reduction of signal. On the contrary, the signal-to-noise ratio (as measured by the distribution of F-statistics) of the present method is on average 3.4-times better than that of AMS, and 1.4-times better than that of Li and Wong.

Algorithms↗

Percolation clustering: a novel approach to the clustering of gene expression patterns in Dictyostelium development.

We present a novel approach to the clustering of gene expression patterns based on the mutual connectivity of the patterns. Unlike certain widely used methods (e.g., self-organizing maps and K-means) which essentially force gene expression data into a fixed number of predetermined clustering structures, our approach aims to reveal the natural tendency of the data to cluster, in analogy to the physical phenomenon of percolation. The approach is probabilistic in nature, and as such accommodates the possibility that one gene participates in multiple clusters. The result is cast in terms of the connectivity of each gene to a certain number of (significant) clusters. A computationally efficient algorithm is developed to implement our approach. Performance of the method is illustrated by clustering both constructed data and gene expression data obtained from Dictyostelium development.

Algorithms↗

Analytical approach to time lag in binary nucleation.

We present an analytical formula for the time required to establish steady state in a nucleating binary system. To test our solution, we evaluate the time lag for a range of activities of both components at the vapor-liquid transition, and show that our result is in much better agreement with a purely numerical simulation than other available analytical formulas, which overestimate the time lag by factors of from 2 to 200.

Journal Article↗