PubMed Health⌕ Search

Biomedical subjects

Paul Fogel

Publications and source records attributed to Paul Fogel.

4 recordsLinked to original sources

Inferential, robust non-negative matrix factorization analysis of microarray data.

MOTIVATION: Modern methods such as microarrays, proteomics and metabolomics often produce datasets where there are many more predictor variables than observations. Research in these areas is often exploratory; even so, there is interest in statistical methods that accurately point to effects that are likely to replicate. Correlations among predictors are used to improve the statistical analysis. We exploit two ideas: non-negative matrix factorization methods that create ordered sets of predictors; and statistical testing within ordered sets which is done sequentially, removing the need for correction for multiple testing within the set. RESULTS: Simulations and theory point to increased statistical power. Computational algorithms are described in detail. The analysis and biological interpretation of a real dataset are given. In addition to the increased power, the benefit of our method is that the organized gene lists are likely to lead better understanding of the biology. AVAILABILITY: An SAS JMP executable script is available from http://www.niss.org/irMF

Algorithms↗

Analysis and prediction of combinatorial chemistry synthesis and screening data.

The goal of combinatorial chemistry is to simultaneously synthesize sets of compounds possessing properties that are then distinguished through screening. As the size of a compound set increases, data analysis becomes more challenging. Analysis of Variance (ANOVA) is an accepted statistical method that offers a straightforward solution to this problem. Two steps encountered by combinatorial scientists appear well suited to ANOVA: the prediction of synthetic outcomes (purity and yield) of set members and the analysis of screening data to identify combinations of reagent inputs that result in molecules with a desired property. To illustrate, a subset of a combinatorial array, referred to as a reaction rehearsal set, is evaluated to create a model predictive of the individual synthetic outcomes of the full matrix. In a second exercise, the biochemical screening data obtained from a combinatorial library is analyzed to identify reagent interactions that result in molecules possessing the sought activity.

Analysis of Variance↗

The Global Error Assessment (GEA) model for the selection of differentially expressed genes in microarray data.

MOTIVATION: Microarray technology has become a powerful research tool in many fields of study; however, the cost of microarrays often results in the use of a low number of replicates (k). Under circumstances where k is low, it becomes difficult to perform standard statistical tests to extract the most biologically significant experimental results. Other more advanced statistical tests have been developed; however, their use and interpretation often remain difficult to implement in routine biological research. The present work outlines a method that achieves sufficient statistical power for selecting differentially expressed genes under conditions of low k, while remaining as an intuitive and computationally efficient procedure. RESULTS: The present study describes a Global Error Assessment (GEA) methodology to select differentially expressed genes in microarray datasets, and was developed using an in vitro experiment that compared control and interferon-gamma treated skin cells. In this experiment, up to nine replicates were used to confidently estimate error, thereby enabling methods of different statistical power to be compared. Gene expression results of a similar absolute expression are binned, so as to enable a highly accurate local estimate of the mean squared error within conditions. The model then relates variability of gene expression in each bin to absolute expression levels and uses this in a test derived from the classical ANOVA. The GEA selection method is compared with both the classical and permutational ANOVA tests, and demonstrates an increased stability, robustness and confidence in gene selection. A subset of the selected genes were validated by real-time reverse transcription-polymerase chain reaction (RT-PCR). All these results suggest that GEA methodology is (i) suitable for selection of differentially expressed genes in microarray data, (ii) intuitive and computationally efficient and (iii) especially advantageous under conditions of low k. AVAILABILITY: The GEA code for R software is freely available upon request to authors.

Algorithms↗

The confirmation rate of primary hits: a predictive model.

HTS data from primary screening are usually analyzed by setting a cutoff for activity, in order to minimize both false-negative and false-positive rates. An alternative approach, based on a calculated probability of being active, is presented here. Given the predicted confirmation rate derived from this probability, the number of primary positives selected for follow-up can be optimized to maximize the number of true positives without picking too many false positives. Typical cutoff-determining methods are more serendipitous in their nature and not easily optimized in an effort to optimize screening efforts. An additional advantage of calculating a probability of being active for each compound screened is that orthogonal mixtures can be deconvoluted without presetting a deconvolution threshold. An important consequence of using the probability of being active with orthogonal mixtures is that individual compound screening results can be recorded irrespective of whether the assays were performed on single compounds or on cocktails.

Biological Assay↗