PubMed Health⌕ Search

Biomedical subjects

Rebecka Jörnsten

Publications and source records attributed to Rebecka Jörnsten.

3 recordsLinked to original sources

DNA microarray data imputation and significance analysis of differential expression.

MOTIVATION: Significance analysis of differential expression in DNA microarray data is an important task. Much of the current research is focused on developing improved tests and software tools. The task is difficult not only owing to the high dimensionality of the data (number of genes), but also because of the often non-negligible presence of missing values. There is thus a great need to reliably impute these missing values prior to the statistical analyses. Many imputation methods have been developed for DNA microarray data, but their impact on statistical analyses has not been well studied. In this work we examine how missing values and their imputation affect significance analysis of differential expression. RESULTS: We develop a new imputation method (LinCmb) that is superior to the widely used methods in terms of normalized root mean squared error. Its estimates are the convex combinations of the estimates of existing methods. We find that LinCmb adapts to the structure of the data: If the data are heterogeneous or if there are few missing values, LinCmb puts more weight on local imputation methods; if the data are homogeneous or if there are many missing values, LinCmb puts more weight on global imputation methods. Thus, LinCmb is a useful tool to understand the merits of different imputation methods. We also demonstrate that missing values affect significance analysis. Two datasets, different amounts of missing values, different imputation methods, the standard t-test and the regularized t-test and ANOVA are employed in the simulations. We conclude that good imputation alleviates the impact of missing values and should be an integral part of microarray data analysis. The most competitive methods are LinCmb, GMC and BPCA. Popular imputation schemes such as SVD, row mean, and KNN all exhibit high variance and poor performance. The regularized t-test is less affected by missing values than the standard t-test. AVAILABILITY: Matlab code is available on request from the authors.

Algorithms↗

Screening anti-inflammatory compounds in injured spinal cord with microarrays: a comparison of bioinformatics analysis approaches.

Inflammatory responses contribute to secondary tissue damage following spinal cord injury (SCI). A potent anti-inflammatory glucocorticoid, methylprednisolone (MP), is the only currently accepted therapy for acute SCI but its efficacy has been questioned. To search for additional anti-inflammatory compounds, we combined microarray analysis with an explanted spinal cord slice culture injury model. We compared gene expression profiles after treatment with MP, acetaminophen, indomethacin, NS398, and combined cytokine inhibitors (IL-1ra and soluble TNFR). Multiple gene filtering methods and statistical clustering analyses were applied to the multi-dimensional data set and results were compared. Our analysis showed a consistent and unique gene expression profile associated with NS398, the selective cyclooxygenase-2 (COX-2) inhibitor, in which the overall effect of these upregulated genes could be interpreted as neuroprotective. In vivo testing demonstrated that NS398 reduced lesion volumes, unlike MP or acetaminophen, consistent with a predicted physiological effect in spinal cord. Combining explanted spinal cultures, microarrays, and flexible clustering algorithms allows us to accelerate selection of compounds for in vivo testing.

Algorithms↗

Simultaneous gene clustering and subset selection for sample classification via MDL.

MOTIVATION: The microarray technology allows for the simultaneous monitoring of thousands of genes for each sample. The high-dimensional gene expression data can be used to study similarities of gene expression profiles across different samples to form a gene clustering. The clusters may be indicative of genetic pathways. Parallel to gene clustering is the important application of sample classification based on all or selected gene expressions. The gene clustering and sample classification are often undertaken separately, or in a directional manner (one as an aid for the other). However, such separation of these two tasks may occlude informative structure in the data. Here we present an algorithm for the simultaneous clustering of genes and subset selection of gene clusters for sample classification. We develop a new model selection criterion based on Rissanen's MDL (minimum description length) principle. For the first time, an MDL code length is given for both explanatory variables (genes) and response variables (sample class labels). The final output of the proposed algorithm is a sparse and interpretable classification rule based on cluster centroids or the closest genes to the centroids. RESULTS: Our algorithm for simultaneous gene clustering and subset selection for classification is applied to three publicly available data sets. For all three data sets, we obtain sparse and interpretable classification models based on centroids of clusters. At the same time, these models give competitive test error rates as the best reported methods. Compared with classification models based on single gene selections, our rules are stable in the sense that the number of clusters has a small variability and the centroids of the clusters are well correlated (or consistent) across different cross validation samples. We also discuss models where the centroids of clusters are replaced with the genes closest to the centroids. These models show comparable test error rates to models based on single gene selection, but are more sparse as well as more stable. Moreover, we comment on how the inclusion of a classification criterion affects the gene clustering, bringing out class informative structure in the data. AVAILABILITY: The methods presented in this paper have been implemented in the R language. The source code is available from the first author.

Algorithms↗