PubMed Health⌕ Search

Biomedical subjects

Ka Yee Yeung

Publications and source records attributed to Ka Yee Yeung.

8 recordsLinked to original sources

Bayesian robust inference for differential gene expression in microarrays with multiple samples.

We consider the problem of identifying differentially expressed genes under different conditions using gene expression microarrays. Because of the many steps involved in the experimental process, from hybridization to image analysis, cDNA microarray data often contain outliers. For example, an outlying data value could occur because of scratches or dust on the surface, imperfections in the glass, or imperfections in the array production. We develop a robust Bayesian hierarchical model for testing for differential expression. Errors are modeled explicitly using a t-distribution, which accounts for outliers. The model includes an exchangeable prior for the variances, which allows different variances for the genes but still shrinks extreme empirical variances. Our model can be used for testing for differentially expressed genes among multiple samples, and it can distinguish between the different possible patterns of differential expression when there are three or more samples. Parameter estimation is carried out using a novel version of Markov chain Monte Carlo that is appropriate when the model puts mass on subspaces of the full parameter space. The method is illustrated using two publicly available gene expression data sets. We compare our method to six other baseline and commonly used techniques, namely the t-test, the Bonferroni-adjusted t-test, significance analysis of microarrays (SAM), Efron's empirical Bayes, and EBarrays in both its lognormal-normal and gamma-gamma forms. In an experiment with HIV data, our method performed better than these alternatives, on the basis of between-replicate agreement and disagreement.

Bayes Theorem↗

Donuts, scratches and blanks: robust model-based segmentation of microarray images.

MOTIVATION: Inner holes, artifacts and blank spots are common in microarray images, but current image analysis methods do not pay them enough attention. We propose a new robust model-based method for processing microarray images so as to estimate foreground and background intensities. The method starts with a very simple but effective automatic gridding method, and then proceeds in two steps. The first step applies model-based clustering to the distribution of pixel intensities, using the Bayesian Information Criterion (BIC) to choose the number of groups up to a maximum of three. The second step is spatial, finding the large spatially connected components in each cluster of pixels. The method thus combines the strengths of the histogram-based and spatial approaches. It deals effectively with inner holes in spots and with artifacts. It also provides a formal inferential basis for deciding when the spot is blank, namely when the BIC favors one group over two or three. RESULTS: We apply our methods for gridding and segmentation to cDNA microarray images from an HIV infection experiment. In these experiments, our method had better stability across replicates than a fixed-circle segmentation method or the seeded region growing method in the SPOT software, without introducing noticeable bias when estimating the intensities of differentially expressed genes. AVAILABILITY: spotSegmentation, an R language package implementing both the gridding and segmentation methods is available through the Bioconductor project (http://www.bioconductor.org). The segmentation method requires the contributed R package MCLUST for model-based clustering (http://cran.us.r-project.org). CONTACT: fraley@stat.washington.edu.

Algorithms↗

Bayesian model averaging: development of an improved multi-class, gene selection and classification tool for microarray data.

MOTIVATION: Selecting a small number of relevant genes for accurate classification of samples is essential for the development of diagnostic tests. We present the Bayesian model averaging (BMA) method for gene selection and classification of microarray data. Typical gene selection and classification procedures ignore model uncertainty and use a single set of relevant genes (model) to predict the class. BMA accounts for the uncertainty about the best set to choose by averaging over multiple models (sets of potentially overlapping relevant genes). RESULTS: We have shown that BMA selects smaller numbers of relevant genes (compared with other methods) and achieves a high prediction accuracy on three microarray datasets. Our BMA algorithm is applicable to microarray datasets with any number of classes, and outputs posterior probabilities for the selected genes and models. Our selected models typically consist of only a few genes. The combination of high accuracy, small numbers of genes and posterior probabilities for the predictions should make BMA a powerful tool for developing diagnostics from expression data. AVAILABILITY: The source codes and datasets used are available from our Supplementary website.

Algorithms↗

From co-expression to co-regulation: how many microarray experiments do we need?

BACKGROUND: Cluster analysis is often used to infer regulatory modules or biological function by associating unknown genes with other genes that have similar expression patterns and known regulatory elements or functions. However, clustering results may not have any biological relevance. RESULTS: We applied various clustering algorithms to microarray datasets with different sizes, and we evaluated the clustering results by determining the fraction of gene pairs from the same clusters that share at least one known common transcription factor. We used both yeast transcription factor databases (SCPD, YPD) and chromatin immunoprecipitation (ChIP) data to evaluate our clustering results. We showed that the ability to identify co-regulated genes from clustering results is strongly dependent on the number of microarray experiments used in cluster analysis and the accuracy of these associations plateaus at between 50 and 100 experiments on yeast data. Moreover, the model-based clustering algorithm MCLUST consistently outperforms more traditional methods in accurately assigning co-regulated genes to the same clusters on standardized data. CONCLUSIONS: Our results are consistent with respect to independent evaluation criteria that strengthen our confidence in our results. However, when one compares ChIP data to YPD, the false-negative rate is approximately 80% using the recommended p-value of 0.001. In addition, we showed that even with large numbers of experiments, the false-positive rate may exceed the true-positive rate. In particular, even when all experiments are included, the best results produce clusters with only a 28% true-positive rate using known gene transcription factor interactions.

Algorithms↗

Bcl-2 overexpression leads to increases in suppressor of cytokine signaling-3 expression in B cells and de novo follicular lymphoma.

The t(14;18)(q32;q21), resulting in deregulated expression of B-cell-leukemia/lymphoma-2 (Bcl-2), represents the genetic hallmark in human follicular lymphomas. Substantial evidence supports the hypothesis that the t(14;18) and Bcl-2 overexpression are necessary but not solely responsible for neoplastic transformation and require cooperating genetic derangements for neoplastic transformation to occur. To investigate genes that cooperate with Bcl-2 to influence cellular signaling pathways important for neoplastic transformation, we used oligonucleotide microarrays to determine differential gene expression patterns in CD19+ B cells isolated from Emu-Bcl-2 transgenic mice and wild-type littermate control mice. Fifty-seven genes were induced and 94 genes were repressed by > or =2-fold in Emu-Bcl-2 transgenic mice (P < 0.05). The suppressor of cytokine signaling-3 (SOCS3) gene was found to be overexpressed 5-fold in B cells from Emu-Bcl-2 transgenic mice. Overexpression of Bcl-2 in both mouse embryo fibroblast-1 and hematopoietic cell lines resulted in induction of SOCS3 protein, suggesting a Bcl-2-associated mechanism underlying SOCS3 induction. Immunohistochemistry with SOCS3 antisera on tissue from a cohort of patients with de novo follicular lymphoma revealed marked overexpression of SOCS3 protein that, within the follicular center cell region, was limited to neoplastic follicular lymphoma cells and colocalized with Bcl-2 expression in 9 of 12 de novo follicular lymphoma cases examined. In contrast, SOCS3 protein expression was not detected in the follicular center cell region of benign hyperplastic tonsil tissue. These data suggest that Bcl-2 overexpression leads to the induction of activated signal transducer and activator of transcription 3 (STAT3) and to the induction of SOCS3, which may contribute to the pathogenesis of follicular lymphoma.

Animals↗

Multiclass classification of microarray data with repeated measurements: application to cancer.

Prediction of the diagnostic category of a tissue sample from its gene-expression profile and selection of relevant genes for class prediction have important applications in cancer research. We have developed the uncorrelated shrunken centroid (USC) and error-weighted, uncorrelated shrunken centroid (EWUSC) algorithms that are applicable to microarray data with any number of classes. We show that removing highly correlated genes typically improves classification results using a small set of genes.

Algorithms↗

Clustering gene-expression data with repeated measurements.

Clustering is a common methodology for the analysis of array data, and many research laboratories are generating array data with repeated measurements. We evaluated several clustering algorithms that incorporate repeated measurements, and show that algorithms that take advantage of repeated measurements yield more accurate and more stable clusters. In particular, we show that the infinite mixture model-based approach with a built-in error model produces superior results.

Algorithms↗

Transcriptional analyses of Barrett's metaplasia and normal upper GI mucosae.

Over the last two decades, the incidence of esophageal adenocarcinoma (EA) has increased dramatically in the US and Western Europe. It has been shown that EAs evolve from premalignant Barrett's esophagus (BE) tissue by a process of clonal expansion and evolution. However, the molecular phenotype of the premalignant metaplasia, and its relationship to those of the normal upper gastrointestinal (GI) mucosae, including gastric, duodenal, and squamous epithelium of the esophagus, has not been systematically characterized. Therefore, we used oligonucleotide-based microarrays to characterize gene expression profiles in each of these tissues. The similarity of BE to each of the normal tissues was compared using a series of computational approaches. Our analyses included esophageal squamous epithelium, which is present at the same anatomic site and exposed to similar conditions as Barrett's epithelium, duodenum that shares morphologic similarity to Barrett's epithelium, and adjacent gastric epithelium. There was a clear distinction among the expression profiles of gastric, duodenal, and squamous epithelium whereas the BE profiles showed considerable overlap with normal tissues. Furthermore, we identified clusters of genes that are specific to each of the tissues, to the Barrett's metaplastic epithelia, and a cluster of genes that was distinct between squamous and non-squamous epithelia.

Barrett Esophagus↗