PubMed Health⌕ Search

Biomedical subjects

Edward Suh

Publications and source records attributed to Edward Suh.

9 recordsLinked to original sources

Confidence intervals for the true classification error conditioned on the estimated error.

Bias and variance for small-sample error estimation are typically posed in terms of statistics for the distributions of the true and estimated errors. On the other hand, a salient practical issue asks, given an error estimate, what can be said about the true error? This question relates to the joint distribution of the true and estimated errors, specifically, the conditional expectation of the true error given the error estimate. A critical issue is that of confidence bounds for the true error given the estimate. We consider the joint distribution of the true error and the estimated error, assuming a random feature-label distribution. From it, we derive the marginal distributions, the conditional expectation of the estimated error given the true error, the conditional expectation of the true error given the estimated error, the conditional variance of the true error given the estimated error, and the 95% upper confidence bound for the true error given the estimated error. Numerous classification and estimation rules are considered across a number of models. Massive simulation is used for continuous models and analytic results are derived for discrete classification. We also consider a breast-cancer study to illustrate how the theory might be applied in practice. Although specific results depend on the classification rule, error-estimation rule, and model, some general trends are seen: (I) if the true error is small (large), then the conditional estimated error is generally high (low)-biased; (II) the conditional expected true error tends to be larger (smaller) than the estimated error for small (large) estimated errors; and (III) the confidence bounds tend to be well above the estimated error for low error estimates, becoming much less so for large estimates.

Algorithms↗

Two-locus genome-wide linkage scan for prostate cancer susceptibility genes with an interaction effect.

Prostate cancer represents a significant worldwide public health burden. Epidemiological and genetic epidemiological studies have consistently provided data supporting the existence of inherited prostate cancer susceptibility genes. Segregation analyses of prostate cancer suggest that a multigene model may best explain familial clustering of this disease. Therefore, modeling gene-gene interactions in linkage analysis may improve the power to detect chromosomal regions harboring these disease susceptibility genes. In this study, we systematically screened for prostate cancer linkage by modeling two-locus gene-gene interactions for all possible pairs of loci across the genome in 426 prostate cancer families from Johns Hopkins Hospital, University of Michigan, University of Umeå, and University of Tampere. We found suggestive evidence for an epistatic interaction for six sets of loci (target chromosome-wide/reference marker-specific P< or =0.0001). Evidence for these interactions was found in two independent subsets from within the 426 families. While the validity of these results requires confirmation from independent studies and the identification of the specific genes underlying this linkage evidence, our approach of systematically assessing gene-gene interactions across the entire genome represents a promising alternative approach for gene identification for prostate cancer.

Aged↗

The interaction of four genes in the inflammation pathway significantly predicts prostate cancer risk.

It is widely hypothesized that the interactions of multiple genes influence individual risk to prostate cancer. However, current efforts at identifying prostate cancer risk genes primarily rely on single-gene approaches. In an attempt to fill this gap, we carried out a study to explore the joint effect of multiple genes in the inflammation pathway on prostate cancer risk. We studied 20 genes in the Toll-like receptor signaling pathway as well as several cytokines. For each of these genes, we selected and genotyped haplotype-tagging single nucleotide polymorphisms (SNP) among 1,383 cases and 780 controls from the CAPS (CAncer Prostate in Sweden) study population. A total of 57 SNPs were included in the final analysis. A data mining method, multifactor dimensionality reduction, was used to explore the interaction effects of SNPs on prostate cancer risk. Interaction effects were assessed for all possible n SNP combinations, where n = 2, 3, or 4. For each n SNP combination, the model providing lowest prediction error among 100 cross-validations was chosen. The statistical significance levels of the best models in each n SNP combination were determined using permutation tests. A four-SNP interaction (one SNP each from IL-10, IL-1RN, TIRAP, and TLR5) had the lowest prediction error (43.28%, P = 0.019). Our ability to analyze a large number of SNPs in a large sample size is one of the first efforts in exploring the effect of high-order gene-gene interactions on prostate cancer risk, and this is an important contribution to this new and quickly evolving field.

Case-Control Studies↗

Optimal number of features as a function of sample size for various classification rules.

MOTIVATION: Given the joint feature-label distribution, increasing the number of features always results in decreased classification error; however, this is not the case when a classifier is designed via a classification rule from sample data. Typically (but not always), for fixed sample size, the error of a designed classifier decreases and then increases as the number of features grows. The potential downside of using too many features is most critical for small samples, which are commonplace for gene-expression-based classifiers for phenotype discrimination. For fixed sample size and feature-label distribution, the issue is to find an optimal number of features. RESULTS: Since only in rare cases is there a known distribution of the error as a function of the number of features and sample size, this study employs simulation for various feature-label distributions and classification rules, and across a wide range of sample and feature-set sizes. To achieve the desired end, finding the optimal number of features as a function of sample size, it employs massively parallel computation. Seven classifiers are treated: 3-nearest-neighbor, Gaussian kernel, linear support vector machine, polynomial support vector machine, perceptron, regular histogram and linear discriminant analysis. Three Gaussian-based models are considered: linear, nonlinear and bimodal. In addition, real patient data from a large breast-cancer study is considered. To mitigate the combinatorial search for finding optimal feature sets, and to model the situation in which subsets of genes are co-regulated and correlation is internal to these subsets, we assume that the covariance matrix of the features is blocked, with each block corresponding to a group of correlated features. Altogether there are a large number of error surfaces for the many cases. These are provided in full on a companion website, which is meant to serve as resource for those working with small-sample classification. AVAILABILITY: For the companion website, please visit http://public.tgen.org/tamu/ofs/ CONTACT: e-dougherty@ee.tamu.edu.

Algorithms↗

Gene clustering based on clusterwide mutual information.

Cluster analysis of gene-wide expression data from DNA microarray hybridization studies has proved to be a useful tool for identifying biologically relevant groupings of genes and constructing gene regulatory networks. The motivation for considering mutual information is its capacity to measure a general dependence among gene random variables. We propose a novel clustering strategy based on minimizing mutual information among gene clusters. Simulated annealing is employed to solve the optimization problem. Bootstrap techniques are employed to get more accurate estimates of mutual information when the data sample size is small. Moreover, we propose to combine the mutual information criterion and traditional distance criteria such as the Euclidean distance and the fuzzy membership metric in designing the clustering algorithm. The performances of the new clustering methods are compared with those of some existing methods, using both synthesized data and experimental data. It is seen that the clustering algorithm based on a combined metric of mutual information and fuzzy membership achieves the best performance. The supplemental material is available at www.gspsnap.tamu.edu/gspweb/zxb/glioma_zxb.

Algorithms↗

Early origin and recent expansion of Plasmodium falciparum.

The emergence of virulent Plasmodium falciparum in Africa within the past 6000 years as a result of a cascade of changes in human behavior and mosquito transmission has recently been hypothesized. Here, we provide genetic evidence for a sudden increase in the African malaria parasite population about 10,000 years ago, followed by migration to other regions on the basis of variation in 100 worldwide mitochondrial DNA sequences. However, both the world and some regional populations appear to be older (50,000 to 100,000 years old), suggesting an earlier wave of migration out of Africa, perhaps during the Pleistocene migration of human beings.

Africa↗

Microarray reveals differences in both tumors and vascular specific gene expression in de novo CD5+ and CD5- diffuse large B-cell lymphomas.

Malignant lymphoma is a heterogeneous disease with different clinical features. Among diffuse large B-cell lymphomas (DLBCLs), a unique subtype has been identified recently based on cell surface marker CD5 and clinicopathological features. These de novo CD5(+) DLBCLs account for approximately 10% of all of the DLBCLs and have poorer prognosis. To additionally understand this subtype of DLBCLs at the molecular level and to find genes that are differentially expressed in de novo CD5(+) DLBCLs, CD5(-) DLBCLs, and mantle cell lymphomas, which also have poor prognosis, we performed gene expression profiling using cDNA microarray technology. Data from a total of 9 samples of CD5(-) DLBCLs, 11 samples of de novo CD5(+) DLBCLs, and 10 samples of mantle cell lymphomas were acquired. A series of genes were identified that distinguish these three types of lymphomas. Among DLBCL cases, integrin beta1 and/or CD36 adhesion molecules were overexpressed in most cases of CD5(+) DLBCL. An immunohistochemical confirmation study revealed that integrin beta1 was expressed on lymphoma cells, which may account for the high extranodal involvement and poor prognosis of CD5(+) DLBCLs. In contrast, CD36 was overexpressed on vascular endothelia in CD5(+) DLBCLs, although there was no difference in vascularity detected by von Wilbrand factor antibody between CD5(+) and CD5(-) DLBCLs. Those results suggest that CD5(+) and CD5(-) DLBCLs have different gene expression signatures in both tumor cells and their vascular systems.

Aged↗

Identification of signature genes by microarray for acute myeloid leukemia without maturation and acute promyelocytic leukemia with t(15;17)(q22;q12)(PML/RARalpha).

Acute myeloid leukemia (AML) has distinct subgroups characterized by different maturation and specific chromosomal translocation. In order to gain insight into the gene expression activities in AML, we carried out a gene expression profiling study with 21 AML samples using cDNA microarrays, focusing on acute promyelocytic leukemia with specific translocation t(15;17)(q22;q12) [French-American-British or FAB-M3 with t(15;17)] and AML without maturation (FAB-M1) characterized by morphologically and phenotypically immature AML blasts and no recurrent chromosomal abnormalities. Using a multivariate sigma-classifier algorithm, we identified 33 strong feature genes that distinguish FAB-M3 with t(15;17) from other AML samples, and 24 strong feature genes that classify FAB-M1. A direct comparison between FAB-M3 with t(15;17) and FAB-M1 led to selection of 13 strong feature genes. Those genes include some known to be related to leukemogenesis and cell differentiation. RIN1, a gene in the ras pathway, was up-regulated in FAB-M3 with t(15;17). Growth factor-binding protein 2 gene was down-regulated in FAB-M1. Huntingtin gene was up-regulated in FAB-M1. Others include syndecan 4, interleukin-2 receptor beta, folate receptor beta, low affinity immunoglobulin gamma, Fc receptor IIC precursor, insulin-like growth factor binding protein 2, and myeloperoxidase, which are involved in cell differentiation. Overexpression of myeloperoxidase in FAB-M3 cells with t(15;17) compared to FAB-M1 cells is consistent with the conventional cytochemical staining pattern. Thus, the study revealed that a morphologically-defined FAB-M1 subtype has a distinct gene expression signature that contributes to its cell differentiation and proliferation as well as FAB-M3 with a recurrent cytogenetic abnormality t(15;17)(q22;q12).

Algorithms↗

Parametric and semiparametric approaches to testing for seasonal trend in serial count data.

We present two tests for seasonal trend in monthly incidence data. The first approach uses a penalized likelihood to choose the number of harmonic terms to include in a parametric harmonic model (which includes time trends and autogression as well as seasonal harmonic terms) and then tests for seasonality using a parametric bootstrap test. The second approach uses a semiparametric regression model to test for seasonal trend. In the semiparametric model, the seasonal pattern is modeled nonparametrically, parametric terms are included for autoregressive effects and a linear time trend, and a parametric bootstrap test is used to test for seasonality. For both procedures, a null distribution is generated under a null Poisson model with time trends and autoregression parameters. We apply the methods to skin melanoma incidence rates collected by the surveillance, epidemiology, and end results (SEER) program of the National Cancer Institute, and perform simulation studies to evaluate the type I error rate and power for the two procedures. These simulations suggest that both procedures are alpha-level procedures. In addition, the harmonic model/bootstrap test had similar or larger power than the semiparametric model/bootstrap test for a wide range of alternatives, and the harmonic model/bootstrap test is much easier to implement. Thus, we recommend the harmonic model/bootstrap test for the analysis of seasonal incidence data.

Journal Article↗