PubMed Health⌕ Search

Biomedical subjects

Seungchan Kim

Publications and source records attributed to Seungchan Kim.

9 recordsLinked to original sources

Gene expression profile in multiple sclerosis patients and healthy controls: identifying pathways relevant to disease.

Multiple sclerosis (MS) and other T cell-mediated autoimmune diseases develop in individuals carrying a complex susceptibility trait, probably following exposure to various environmental triggers. Owing to the presumed weak influence of single genes on disease predisposition and the recognized genetic heterogeneity of autoimmune disorders in humans, candidate gene searches in MS have been difficult. In an attempt to identify molecular markers indicative of disease status rather than susceptibility genes for MS, we show that gene expression profiling of peripheral blood mononuclear cells by cDNA microarrays can distinguish MS patients from healthy controls. Our findings support the concept that the activation of autoreactive T cells is of primary importance for this complex organ-specific disorder and prompt further investigations on gene expression in peripheral blood cells aimed at characterizing disease phenotypes.

Adult↗

Corrected small-sample estimation of the Bayes error.

MOTIVATION: A major problem of pattern classification is estimation of the Bayes error when only small samples are available. One way to estimate the Bayes error is to design a classifier based on some classification rule applied to sample data, estimate the error of the designed classifier, and then use this estimate as an estimate of the Bayes error. Relative to the Bayes error, the expected error of the designed classifier is biased high, and this bias can be severe with small samples. RESULTS: This paper provides a correction for the bias by subtracting a term derived from the representation of the estimation error. It does so for Boolean classifiers, these being defined on binary features. Although the general theory applies to any Boolean classifier, a model is introduced to reduce the number of parameters. A key point is that the expected correction is conservative. Properties of the corrected estimate are studied via simulation. The correction applies to binary predictors because they are mathematically identical to Boolean classifiers. In this context the correction is adapted to the coefficient of determination, which has been used to measure nonlinear multivariate relations between genes and design genetic regulatory networks. An application using gene-expression data from a microarray experiment is provided on the website http://gspsnap.tamu.edu/smallsample/ (user:'smallsample', password:'smallsample)').

Algorithms↗

Microarray reveals differences in both tumors and vascular specific gene expression in de novo CD5+ and CD5- diffuse large B-cell lymphomas.

Malignant lymphoma is a heterogeneous disease with different clinical features. Among diffuse large B-cell lymphomas (DLBCLs), a unique subtype has been identified recently based on cell surface marker CD5 and clinicopathological features. These de novo CD5(+) DLBCLs account for approximately 10% of all of the DLBCLs and have poorer prognosis. To additionally understand this subtype of DLBCLs at the molecular level and to find genes that are differentially expressed in de novo CD5(+) DLBCLs, CD5(-) DLBCLs, and mantle cell lymphomas, which also have poor prognosis, we performed gene expression profiling using cDNA microarray technology. Data from a total of 9 samples of CD5(-) DLBCLs, 11 samples of de novo CD5(+) DLBCLs, and 10 samples of mantle cell lymphomas were acquired. A series of genes were identified that distinguish these three types of lymphomas. Among DLBCL cases, integrin beta1 and/or CD36 adhesion molecules were overexpressed in most cases of CD5(+) DLBCL. An immunohistochemical confirmation study revealed that integrin beta1 was expressed on lymphoma cells, which may account for the high extranodal involvement and poor prognosis of CD5(+) DLBCLs. In contrast, CD36 was overexpressed on vascular endothelia in CD5(+) DLBCLs, although there was no difference in vascularity detected by von Wilbrand factor antibody between CD5(+) and CD5(-) DLBCLs. Those results suggest that CD5(+) and CD5(-) DLBCLs have different gene expression signatures in both tumor cells and their vascular systems.

Aged↗

Identification of signature genes by microarray for acute myeloid leukemia without maturation and acute promyelocytic leukemia with t(15;17)(q22;q12)(PML/RARalpha).

Acute myeloid leukemia (AML) has distinct subgroups characterized by different maturation and specific chromosomal translocation. In order to gain insight into the gene expression activities in AML, we carried out a gene expression profiling study with 21 AML samples using cDNA microarrays, focusing on acute promyelocytic leukemia with specific translocation t(15;17)(q22;q12) [French-American-British or FAB-M3 with t(15;17)] and AML without maturation (FAB-M1) characterized by morphologically and phenotypically immature AML blasts and no recurrent chromosomal abnormalities. Using a multivariate sigma-classifier algorithm, we identified 33 strong feature genes that distinguish FAB-M3 with t(15;17) from other AML samples, and 24 strong feature genes that classify FAB-M1. A direct comparison between FAB-M3 with t(15;17) and FAB-M1 led to selection of 13 strong feature genes. Those genes include some known to be related to leukemogenesis and cell differentiation. RIN1, a gene in the ras pathway, was up-regulated in FAB-M3 with t(15;17). Growth factor-binding protein 2 gene was down-regulated in FAB-M1. Huntingtin gene was up-regulated in FAB-M1. Others include syndecan 4, interleukin-2 receptor beta, folate receptor beta, low affinity immunoglobulin gamma, Fc receptor IIC precursor, insulin-like growth factor binding protein 2, and myeloperoxidase, which are involved in cell differentiation. Overexpression of myeloperoxidase in FAB-M3 cells with t(15;17) compared to FAB-M1 cells is consistent with the conventional cytochemical staining pattern. Thus, the study revealed that a morphologically-defined FAB-M1 subtype has a distinct gene expression signature that contributes to its cell differentiation and proliferation as well as FAB-M3 with a recurrent cytogenetic abnormality t(15;17)(q22;q12).

Algorithms↗

Inference from clustering with application to gene-expression microarrays.

There are many algorithms to cluster sample data points based on nearness or a similarity measure. Often the implication is that points in different clusters come from different underlying classes, whereas those in the same cluster come from the same class. Stochastically, the underlying classes represent different random processes. The inference is that clusters represent a partition of the sample points according to which process they belong. This paper discusses a model-based clustering toolbox that evaluates cluster accuracy. Each random process is modeled as its mean plus independent noise, sample points are generated, the points are clustered, and the clustering error is the number of points clustered incorrectly according to the generating random processes. Various clustering algorithms are evaluated based on process variance and the key issue of the rate at which algorithmic performance improves with increasing numbers of experimental replications. The model means can be selected by hand to test the separability of expected types of biological expression patterns. Alternatively, the model can be seeded by real data to test the expected precision of that output or the extent of improvement in precision that replication could provide. In the latter case, a clustering algorithm is used to form clusters, and the model is seeded with the means and variances of these clusters. Other algorithms are then tested relative to the seeding algorithm. Results are averaged over various seeds. Output includes error tables and graphs, confusion matrices, principal-component plots, and validation measures. Five algorithms are studied in detail: K-means, fuzzy C-means, self-organizing maps, hierarchical Euclidean-distance-based and correlation-based clustering. The toolbox is applied to gene-expression clustering based on cDNA microarrays using real data. Expression profile graphics are generated and error analysis is displayed within the context of these profile graphics. A large amount of generated output is available over the web.

Computational Biology↗

Strong feature sets from small samples.

For small samples, classifier design algorithms typically suffer from overfitting. Given a set of features, a classifier must be designed and its error estimated. For small samples, an error estimator may be unbiased but, owing to a large variance, often give very optimistic estimates. This paper proposes mitigating the small-sample problem by designing classifiers from a probability distribution resulting from spreading the mass of the sample points to make classification more difficult, while maintaining sample geometry. The algorithm is parameterized by the variance of the spreading distribution. By increasing the spread, the algorithm finds gene sets whose classification accuracy remains strong relative to greater spreading of the sample. The error gives a measure of the strength of the feature set as a function of the spread. The algorithm yields feature sets that can distinguish the two classes, not only for the sample data, but for distributions spread beyond the sample data. For linear classifiers, the topic of the present paper, the classifiers are derived analytically from the model, thereby providing an enormous savings in computation time. The algorithm is applied to cancer classification via cDNA microarrays. In particular, the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the algorithm is used to find gene sets whose expressions can be used to classify BRCA1 and BRCA2 tumors.

Breast Neoplasms↗

Probabilistic Boolean Networks: a rule-based uncertainty model for gene regulatory networks.

MOTIVATION: Our goal is to construct a model for genetic regulatory networks such that the model class: (i) incorporates rule-based dependencies between genes; (ii) allows the systematic study of global network dynamics; (iii) is able to cope with uncertainty, both in the data and the model selection; and (iv) permits the quantification of the relative influence and sensitivity of genes in their interactions with other genes. RESULTS: We introduce Probabilistic Boolean Networks (PBN) that share the appealing rule-based properties of Boolean networks, but are robust in the face of uncertainty. We show how the dynamics of these networks can be studied in the probabilistic context of Markov chains, with standard Boolean networks being special cases. Then, we discuss the relationship between PBNs and Bayesian networks--a family of graphical models that explicitly represent probabilistic relationships between variables. We show how probabilistic dependencies between a gene and its parent genes, constituting the basic building blocks of Bayesian networks, can be obtained from PBNs. Finally, we present methods for quantifying the influence of genes on other genes, within the context of PBNs. Examples illustrating the above concepts are presented throughout the paper.

Cell Cycle↗

Identification of combination gene sets for glioma classification.

One goal for the gene expression profiling of cancer tissues is to identify signature genes that robustly distinguish different types or grades of tumors. Such signature genes would ideally provide a molecular basis for classification and also yield insight into the molecular events underlying different cancer phenotypes. This study applies a recently developed algorithm to identify not only single classifier genes but also gene sets (combinations) for use as glioma classifiers. Classifier genes identified by this algorithm are shown to be strong features by conservatively and collectively considering the misclassification errors of the feature sets. Applying this approach to a test set of 25 patients, we have identified the best single genes and two- to three-gene combinations for distinguishing four types of glioma: (a) oligodendroglioma; (b) anaplastic oligodendroglioma; (c) anaplastic astrocytoma; and (d) glioblastoma multiforme. Some of the identified genes, such as insulin-like growth factor-binding protein 2, have been confirmed to be associated with one of the tumor types. Using combinations of genes, the classification error rate can be significantly lowered. In many instances, neither of the individual genes of a two-gene set performs well as an accurate classifier, but the combination of the two genes forms a robust classifier with a small error rate. Two-gene and three-gene combinations thus provide robust classifiers possessing the potential to translate expression microarray results into diagnostic histopathological assays for clinical utilization.

Algorithms↗