PubMed Health⌕ Search

Biomedical subjects

In-Beum Lee

Publications and source records attributed to In-Beum Lee.

10 recordsLinked to original sources

Multi-model statistical process monitoring and diagnosis of a sequencing batch reactor.

Biological processes exhibit different behavior depending on the influent loads, temperature, microorganism activity, and so on. It has been shown that a combination of several models can provide a suitable approach to model such processes. In the present study, we developed a multiple statistical model approach for the monitoring of biological batch processes. The proposed method consists of four main components: (1) multiway principal component analysis (MPCA) to reduce the dimensionality of data and to remove collinearity; (2) multiple models with a posterior probability for modeling different operating regions; (3) local batch monitoring by the T(2)- and Q-statistics of the specific local model; and (4) a new discrimination measure (DM) to identify when the system has shifted to a new operating condition. Under this approach, local monitoring by multiple models divides the entire historical data set into separate regions, which are then modeled separately. Then, these local regions can be supervised separately, leading to more effective batch monitoring. The proposed method is applied to a pilot-scale 80-L sequencing batch reactor (SBR) for biological wastewater treatment. This SBR is characterized by nonstationary, batchwise, and multiple operation modes. The results obtained for the pilot-scale SBR indicate that the proposed method has the ability to model multiple operating conditions, to identify various operating regions, and also to determine whether the biosystem has shifted to a new operating condition. Our findings show that the local monitoring approach can give more reliable and higher resolution monitoring results than the global model.

Biodegradation, Environmental↗

Integrated framework of nonlinear prediction and process monitoring for complex biological processes.

Bioprocesses and biosystems have nonlinear and multiple operation patterns depending on the influent loads, temperatures, the activity of microorganisms, and other factors. In this paper, an integrated framework of nonlinear modeling and process monitoring methods is developed for a complex biological process. The proposed method is based on modeling by fuzzy partial least squares (FPLS) and on process monitoring by a statistical decomposition, which is suitable for predicting and supervising a nonlinear biological process. Case studies in the bio-simulated process and industrial biological plant show that the proposed method can give superior prediction and monitoring performance in complex biological plants compared to other linear and nonlinear methods, since it can effectively capture the nonlinear causal relationship within the biosystem. This gives us the integrated framework that is able to both model and monitor the nonlinear bioprocess simultaneously.

Algorithms↗

On-line adaptive and nonlinear process monitoring of a pilot-scale sequencing batch reactor.

This article describes the application of on-line nonlinear monitoring of a sequencing batch reactor (SBR). Three-way batch data of SBR are unfolded batch-wisely, and then a adaptive and nonlinear multivariate monitoring method is used to capture the nonlinear characteristics of normal batches. The approach is successfully applied to an 80 L SBR for biological wastewater treatment, where the SBR poses an interesting challenge in view of process monitoring since it is characterized by nonstationary, batchwise, multistage, and nonlinear dynamics. In on-line batch monitoring, the developed adaptive and nonlinear process monitoring method can effectively capture the nonlinear relationship among process variables of a biological process in a SBR. The results of this pilot-scale SBR monitoring system using simple on-line measurements clearly demonstrated that the adaptive and nonlinear monitoring technique showed lower false alarm rate and physically meaningful, that is, robust monitoring results.

Bioreactors↗

Multiple detection of food-borne pathogenic bacteria using a novel 16S rDNA-based oligonucleotide signature chip.

There have been many attempts to develop sensitive and accurate techniques for the detection and diagnosis of pathogenic bacteria using nucleic acid-based technology. To achieve efficient multiple detection of seven selected food-borne pathogens, we assessed the respective 16S rDNA pathogen specific sequences using an oligonucleotide-based signature array. Strategic optimal design of specific capture probes was achieved by using the characteristic first variable region. To assess the specificity of this pathogen detection system, we employed a two-step experimental strategy. Under conditions established through experiments with chemically synthesized model targets comprising both conserved and variable regions of 16S rDNA, we confirmed the validity of this system using real 16S rDNA targets. Detection with real targets was successfully performed using our system, and better specificity was obtained compared to experiments with model targets. Moreover, the subtypes of Vibrio pathogens were successfully classified. We developed a two-dimensional visualization plot tool for positive control and specific spots, which allowed facile and minute differentiation between spot intensities. Repeated array formats were employed to ensure experimental uniformity, and included the statistical p-value criterion for pathogen discrimination. The present results thus indicate that our novel oligonucleotide-based signature chip detection system can be employed for the effective detection of multiple pathogens.

Bacteria↗

Gene selection and classification from microarray data using kernel machine.

The discrimination of cancer patients (including subtypes) based on gene expression data is a critical problem with clinical ramifications. Central to solving this problem is the issue of how to extract the most relevant genes from the several thousand genes on a typical microarray. Here, we propose a methodology that can effectively select an informative subset of genes and classify the subtypes (or patients) of disease using the selected genes. We employ a kernel machine, kernel Fisher discriminant analysis (KFDA), for discrimination and use the derivatives of the kernel function to perform gene selection. Using a modified form of KFDA in the minimum squared error (MSE) sense and the gradients of the kernel functions, we construct an effective gene selection criterion. We assess the performance of the proposed methodology by applying it to three gene expression datasets: leukemia dataset, breast cancer dataset and colon cancer dataset. Using a few informative genes, the proposed method accurately and reliably classified cancer subtypes (or patients). Also, through a comparison study, we verify the reliability of the gene selection and discrimination results.

Algorithms↗

Enhanced process monitoring of fed-batch penicillin cultivation using time-varying and multivariate statistical analysis.

On-line monitoring of penicillin cultivation processes is crucial to the safe production of high-quality products. In the past, multiway principal component analysis (MPCA), a multivariate projection method, has been widely used to monitor batch and fed-batch processes. However, when MPCA is used for on-line batch monitoring, the future behavior of each new batch must be inferred up to the end of the batch operation at each time and the batch lengths must be equalized. This represents a major shortcoming because predicting the future observations without considering the dynamic relationships may distort the data information, leading to false alarms. In this paper, a new statistical batch monitoring approach based on variable-wise unfolding and time-varying score covariance structures is proposed in order to overcome the drawbacks of conventional MPCA and obtain better monitoring performance. The proposed method does not require prediction of the future values while the dynamic relations of data are preserved by using time-varying score covariance structures, and can be used to monitor batch processes in which the batch length varies. The proposed method was used to detect and identify faults in the fed-batch penicillin cultivation process, for four different fault scenarios. The simulation results clearly demonstrate the power and advantages of the proposed method in comparison to MPCA.

Algorithms↗

Nonlinear modeling and adaptive monitoring with fuzzy and multivariate statistical methods in biological wastewater treatment plants.

A new approach to nonlinear modeling and adaptive monitoring using fuzzy principal component regression (FPCR) is proposed and then applied to a real wastewater treatment plant (WWTP) data set. First, principal component analysis (PCA) is used to reduce the dimensionality of data and to remove collinearity. Second, the adaptive credibilistic fuzzy-c-means method is used to appropriately monitor diverse operating conditions based on the PCA score values. Then a new adaptive discrimination monitoring method is proposed to distinguish between a large process change and a simple fault. Third, a FPCR method is proposed, where the Takagi-Sugeno-Kang (TSK) fuzzy model is employed to model the relation between the PCA score values and the target output to avoid the over-fitting problem with original variables. Here, the rule bases, the centers and the widths of TSK fuzzy model are found by heuristic methods. The proposed FPCR method is applied to predict the output variable, the reduction of chemical oxygen demand in the full-scale WWTP. The result shows that it has the ability to model the nonlinear process and multiple operating conditions and is able to identify various operating regions and discriminate between a sustained fault and a simple fault (or abnormalities) occurring within the process data.

Algorithms↗

New gene selection method for classification of cancer subtypes considering within-class variation.

In this work we propose a new method for finding gene subsets of microarray data that effectively discriminates subtypes of disease. We developed a new criterion for measuring the relevance of individual genes by using mean and standard deviation of distances from each sample to the class centroid in order to treat the well-known problem of gene selection, large within-class variation. Also this approach has the advantage that it is applicable not only to binary classification but also to multiple classification problems. We demonstrated the performance of the method by applying it to the publicly available microarray datasets, leukemia (two classes) and small round blue cell tumors (four classes). The proposed method provides a very small number of genes compared with the previous methods without loss of discriminating power and thus it can effectively facilitate further biological and clinical researches.

Acute Disease↗

Optimal approach for classification of acute leukemia subtypes based on gene expression data.

The classification of cancer subtypes, which is critical for successful treatment, has been studied extensively with the use of gene expression profiles from oligonucleotide chips or cDNA microarrays. Various pattern recognition methods have been successfully applied to gene expression data. However, these methods are not optimal, rather they are high-performance classifiers that emphasize only classification accuracy. In this paper, we propose an approach for the construction of the optimal linear classifier using gene expression data. Two linear classification methods, linear discriminant analysis (LDA) and discriminant partial least-squares (DPLS), are applied to distinguish acute leukemia subtypes. These methods are shown to give satisfactory accuracy. Moreover, we determined optimally the number of genes participating in the classification (a remarkably small number compared to previous results) on the basis of the statistical significance test. Thus, the proposed method constructs the optimal classifier that is composed of a small size predictor and provides high accuracy.

Algorithms↗

Discovery of differentially expressed genes related to histological subtype of hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is one of the most common human malignancies in the world. To identify the histological subtype-specific genes of HCC, we analyzed the gene expression profile of 10 HCC patients by means of cDNA microarray. We proposed a systematic approach for determining the discriminatory genes and revealing the biological phenomena of HCC with cDNA microarray data. First, normalization of cDNA microarray data was performed to reduce or minimize systematic variations. On the basis of the suitably normalized data, we identified specific genes involved in histological subtype of HCC. Two classification methods, Fisher's discriminant analysis (FDA) and support vector machine (SVM), were used to evaluate the reliability of the selected genes and discriminate the histological subtypes of HCC. This study may provide a clue for the needs of different chemotherapy and the reason for heterogeneity of the clinical responses according to histological subtypes.

Algorithms↗