PubMed Health⌕ Search

Biomedical subjects

Johan A Westerhuis

Publications and source records attributed to Johan A Westerhuis.

10 recordsLinked to original sources

Centering, scaling, and transformations: improving the biological information content of metabolomics data.

BACKGROUND: Extracting relevant biological information from large data sets is a major challenge in functional genomics research. Different aspects of the data hamper their biological interpretation. For instance, 5000-fold differences in concentration for different metabolites are present in a metabolomics data set, while these differences are not proportional to the biological relevance of these metabolites. However, data analysis methods are not able to make this distinction. Data pretreatment methods can correct for aspects that hinder the biological interpretation of metabolomics data sets by emphasizing the biological information in the data set and thus improving their biological interpretability. RESULTS: Different data pretreatment methods, i.e. centering, autoscaling, pareto scaling, range scaling, vast scaling, log transformation, and power transformation, were tested on a real-life metabolomics data set. They were found to greatly affect the outcome of the data analysis and thus the rank of the, from a biological point of view, most important metabolites. Furthermore, the stability of the rank, the influence of technical errors on data analysis, and the preference of data analysis methods for selecting highly abundant metabolites were affected by the data pretreatment method used prior to data analysis. CONCLUSION: Different pretreatment methods emphasize different aspects of the data and each pretreatment method has its own merits and drawbacks. The choice for a pretreatment method depends on the biological question to be answered, the properties of the data set and the data analysis method selected. For the explorative analysis of the validation data set used in this study, autoscaling and range scaling performed better than the other pretreatment methods. That is, range scaling and autoscaling were able to remove the dependence of the rank of the metabolites on the average concentration and the magnitude of the fold changes and showed biologically sensible results after PCA (principal component analysis).In conclusion, selecting a proper data pretreatment method is an essential step in the analysis of metabolomics data and greatly affects the metabolites that are identified to be the most important.

Cluster Analysis↗

Inline monitoring of butane isomers adsorption on MFI using near-infrared spectroscopy: drift correction in time based experiments.

Near-infrared (NIR) spectroscopy is used to monitor online a large variety of processes. Hydrocarbons with their strong NIR spectral signature are good candidate analytes. For this work, the sorption data are measured in a manometric setup coupled with online NIR spectroscopy, to monitor the bulk composition. The assessment of time based results faces a baseline stability problem. The goal of this article is to study the robustness of different spectral preprocessing methods when dealing with time based data. In this study, it was found that for time based experiments it is necessary to perform drift correction on the spectra combined with a water band correction. For the calibration experiments, which only last few seconds, offset correction and drift correction performed equally well.

Adsorption↗

Tackling calibration problems of spectroscopic analysis in high-throughput experimentation.

High-throughput experimentation and screening methods are changing work flows and creating new possibilities in biochemistry, organometallic chemistry, and catalysis. However, many high-throughput systems rely on off-line chromatography methods that shift the bottleneck to the analysis stage. On-line or at-line spectroscopic analysis is an attractive alternative. It is fast, noninvasive, and nondestructive and requires no sample handling. The disadvantage is that spectroscopic calibration is time-consuming and complex. Ideally, the calibration model should give reliable predictions while keeping the number of calibration samples to a minimum. In this paper, we employ the net analyte signal approach to build a calibration model for Fourier transform near-infrared measurements, using a minimum number of calibration samples based on blank samples. This approach fits very well to high-throughput setups. With this approach, we can reduce the number of calibration samples to the number of chemical components in the system. Thus, the question is no longer how many but which type of calibration samples should one include in the model to obtain reliable predictions. Various calibration models are tested using Monte Carlo simulations, and the results are compared with experimental data for palladium-catalyzed Heck cross-coupling.

Journal Article↗

Quantitative Raman reaction monitoring using the solvent as internal standard.

Despite its potential, the use of Raman spectroscopy for real-time quantitative reaction monitoring is still rather limited. The problems of fluorescence, laser instability, low intensities, and the inner filter effect often outscore the advantages as narrow bands, the use of glass fibers, and low scattering of water and glass. In this paper, we present real-time quantitative monitoring of the catalyzed Heck reaction by using the solvent as internal standard. In this way, all multiplicative distortions, e.g., laser intensity variations or absorbance of the laser light, can be corrected for. We also show that a limited amount of fluorescence does not hamper the analysis. Finally, we present a new method to correct for the inner filter effect, i.e., the absorbance of Raman scattered light by the reaction medium. Simultaneous absorption measurements of the reaction mixture enable accurate correction of Raman signals for the inner filter effect. Thus, for reaction monitoring applications, a Raman spectrometer should be equipped with an absorbance measurement device.

Acrylates↗

New background correction method for liquid chromatography with diode array detection, infrared spectroscopic detection and Raman spectroscopic detection.

A new method to eliminate the background spectrum (EBS) during analyte elution in column liquid chromatography (LC) coupled to spectroscopic techniques is proposed. This method takes into account the shape and also intensity differences of the background eluent spectrum. This allows the EBS method to make a better estimation of the background eluent spectrum during analyte elution. This is an advantage for quantification as well as for identification of analytes. The EBS method uses a two-step procedure. First, the baseline spectra are modeled using a limited number of principal components (PCs). Subsequently, an asymmetric least squares (asLS) regression method is applied using these principal components to correct the measured spectra during elution for the background contribution. The asymmetric least squares regression needs one parameter, the asymmetry factor p. This asymmetry factor determines relative weight of positive and negative residuals. Simulations are performed to test the EBS method in well-defined situations. The effect of spectral noise on the performance and the sensitivity of the EBS method for the value of the asymmetry factorp is tested. Two applications of the EBS method are discussed. In the first application, the goal is to extract the analyte spectrum from an LC-Raman analysis. In this case, the EBS method facilitates easy identification of unknown analytes using spectral libraries. In a second application, the EBS method is used for baseline correction in LC-diode array detection (DAD) analysis of polymeric standards during a gradient elution separation. It is shown that the EBS method yields a good baseline correction, without the need to perform a blank chromatographic run.

Spectrophotometry, Infrared↗

Model selection and optimal sampling in high-throughput experimentation.

The practical difficulties encountered in analyzing the kinetics of new reactions are considered from the viewpoint of the capabilities of state-of-the-art high-throughput systems. There are three problems. The first problem is that of model selection, i.e., choosing the correct reaction rate law. The second problem is how to obtain good estimates of the reaction parameters using only a small number of samples once a kinetic model is selected. The third problem is how to perform both functions using just one small set of measurements. To solve the first problem, we present an optimal sampling protocol to choose the correct kinetic model for a given reaction, based on T-optimal design. This protocol is then tested for the case of second-order and pseudo-first-order reactions using both experiments and computer simulations. To solve the second problem, we derive the information function for second-order reactions and use this function to find the optimal sampling points for estimating the kinetic constants. The third problem is further complicated by the fact that the optimal measurement times for determining the correct kinetic model differ from those needed to obtain good estimates of the kinetic constants. To solve this problem, we propose a Pareto optimal approach that can be tuned to give the set of best possible solutions for the two criteria. One important advantage of this approach is that it enables the integration of a priori knowledge into the workflow.

Journal Article↗

Kinetic studies of cascade reactions in high-throughput systems.

The application of robotic systems to the study of complex reaction kinetics is considered, using the cascade reaction A --> B --> C as a working example. Practical problems in calculating the rate constants k1 and k2 for the reactions A --> B and B --> C from concentration measurements of CA, CB, or CC are discussed in the light of the symmetry and invertability of the rate equations. A D-optimal analysis is used to determine the points in time and the species that will give the best (i.e., most accurate) results. When exact data are used, the most robust solution results from measuring the pair of concentrations (CA, CC). The system's information function is computed using numeric methods. This function is then used to estimate the amount of information obtainable from a given cascade reaction at any given time. The theoretical findings are compared with experimental results from a set of two-stage cascade experiments monitored using UV-visible spectroscopy. Finally, the pros and cons of using a single reaction sample to estimate both k1 and k2 are discussed.

Journal Article↗

Tracking chemical kinetics in high-throughput systems.

Combinatorial chemistry and high-throughput experimentation (HTE) have revolutionized the pharmaceutical industry-but can chemists truly repeat this success in the fields of catalysis and materials science? We propose to bridge the traditional "discovery" and "optimization" stages in HTE by enabling parallel kinetic analysis of an array of chemical reactions. We present here the theoretical basis to extract concentration profiles from reaction arrays and derive the optimal criteria to follow (pseudo)first-order reactions in time in parallel systems. We use the information vector f and introduce in this context the information gain ratio, chi(r), to quantify the amount of useful information that can be obtained by measuring the extent of a specified reaction r in the array at any given time. Our method is general and independent of the analysis technique, but it is more effective if the analysis is performed on-line. The feasibility of this new approach is demonstrated in the fast kinetic analysis of the carbon-sulfur coupling between 3-chlorophenylhydrazonopropane dinitrile and beta-mercaptoethanol. The theory agrees well with the results obtained from 31 repeated C-S coupling experiments.

Journal Article↗

Batch process monitoring using on-line MIR spectroscopy.

Many high quality products are produced in a batch wise manner. One of the characteristics of a batch process is the recipe driven nature. By repeating the recipe in an identical manner a desired end-product is obtained. However, in spite of repeating the recipe in an identical manner, process differences occur. These differences can be caused by a change of feed stock supplier or impurities in the process. Because of this, differences might occur in the end-product quality or unsafe process situations arise. Therefore, the need to monitor an industrial batch process exists. An industrial process is usually monitored by process measurements such as pressures and temperatures. Nowadays, due to technical developments, spectroscopy is more and more used for process monitoring. Spectroscopic measurements have the advantage of giving a direct chemical insight in the process. Multivariate statistical process control (MSPC) is a statistical way of monitoring the behaviour of a process. Combining spectroscopic measurements with MSPC will notice process perturbations or process deviations from normal operating conditions in a very simple manner. In the following an application is given of batch process monitoring. It is shown how a calibration model is developed and used with the principles of MSPC. Statistical control charts are developed and used to detect batches with a process upset.

Journal Article↗

Near-infrared spectroscopic monitoring of a series of industrial batch processes using a bilinear grey model.

A good process understanding is the foundation for process optimization, process monitoring, end-point detection, and estimation of the end-product quality. Performing good process measurements and the construction of process models will contribute to a better process understanding. To improve the process knowledge it is common to build process models. These models are often based on first principles such as kinetic rates or mass balances. These types of models are also known as hard or white models. White models are characterized by being generally applicable but often having only a reasonable fit to real process data. Other commonly used types of models are empirical or black-box models such as regression and neural nets. Black-box models are characterized by having a good data fit but they lack a chemically meaningful model interpretation. Alternative models are grey models, which are combinations of white models and black models. The aim of a grey model is to combine the advantages of both black-box models and white models. In a qualitative case study of monitoring industrial batches using near-infrared (NIR) spectroscopy, it is shown that grey models are a good tool for detecting batch-to-batch variations and an excellent tool for process diagnosis compared to common spectroscopic monitoring tools.

Chemical Industry↗