PubMed Health⌕ Search

Biomedical subjects

Age K Smilde

Publications and source records attributed to Age K Smilde.

13 recordsLinked to original sources

Centering, scaling, and transformations: improving the biological information content of metabolomics data.

BACKGROUND: Extracting relevant biological information from large data sets is a major challenge in functional genomics research. Different aspects of the data hamper their biological interpretation. For instance, 5000-fold differences in concentration for different metabolites are present in a metabolomics data set, while these differences are not proportional to the biological relevance of these metabolites. However, data analysis methods are not able to make this distinction. Data pretreatment methods can correct for aspects that hinder the biological interpretation of metabolomics data sets by emphasizing the biological information in the data set and thus improving their biological interpretability. RESULTS: Different data pretreatment methods, i.e. centering, autoscaling, pareto scaling, range scaling, vast scaling, log transformation, and power transformation, were tested on a real-life metabolomics data set. They were found to greatly affect the outcome of the data analysis and thus the rank of the, from a biological point of view, most important metabolites. Furthermore, the stability of the rank, the influence of technical errors on data analysis, and the preference of data analysis methods for selecting highly abundant metabolites were affected by the data pretreatment method used prior to data analysis. CONCLUSION: Different pretreatment methods emphasize different aspects of the data and each pretreatment method has its own merits and drawbacks. The choice for a pretreatment method depends on the biological question to be answered, the properties of the data set and the data analysis method selected. For the explorative analysis of the validation data set used in this study, autoscaling and range scaling performed better than the other pretreatment methods. That is, range scaling and autoscaling were able to remove the dependence of the rank of the metabolites on the average concentration and the magnitude of the fold changes and showed biologically sensible results after PCA (principal component analysis).In conclusion, selecting a proper data pretreatment method is an essential step in the analysis of metabolomics data and greatly affects the metabolites that are identified to be the most important.

Cluster Analysis↗

Large-scale human metabolomics studies: a strategy for data (pre-) processing and validation.

A large metabolomics study was performed on 600 plasma samples taken at four time points before and after a single intake of a high fat test meal by obese and lean subjects. All samples were analyzed by a liquid chromatography-mass spectrometry (LC-MS) lipidomic method for metabolic profiling. A pragmatic approach combining several well-established statistical methods was developed for processing this large data set in order to detect small differences in metabolic profiles in combination with a large biological variation. Such metabolomics studies require a careful analytical and statistical protocol. The strategy included data preprocessing, data analysis, and validation of statistical models. After several data preprocessing steps, partial least-squares discriminant analysis (PLS-DA) was used for finding biomarkers. To validate the found biomarkers statistically, the PLS-DA models were validated by means of a permutation test, biomarker models, and noninformative models. Univariate plots of potential biomarkers were used to obtain insight in up- or downregulation. The strategy proposed proved to be applicable for dealing with large-scale human metabolomics studies.

Chromatography, Liquid↗

Fusion of mass spectrometry-based metabolomics data.

A general method is presented for combining mass spectrometry-based metabolomics data. Such data are becoming more and more abundant, and proper tools for fusing these types of data sets are needed. Fusion of metabolomics data leads to a comprehensive view on the metabolome of an organism or biological system. The ideas presented draw upon established techniques in data analysis. Hence, they are also widely applicable to other types of X-omics data provided there is a proper pretreatment of the data. These issues are discussed using a real-life metabolomics data set from a microbial fermentation process.

Databases as Topic↗

ANOVA-simultaneous component analysis (ASCA): a new tool for analyzing designed metabolomics data.

MOTIVATION: Datasets resulting from metabolomics or metabolic profiling experiments are becoming increasingly complex. Such datasets may contain underlying factors, such as time (time-resolved or longitudinal measurements), doses or combinations thereof. Currently used biostatistics methods do not take the structure of such complex datasets into account. However, incorporating this structure into the data analysis is important for understanding the biological information in these datasets. RESULTS: We describe ASCA, a new method that can deal with complex multivariate datasets containing an underlying experimental design, such as metabolomics datasets. It is a direct generalization of analysis of variance (ANOVA) for univariate data to the multivariate case. The method allows for easy interpretation of the variation induced by the different factors of the design. The method is illustrated with a dataset from a metabolomics experiment with time and dose factors.

Algorithms↗

Tackling calibration problems of spectroscopic analysis in high-throughput experimentation.

High-throughput experimentation and screening methods are changing work flows and creating new possibilities in biochemistry, organometallic chemistry, and catalysis. However, many high-throughput systems rely on off-line chromatography methods that shift the bottleneck to the analysis stage. On-line or at-line spectroscopic analysis is an attractive alternative. It is fast, noninvasive, and nondestructive and requires no sample handling. The disadvantage is that spectroscopic calibration is time-consuming and complex. Ideally, the calibration model should give reliable predictions while keeping the number of calibration samples to a minimum. In this paper, we employ the net analyte signal approach to build a calibration model for Fourier transform near-infrared measurements, using a minimum number of calibration samples based on blank samples. This approach fits very well to high-throughput setups. With this approach, we can reduce the number of calibration samples to the number of chemical components in the system. Thus, the question is no longer how many but which type of calibration samples should one include in the model to obtain reliable predictions. Various calibration models are tested using Monte Carlo simulations, and the results are compared with experimental data for palladium-catalyzed Heck cross-coupling.

Journal Article↗

Profiling of liquid crystal displays with Raman spectroscopy: preprocessing of spectra.

Raman spectroscopy is applied for characterizing paintable displays. Few other options than Raman spectroscopy exist for doing so because of the liquid nature of functional materials. The challenge is to develop a method that can be used for estimating the composition of a single display cell on the basis of the collected three-dimensional Raman spectra. A classical least squares (CLS) model is used to model the measured spectra. It is shown that spectral preprocessing is a necessary and critical step for obtaining a good CLS model and reliable compositional profiles. Different kinds of preprocessing are explained. For each data set the type and amount of preprocessing may be different. This is shown using two data sets measured on essentially the same type of display cell, but under different experimental conditions. For model validation three criteria are introduced: mean sum of squares of residuals, percentage of unexplained information (PUN), and average residual curve. It is shown that the decision about the best combination of preprocessing techniques cannot be based only on overall error indicators (such as PUN). In addition, local residual analysis must be done and the feasibility of the extracted profiles should be taken into account.

Algorithms↗

Performance optimization of spectroscopic process analyzers.

To increase the power and the robustness of spectroscopic process analyzers, methods are needed that suppress the spectral variation that is not related to the property of interest in the process stream. An approach for the selection of a suitable method is presented. The approach uses the net analyte signal (NAS) to analyze the situation and to select methods to suppress the nonrelevant spectral variation. The empirically determined signal-to-noise of the NAS is used as a figure of merit. The advantages of the approach are (i). that the error of the reference method does not affect method selection and (ii). that only a few spectral measurements are needed. A diagnostic plot is proposed that guides the user in the evaluation of the particular suppression method. As an example, NIR spectroscopic monitoring of a mol-sieve separation process is used.

Journal Article↗

Analysis of longitudinal metabolomics data.

MOTIVATION: Metabolomics datasets are generally large and complex. Using principal component analysis (PCA), a simplified view of the variation in the data is obtained. The PCA model can be interpreted and the processes underlying the variation in the data can be analysed. In metabolomics, often a priori information is present about the data. Various forms of this information can be used in an unsupervised data analysis with weighted PCA (WPCA). A WPCA model will give a view on the data that is different from the view obtained using PCA, and it will add to the interpretation of the information in a metabolomics dataset. RESULTS: A method is presented to translate spectra of repeated measurements into weights describing the experimental error. These weights are used in the data analysis with WPCA. The WPCA model will give a view on the data where the non-uniform experimental error is accounted for. Therefore, the WPCA model will focus more on the natural variation in the data. AVAILABILITY: M-files for MATLAB for the algorithm used in this research are available at http://www-its.chem.uva.nl/research/pac/Software/pcaw.zip.

Algorithms↗

Quantitative analysis of target components by comprehensive two-dimensional gas chromatography.

Quantitative analysis using comprehensive two-dimensional (2D) gas chromatography (GC) is still rarely reported. This is largely due to a lack of suitable software. The objective of the present study is to generate quantitative results from a large GC x GC data set, consisting of 32 chromatograms. In this data set, six target components need to be quantified. We compare the results of conventional integration with those obtained using so-called "multiway analysis methods". With regard to accuracy and precision, integration performs slightly better than Parallel Factor (PARAFAC) analysis. In terms of speed and possibilities for automation, multiway methods in general are far superior to traditional integration.

Chromatography, Gas↗

Near-infrared spectroscopic monitoring of a series of industrial batch processes using a bilinear grey model.

A good process understanding is the foundation for process optimization, process monitoring, end-point detection, and estimation of the end-product quality. Performing good process measurements and the construction of process models will contribute to a better process understanding. To improve the process knowledge it is common to build process models. These models are often based on first principles such as kinetic rates or mass balances. These types of models are also known as hard or white models. White models are characterized by being generally applicable but often having only a reasonable fit to real process data. Other commonly used types of models are empirical or black-box models such as regression and neural nets. Black-box models are characterized by having a good data fit but they lack a chemically meaningful model interpretation. Alternative models are grey models, which are combinations of white models and black models. The aim of a grey model is to combine the advantages of both black-box models and white models. In a qualitative case study of monitoring industrial batches using near-infrared (NIR) spectroscopy, it is shown that grey models are a good tool for detecting batch-to-batch variations and an excellent tool for process diagnosis compared to common spectroscopic monitoring tools.

Chemical Industry↗

Selection of optimal process analyzers for plant-wide monitoring.

In this paper, the effect of process analyzer selection and positioning on plant-wide process monitoring is investigated. A fundamental problem in process analytical chemistry is the incomparability of different instrument characteristics. A fast but imprecise instrument is incomparable to a slow but precise instrument. Theory is developed to overcome this problem by using an abstract definition of a process analyzer. This definition allows us to put all instrument characteristics for a particular monitoring task on an equal footing. This results in a measurability factor M that expresses monitoring performance of any process measurement by combining instrument characteristics such as precision, sampling rate, grab size, response correlation, and delay time. Both the choice of location and the performance characteristics of different process analyzers can be evaluated using the measurability factor. The unifying nature of the measurability factor allows for a rational decision between completely different process analyzers and locations (Smilde et al., in this issue). The theory is illustrated and validated with an experiment. A tubular reactor for free-radical bulk polymerization of styrene is monitored by in-line short-wave near-infrared spectroscopy at different positions. Alternatively, product samples are collected for at-line near-infrared analysis. Both analyzers measure styrene monomer concentration. The analysis results are used to predict conversion as well as number and weight average molecular mass of the polystyrene reactor product. The theoretical measurability factors for this case study correspond well with the experimental findings.

Journal Article↗

Direct sampling tandem mass spectrometry (MS/MS) and multiway calibration for isomer quantitation.

Direct sampling tandem mass spectrometry (MS/MS) was used for the quantitation of mixtures of the isomers 2-, 3- and 4-ethyl pyridine. The similarity between the analytes and the second-order nature of MS/MS data require the use of multivariate calibration techniques capable of handling multiway data. Multilinear PLS (N-PLS) was applied here, as well as the alternative technique of unfolding the data and using standard two-way PLS. Particular attention was paid to the optimal type of spectral preprocessing. Due to the presence of heteroscedastic noise the logarithmic transform of the spectra prior to calibration gives the best results. Predictions errors of the order of 10-15% were obtained, which compare well with other results found in the literature.

Calibration↗