PubMed Health⌕ Search

Biomedical subjects

Huub C J Hoefsloot

Publications and source records attributed to Huub C J Hoefsloot.

6 recordsLinked to original sources

Centering, scaling, and transformations: improving the biological information content of metabolomics data.

BACKGROUND: Extracting relevant biological information from large data sets is a major challenge in functional genomics research. Different aspects of the data hamper their biological interpretation. For instance, 5000-fold differences in concentration for different metabolites are present in a metabolomics data set, while these differences are not proportional to the biological relevance of these metabolites. However, data analysis methods are not able to make this distinction. Data pretreatment methods can correct for aspects that hinder the biological interpretation of metabolomics data sets by emphasizing the biological information in the data set and thus improving their biological interpretability. RESULTS: Different data pretreatment methods, i.e. centering, autoscaling, pareto scaling, range scaling, vast scaling, log transformation, and power transformation, were tested on a real-life metabolomics data set. They were found to greatly affect the outcome of the data analysis and thus the rank of the, from a biological point of view, most important metabolites. Furthermore, the stability of the rank, the influence of technical errors on data analysis, and the preference of data analysis methods for selecting highly abundant metabolites were affected by the data pretreatment method used prior to data analysis. CONCLUSION: Different pretreatment methods emphasize different aspects of the data and each pretreatment method has its own merits and drawbacks. The choice for a pretreatment method depends on the biological question to be answered, the properties of the data set and the data analysis method selected. For the explorative analysis of the validation data set used in this study, autoscaling and range scaling performed better than the other pretreatment methods. That is, range scaling and autoscaling were able to remove the dependence of the rank of the metabolites on the average concentration and the magnitude of the fold changes and showed biologically sensible results after PCA (principal component analysis).In conclusion, selecting a proper data pretreatment method is an essential step in the analysis of metabolomics data and greatly affects the metabolites that are identified to be the most important.

Cluster Analysis↗

ANOVA-simultaneous component analysis (ASCA): a new tool for analyzing designed metabolomics data.

MOTIVATION: Datasets resulting from metabolomics or metabolic profiling experiments are becoming increasingly complex. Such datasets may contain underlying factors, such as time (time-resolved or longitudinal measurements), doses or combinations thereof. Currently used biostatistics methods do not take the structure of such complex datasets into account. However, incorporating this structure into the data analysis is important for understanding the biological information in these datasets. RESULTS: We describe ASCA, a new method that can deal with complex multivariate datasets containing an underlying experimental design, such as metabolomics datasets. It is a direct generalization of analysis of variance (ANOVA) for univariate data to the multivariate case. The method allows for easy interpretation of the variation induced by the different factors of the design. The method is illustrated with a dataset from a metabolomics experiment with time and dose factors.

Algorithms↗

Profiling of liquid crystal displays with Raman spectroscopy: preprocessing of spectra.

Raman spectroscopy is applied for characterizing paintable displays. Few other options than Raman spectroscopy exist for doing so because of the liquid nature of functional materials. The challenge is to develop a method that can be used for estimating the composition of a single display cell on the basis of the collected three-dimensional Raman spectra. A classical least squares (CLS) model is used to model the measured spectra. It is shown that spectral preprocessing is a necessary and critical step for obtaining a good CLS model and reliable compositional profiles. Different kinds of preprocessing are explained. For each data set the type and amount of preprocessing may be different. This is shown using two data sets measured on essentially the same type of display cell, but under different experimental conditions. For model validation three criteria are introduced: mean sum of squares of residuals, percentage of unexplained information (PUN), and average residual curve. It is shown that the decision about the best combination of preprocessing techniques cannot be based only on overall error indicators (such as PUN). In addition, local residual analysis must be done and the feasibility of the extracted profiles should be taken into account.

Algorithms↗

Analysis of longitudinal metabolomics data.

MOTIVATION: Metabolomics datasets are generally large and complex. Using principal component analysis (PCA), a simplified view of the variation in the data is obtained. The PCA model can be interpreted and the processes underlying the variation in the data can be analysed. In metabolomics, often a priori information is present about the data. Various forms of this information can be used in an unsupervised data analysis with weighted PCA (WPCA). A WPCA model will give a view on the data that is different from the view obtained using PCA, and it will add to the interpretation of the information in a metabolomics dataset. RESULTS: A method is presented to translate spectra of repeated measurements into weights describing the experimental error. These weights are used in the data analysis with WPCA. The WPCA model will give a view on the data where the non-uniform experimental error is accounted for. Therefore, the WPCA model will focus more on the natural variation in the data. AVAILABILITY: M-files for MATLAB for the algorithm used in this research are available at http://www-its.chem.uva.nl/research/pac/Software/pcaw.zip.

Algorithms↗

Selection of optimal process analyzers for plant-wide monitoring.

In this paper, the effect of process analyzer selection and positioning on plant-wide process monitoring is investigated. A fundamental problem in process analytical chemistry is the incomparability of different instrument characteristics. A fast but imprecise instrument is incomparable to a slow but precise instrument. Theory is developed to overcome this problem by using an abstract definition of a process analyzer. This definition allows us to put all instrument characteristics for a particular monitoring task on an equal footing. This results in a measurability factor M that expresses monitoring performance of any process measurement by combining instrument characteristics such as precision, sampling rate, grab size, response correlation, and delay time. Both the choice of location and the performance characteristics of different process analyzers can be evaluated using the measurability factor. The unifying nature of the measurability factor allows for a rational decision between completely different process analyzers and locations (Smilde et al., in this issue). The theory is illustrated and validated with an experiment. A tubular reactor for free-radical bulk polymerization of styrene is monitored by in-line short-wave near-infrared spectroscopy at different positions. Alternatively, product samples are collected for at-line near-infrared analysis. Both analyzers measure styrene monomer concentration. The analysis results are used to predict conversion as well as number and weight average molecular mass of the polystyrene reactor product. The theoretical measurability factors for this case study correspond well with the experimental findings.

Journal Article↗