PubMed Health⌕ Search

Biomedical subjects

Harald Martens

Publications and source records attributed to Harald Martens.

8 recordsLinked to original sources

Parameter estimation in biochemical systems models with alternating regression.

BACKGROUND: The estimation of parameter values continues to be the bottleneck of the computational analysis of biological systems. It is therefore necessary to develop improved methods that are effective, fast, and scalable. RESULTS: We show here that alternating regression (AR), applied to S-system models and combined with methods for decoupling systems of differential equations, provides a fast new tool for identifying parameter values from time series data. The key feature of AR is that it dissects the nonlinear inverse problem of estimating parameter values into iterative steps of linear regression. We show with several artificial examples that the method works well in many cases. In cases of no convergence, it is feasible to dedicate some computational effort to identifying suitable start values and search settings, because the method is fast in comparison to conventional methods that the search for suitable initial values is easily recouped. Because parameter estimation and the identification of system structure are closely related in S-system modeling, the AR method is beneficial for the latter as well. Specifically, we show with an example from the literature that AR is three to five orders of magnitudes faster than direct structure identifications in systems of nonlinear differential equations. CONCLUSION: Alternating regression provides a strategy for the estimation of parameter values and the identification of structure and regulation in S-systems that is genuinely different from all existing methods. Alternating regression is usually very fast, but its convergence patterns are complex and will require further investigation. In cases where convergence is an issue, the enormous speed of the method renders it feasible to select several initial guesses and search settings as an effective countermeasure.

Computational Biology↗

Challenges related to analysis of protein spot volumes from two-dimensional gel electrophoresis as revealed by replicate gels.

Assumptions that need to be considered prior to statistical analysis of protein spot volumes from two-dimensional gel electrophoresis (2-DE) data are studied using replicate gels of the same sample. The most important observation is that the data tables of protein spot volumes from 2-DE images contain a large number of missing values, which are not consistent with the presence or absence of the proteins. This implies both loss of information and problems for the subsequent statistical analysis. Challenges with 2-DE protein spot volumes are viewed in light of multiple gel comparisons and multivariate data analysis.

Data Interpretation, Statistical↗

Correcting attenuated total reflection-Fourier transform infrared spectra for water vapor and carbon dioxide.

Fourier transform infrared (FT-IR) spectroscopy is a valuable technique for characterization of biological samples, providing a detailed fingerprint of the major chemical constituents. However, water vapor and CO(2) in the beam path often cause interferences in the spectra, which can hamper the data analysis and interpretation of results. In this paper we present a new method for removal of the spectral contributions due to atmospheric water and CO(2) from attenuated total reflection (ATR)-FT-IR spectra. In the IR spectrum, four separate wavenumber regions were defined, each containing an absorption band from either water vapor or CO(2). From two calibration data sets, gas model spectra were estimated in each of the four spectral regions, and these model spectra were applied for correction of gas absorptions in two independent test sets (spectra of aqueous solutions and a yeast biofilm (C. albicans) growing on an ATR crystal, respectively). The amounts of the atmospheric gases as expressed by the model spectra were estimated by regression, using second-derivative transformed spectra, and the estimated gas spectra could subsequently be subtracted from the sample spectra. For spectra of the growing yeast biofilm, the gas correction revealed otherwise hidden variations of relevance for modeling the growth dynamics. As the presented method improved the interpretation of the principle component analysis (PCA) models, it has proven to be a valuable tool for filtering atmospheric variation in ATR-FT-IR spectra.

Artifacts↗

Analysis of covariance patterns in gene expression data and FT-IR spectra.

The aim of this study was to detect and interpret correlation patterns in several large data matrices from the same biological system using Partial Least Squares Regression (PLSR) in order to get information on the system under investigation. To do this, DNA microarray data and Fourier Transform Infrared (FT-IR) spectra from a designed study where Campylobacter jejuni was exposed to environmental stress conditions, were used. The experimental design included variation in atmospheric conditions, temperature and time. PLSR was first used to analyse each of the two data types separately in order to explore the effect of the experimental parameters on the data. The results showed that both the gene expression and FT-IR spectra were affected by the variations in atmosphere, temperature and time, but that the effect was different for the two types of data. When the DNA microarray data and FT-IR spectra were linked together by PLSR, covariation due to temperature was seen. Both specific genes and ranges in the FT-IR spectra that were connected to the variation in temperature were detected. Some of these are possibly connected to properties of the cell wall of the bacteria. The results in this study show the potential of PLSR for investigation of covariance structures in biological data. By doing this, valuable information about the biological system can be detected and interpreted. It was also shown that the use of FT-IR spectroscopy provided important information about the stress responses in the bacteria, information that was not detected from the DNA microarray data.

Campylobacter jejuni↗

Evaluation of nonstarch polysaccharides and oligosaccharide content of different soybean varieties (Glycine max) by near-infrared spectroscopy and proteomics.

A total of 832 samples of soybeans were screened by near-infrared (NIR) reflectance spectroscopy, to identify soybean samples with a lower content of oligosaccharides and nonstarch polysaccharides (NSP). Of these, 38 samples were identified on the basis of variation in protein content and agronomic value and submitted to high-resolution NIR spectroscopy. On the basis of the NIR data, 12 samples were further selected for chromatographic characterization of carbohydrate composition (mono-, di-, and oligosaccharides and NSP). Their soluble proteins were separated by two-dimensional gel electrophoresis (2DE). Using partial least-squares regression (PLSR), it was possible to predict the content of total NSP from the high-resolution NIR spectra, suggesting that NIR is a suitable and rapid nondestructive method to determine carbohydrate composition in soybeans. The 2DE analyses showed varying intensities of several proteins, including the glycinin G1 precursor. PLSR analysis showed a negative correlation between this protein and insoluble NSP and total uronic acid (UA).

Oligosaccharides↗

Near-infrared spectra of Penicillium camemberti strains separated by extended multiplicative signal correction improved prediction of physical and chemical variations.

Different methods for spectral preprocessing were compared in relation to the ability to distinguish between fungal isolates and growth stages for Penicillium camemberti grown on cheese substrate. The best classification results were obtained by temperatureand wavelength-extended multivariate signal correction (TWEMSC) preprocessing, whereby three patterns of variation in nearinfrared (NIR) log(1/R) spectra of fungal colonies could be separated mathematically: (1) physical light scattering and its wavelength dependency, (2) differences in light absorption of water due to varying sample temperature, etc., and (3) differences in light absorption between different fungal isolates. With this preprocessing, discriminant partial least squares (PLS) regression yielded 100% correct classification of three isolates, both within the cross-validated calibration set and in two independent test sets of samples.

Algorithms↗

Analysis of genetic marker-phenotype relationships by jack-knifed partial least squares regression (PLSR).

The utility of a relatively new multivariate method, bi-linear modelling by cross-validated partial least squares regression (PLSR), was investigated in the analysis of QTL. The distinguishing feature of PLSR is to reveal reliable covariance structures in data of different types with regard to the same set objects. Two matrices X (here: genetic markers) and Y (here: phenotypes) are interactively decomposed into latent variables (PLS components, or PCs) in a way which facilitates statistically reliable and graphically interpretable model building. Natural collinearities between input variables are utilized actively to stabilise the modelling, instead of being treated as a statistical problem. The importance of cross-validation/jack-knifing as an intuitively appealing way to avoid overfitting, is emphasized. Two datasets from chromosomal mapping studies of different complexity were chosen for illustration (QTL for tomato yield and for oat heading date). Results from PLSR analysis were compared to published results and to results using the package PLABQTL in these data sets. In all cases PLSR gave at least similar explained validation variances as the reported studies. An attractive feature is that PLSR allows the analysis of several traits/replicates in one analysis, and the direct visual identification of individuals with desirable marker genotypes. It is suggested that PLSR may be useful in structural and functional genomics and in marker assisted selection, particularly in cases with limited number of objects.

Crosses, Genetic↗

Light scattering and light absorbance separated by extended multiplicative signal correction. application to near-infrared transmission analysis of powder mixtures.

The extended multiplicative signal correction (EMSC) preprocessing method allows a separation of physical light-scattering effects from chemical (vibrational) light absorbance effects in spectra from, for example, powders or turbid solutions. It is here applied to diffuse near infrared transmission (NIT) spectra of mixtures of wheat gluten (protein) and starch (carbohydrate) powders, linearized by conventional log(1/T). Without any correction for uncontrolled light scattering variation between the powder samples, these absorbance spectra could give reasonable predictions of the analyte (gluten), but only when using multivariate calibration with a much more complex model than expected. Standard MSC preprocessing did not work for these data at all; it removed too much analyte information. However, the EMSC preprocessing yielded powder spectra that obeyed Beer's Law more or less as if they had been obtained from transparent liquid solutions, apparently by isolating the chemical light absorption from additive, multiplicative, and wavelength-dependent effects of uncontrolled light-scattering variations. The model-based EMSC and its converse, the extended inverted signal correction (EISC), gave rather complete descriptions of the diffuse absorbance spectra and virtually indistinguishable performance in the calibration set and the test set of samples.

Journal Article↗