PubMed Health⌕ Search

Biomedical subjects

Johan Trygg

Publications and source records attributed to Johan Trygg.

12 recordsLinked to original sources

MASQOT-GUI: spot quality assessment for the two-channel microarray platform.

UNLABELLED: MASQOT-GUI provides an open-source, platform-independent software pipeline for two-channel microarray spot quality control. This includes gridding, segmentation, quantification, quality assessment and data visualization. It hosts a set of independent applications, with interactions between the tools as well as import and export support for external software. The implementation of automated multivariate quality control assessment, which is a unique feature of MASQOT-GUI, is based on the previously documented and evaluated MASQOT methodology. Further abilities of the application are outlined and illustrated. AVAILABILITY: MASQOT-GUI is Java-based and licensed under the GNU LGPL. Source code and installation files are available for download at http://masqot-gui.sourceforge.net/

Computer Graphics↗

Predictive metabolite profiling applying hierarchical multivariate curve resolution to GC-MS data--a potential tool for multi-parametric diagnosis.

A method for predictive metabolite profiling based on resolution of GC-MS data followed by multivariate data analysis is presented and applied to three different biofluid data sets (rat urine, aspen leaf extracts, and human blood plasma). Hierarchical multivariate curve resolution (H-MCR) was used to simultaneously resolve the GC-MS data into pure profiles, describing the relative metabolite concentrations between samples, for multivariate analysis. Here, we present an extension of the H-MCR method allowing treatment of independent samples according to processing parameters estimated from a set of training samples. Predictions or inclusion of the new samples, based on their metabolite profiles, into an existing model could then be carried out, which is a requirement for a working application within, e.g., clinical diagnosis. Apart from allowing treatment and prediction of independent samples the proposed method also reduces the time for the curve resolution process since only a subset of representative samples have to be processed while the remaining samples can be treated according to the obtained processing parameters. The time required for resolving the 30 training samples in the rat urine example was approximately 13 h, while the treatment of the 30 test samples according to the training parameters required only approximately 30 s per sample (approximately 15 min in total). In addition, the presented results show that the suggested approach works for describing metabolic changes in different biofluids, indicating that this is a general approach for high-throughput predictive metabolite profiling, which could have important applications in areas such as plant functional genomics, drug toxicity, treatment efficacy and early disease diagnosis.

Animals↗

Statistically integrated metabonomic-proteomic studies on a human prostate cancer xenograft model in mice.

A novel statistically integrated proteometabonomic method has been developed and applied to a human tumor xenograft mouse model of prostate cancer. Parallel 2D-DIGE proteomic and 1H NMR metabolic profile data were collected on blood plasma from mice implanted with a prostate cancer (PC-3) xenograft and from matched control animals. To interpret the xenograft-induced differences in plasma profiles, multivariate statistical algorithms including orthogonal projection to latent structure (OPLS) were applied to generate models characterizing the disease profile. Two approaches to integrating metabonomic data matrices are presented based on OPLS algorithms to provide a framework for generating models relating to the specific and common sources of variation in the metabolite concentrations and protein abundances that can be directly related to the disease model. Multiple correlations between metabolites and proteins were found, including associations between serotransferrin precursor and both tyrosine and 3-D-hydroxybutyrate. Additionally, a correlation between decreased concentration of tyrosine and increased presence of gelsolin was also observed. This approach can provide enhanced recovery of combination candidate biomarkers across multi-omic platforms, thus, enhancing understanding of in vivo model systems studied by multiple omic technologies.

Animals↗

Consensus by democracy. Using meta-analyses of microarray and genomic data to model the cold acclimation signaling pathway in Arabidopsis.

The whole-genome response of Arabidopsis (Arabidopsis thaliana) exposed to different types and durations of abiotic stress has now been described by a wealth of publicly available microarray data. When combined with studies of how gene expression is affected in mutant and transgenic Arabidopsis with altered ability to transduce the low temperature signal, these data can be used to test the interactions between various low temperature-associated transcription factors and their regulons. We quantized a collection of Affymetrix microarray data so that each gene in a particular regulon could vote on whether a cis-element found in its promoter conferred induction (+1), repression (-1), or no transcriptional change (0) during cold stress. By statistically comparing these election results with the voting behavior of all genes on the same gene chip, we verified the bioactivity of novel cis-elements and defined whether they were inductive or repressive. Using in silico mutagenesis we identified functional binding consensus variants for the transcription factors studied. Our results suggest that the previously identified ICEr1 (induction of CBF expression region 1) consensus does not correlate with cold gene induction, while the ICEr3/ICEr4 consensuses identified using our algorithms are present in regulons of genes that were induced coordinate with observed ICE1 transcript accumulation and temporally preceding genes containing the dehydration response element. Statistical analysis of overlap and cis-element enrichment in the ICE1, CBF2, ZAT12, HOS9, and PHYA regulons enabled us to construct a regulatory network supported by multiple lines of evidence that can be used for future hypothesis testing.

Acclimatization↗

Extraction and GC/MS analysis of the human blood plasma metabolome.

Analysis of the entire set of low molecular weight compounds (LMC), the metabolome, could provide deeper insights into mechanisms of disease and novel markers for diagnosis. In the investigation, we developed an extraction and derivatization protocol, using experimental design theory (design of experiment), for analyzing the human blood plasma metabolome by GC/MS. The protocol was optimized by evaluating the data for more than 500 resolved peaks using multivariate statistical tools including principal component analysis and partial least-squares projections to latent structures (PLS). The performance of five organic solvents (methanol, ethanol, acetonitrile, acetone, chloroform), singly and in combination, was investigated to optimize the LMC extraction. PLS analysis demonstrated that methanol extraction was particularly efficient and highly reproducible. The extraction and derivatization conditions were also optimized. Quantitative data for 32 endogenous compounds showed good precision and linearity. In addition, the determined amounts of eight selected compounds agreed well with analyses by independent methods in accredited laboratories, and most of the compounds could be detected at absolute levels of approximately 0.1 pmol injected, corresponding to plasma concentrations between 0.1 and 1 microM. The results suggest that the method could be usefully integrated into metabolomic studies for various purposes, e.g., for identifying biological markers related to diseases.

Acetone↗

A statistical resampling method to calculate biomagnification factors exemplified with organochlorine data from herring (Clupea harengus) muscle and guillemot (Uria aalge) egg from the Baltic sea.

A novel method for calculating biomagnification factors is presented and demonstrated using contaminant concentration data from the Swedish national monitoring program regarding organochlorine contaminants (OCs) in herring (Clupea harengus) muscle and guillemot (Uria aalge) egg, sampled from 1996 to 1999 from the Baltic Sea. With this randomly sampled ratios (RSR) method, biomagnification factors (BMF(RSR)) were generated and denoted with standard deviation (SD) as a measure of the variation. The BMFRsR were calculated by randomly selecting one guillemot egg out of a total of 29 and one herring out of a total of 74, and the ratio was determined between the concentration of a given OC in that egg and the concentration of the same OC in that herring. With the resampling technique, this was performed 50 000 times for any given OC, and from this new distribution of ratios, BMF(RSR) for each OC were calculated and given as geometric mean (GM) with GM standard deviation (GMSD) range, arithmetic mean (AM) with AMSD range, and minimum (BMF(MIN)) as well as maximum (BMF(MAX)) biomagnification factors. The 14 analyzed OCs were p,p'DDT and its metabolites p,p'DDE and p,p'DDD, polychlorinated biphenyls (PCB congeners: CB28, CB52, CB101, CB118, CB138, CB153, and CB180), hexachlorocyclohexane isomers (alpha-, beta-, and gammaHCH), and hexachlorobenzene (HCB). Multivariate data analysis (MVDA) methods, including principal components analysis (PCA), partial least squares regression (PLS), and PLS discriminant analyses (PLS-DA), were first used to extract information from the complex biological and chemical data generated from each individual animal. MVDA were used to model similarities/dissimilarities regarding species (PCA, PLS-DA), sample years (PLS), and sample location (PLS-DA) to give a deeper understanding of the data that the BMF modeling was based upon. Contaminants that biomagnify, that had BMF(RSR) significantly higher than one, were p,p'DDE, CB118, HCB, CB138, CB180, CB153, ,betaHCH, and CB28. The contaminants that did not biomagnifywere p,p'DDT, p,p'DDD, alphaHCH, CB101, and CB52. Eventual biomagnification for gammaHCH could not be determined. The BMF(RSR) for OCs present in herring muscle and guillemot egg showed a broad span with large variations for each contaminant. To be able to make reliable calculations of BMFs for different contaminants, we emphasize the importance of using data based upon large numbers of, as well as well-defined, individuals.

Animals↗

MASQOT: a method for cDNA microarray spot quality control.

BACKGROUND: cDNA microarray technology has emerged as a major player in the parallel detection of biomolecules, but still suffers from fundamental technical problems. Identifying and removing unreliable data is crucial to prevent the risk of receiving illusive analysis results. Visual assessment of spot quality is still a common procedure, despite the time-consuming work of manually inspecting spots in the range of hundreds of thousands or more. RESULTS: A novel methodology for cDNA microarray spot quality control is outlined. Multivariate discriminant analysis was used to assess spot quality based on existing and novel descriptors. The presented methodology displays high reproducibility and was found superior in identifying unreliable data compared to other evaluated methodologies. CONCLUSION: The proposed methodology for cDNA microarray spot quality control generates non-discrete values of spot quality which can be utilized as weights in subsequent analysis procedures as well as to discard spots of undesired quality using the suggested threshold values. The MASQOT approach provides a consistent assessment of spot quality and can be considered an alternative to the labor-intensive manual quality assessment process.

Data Interpretation, Statistical↗

High-throughput data analysis for detecting and identifying differences between samples in GC/MS-based metabolomic analyses.

In metabolomics, the objective is to identify differences in metabolite profiles between samples. A widely used tool in metabolomics investigations is gas chromatography-mass spectrometry (GC/MS). More than 400 compounds can be detected in a single analysis, if overlapping GC/MS peaks are deconvoluted. However, the deconvolution process is time-consuming and difficult to automate, and additional processing is needed in order to compare samples. Therefore, there is a need to improve and automate the data processing strategy for data generated in GC/MS-based metabolomics; if not, the processing step will be a major bottleneck for high-throughput analyses. Here we describe a new semiautomated strategy using a hierarchical multivariate curve resolution approach that processes all samples simultaneously. The presented strategy generates (after appropriate treatment, e.g., multivariate analysis) tables of all the detected metabolites that differ in relative concentrations between samples. The processing of 70 samples took similar time to that of the GC/TOFMS analyses of the samples. The strategy has been validated using two different sets of samples: a complex mixture of standard compounds and Arabidopsis samples.

Arabidopsis↗

Extraction, interpretation and validation of information for comparing samples in metabolic LC/MS data sets.

LC/MS is an analytical technique that, due to its high sensitivity, has become increasingly popular for the generation of metabolic signatures in biological samples and for the building of metabolic data bases. However, to be able to create robust and interpretable (transparent) multivariate models for the comparison of many samples, the data must fulfil certain specific criteria: (i) that each sample is characterized by the same number of variables, (ii) that each of these variables is represented across all observations, and (iii) that a variable in one sample has the same biological meaning or represents the same metabolite in all other samples. In addition, the obtained models must have the ability to make predictions of, e.g. related and independent samples characterized accordingly to the model samples. This method involves the construction of a representative data set, including automatic peak detection, alignment, setting of retention time windows, summing in the chromatographic dimension and data compression by means of alternating regression, where the relevant metabolic variation is retained for further modelling using multivariate analysis. This approach has the advantage of allowing the comparison of large numbers of samples based on their LC/MS metabolic profiles, but also of creating a means for the interpretation of the investigated biological system. This includes finding relevant systematic patterns among samples, identifying influential variables, verifying the findings in the raw data, and finally using the models for predictions. The presented strategy was here applied to a population study using urine samples from two cohorts, Shanxi (People's Republic of China) and Honolulu (USA). The results showed that the evaluation of the extracted information data using partial least square discriminant analysis (PLS-DA) provided a robust, predictive and transparent model for the metabolic differences between the two populations. The presented findings suggest that this is a general approach for data handling, analysis, and evaluation of large metabolic LC/MS data sets.

Chromatography, Liquid↗

Statistical total correlation spectroscopy: an exploratory approach for latent biomarker identification from metabolic 1H NMR data sets.

We describe here the implementation of the statistical total correlation spectroscopy (STOCSY) analysis method for aiding the identification of potential biomarker molecules in metabonomic studies based on NMR spectroscopic data. STOCSY takes advantage of the multicollinearity of the intensity variables in a set of spectra (in this case 1H NMR spectra) to generate a pseudo-two-dimensional NMR spectrum that displays the correlation among the intensities of the various peaks across the whole sample. This method is not limited to the usual connectivities that are deducible from more standard two-dimensional NMR spectroscopic methods, such as TOCSY. Moreover, two or more molecules involved in the same pathway can also present high intermolecular correlations because of biological covariance or can even be anticorrelated. This combination of STOCSY with supervised pattern recognition and particularly orthogonal projection on latent structure-discriminant analysis (O-PLS-DA) offers a new powerful framework for analysis of metabonomic data. In a first step O-PLS-DA extracts the part of NMR spectra related to discrimination. This information is then cross-combined with the STOCSY results to help identify the molecules responsible for the metabolic variation. To illustrate the applicability of the method, it has been applied to 1H NMR spectra of urine from a metabonomic study of a model of insulin resistance based on the administration of a carbohydrate diet to three different mice strains (C57BL/6Oxjr, BALB/cOxjr, and 129S6/SvEvOxjr) in which a series of metabolites of biological importance can be conclusively assigned and identified by use of the STOCSY approach.

Amines↗

Evaluation of the orthogonal projection on latent structure model limitations caused by chemical shift variability and improved visualization of biomarker changes in 1H NMR spectroscopic metabonomic studies.

In general, applications of metabonomics using biofluid NMR spectroscopic analysis for probing abnormal biochemical profiles in disease or due to toxicity have all relied on the use of chemometric techniques for sample classification. However, the well-known variability of some chemical shifts in 1H NMR spectra of biofluids due to environmental differences such as pH variation, when coupled with the large number of variables in such spectra, has led to the situation where it is necessary to reduce the size of the spectra or to attempt to align the shifting peaks, to get more robust and interpretable chemometric models. Here, a new approach that avoids this problem is demonstrated and shows that, moreover, inclusion of variable peak position data can be beneficial and can lead to useful biochemical information. The interpretation of chemometric models using combined back-scaled loading plots and variable weights demonstrates that this peak position variation can be handled successfully and also often provides additional information on the physicochemical variations in metabonomic data sets.

3-Hydroxybutyric Acid↗

Using chemometrics for navigating in the large data sets of genomics, proteomics, and metabonomics (gpm).

This article describes the applicability of multivariate projection techniques, such as principal-component analysis (PCA) and partial least-squares (PLS) projections to latent structures, to the large-volume high-density data structures obtained within genomics, proteomics, and metabonomics. PCA and PLS, and their extensions, derive their usefulness from their ability to analyze data with many, noisy, collinear, and even incomplete variables in both X and Y. Three examples are used as illustrations: the first example is a genomics data set and involves modeling of microarray data of cell cycle-regulated genes in the microorganism Saccharomyces cerevisiae. The second example contains NMR-metabonomics data, measured on urine samples of male rats treated with either of the drugs chloroquine or amiodarone. The third and last data set describes sequence-function classification studies in a set of G-protein-coupled receptors using hierarchical PCA.

Animals↗