PubMed Health⌕ Search

Biomedical subjects

Bob W Wright

Publications and source records attributed to Bob W Wright.

6 recordsLinked to original sources

Unsupervised parameter optimization for automated retention time alignment of severely shifted gas chromatographic data using the piecewise alignment algorithm.

Simulated chromatographic separations were used to study the performance of piecewise retention time alignment and to demonstrate automated unsupervised (without a training set) parameter optimization. The average correlation coefficient between the target chromatogram and all remaining chromatograms in the data set was used to optimize the alignment parameters. This approach frees the user from providing class information and makes the alignment algorithm applicable to classifying completely unknown data sets. The average peak in the raw simulated data set was shifted up to two peak-widths-at-base (average relative shift=2.0) and after alignment the average relative shift was improved to 0.3. Piecewise alignment was applied to severely shifted GC separations of gasolines and reformate distillation fraction samples. The average relative shifts in the raw gasolines and reformates data were 4.7 and 1.5, respectively, but after alignment improved to 0.5 and 0.4, respectively. The effect of piecewise alignment on peak heights and peak areas is also reported. The average relative difference in peak height was -0.20%. The average absolute relative difference in area was 0.15%.

Algorithms↗

Classification of high-speed gas chromatography-mass spectrometry data by principal component analysis coupled with piecewise alignment and feature selection.

A useful methodology is introduced for the analysis of data obtained via gas chromatography with mass spectrometry (GC-MS) utilizing a complete mass spectrum at each retention time interval in which a mass spectrum was collected. Principal component analysis (PCA) with preprocessing by both piecewise retention time alignment and analysis of variance (ANOVA) feature selection is applied to all mass channels collected. The methodology involves concatenating all concurrently measured individual m/z chromatograms from m/z 20 to 120 for each GC-MS separation into a row vector. All of the sample row vectors are incorporated into a matrix where each row is a sample vector. This matrix is piecewise aligned and reduced by ANOVA feature selection. Application of the preprocessing steps (retention time alignment and feature selection) to all mass channels collected during the chromatographic separation allows considerably more selective chemical information to be incorporated in the PCA classification, and is the primary novelty of the report. This methodology is objective and requires no knowledge of the specific analytes of interest, as in selective ion monitoring (SIM), and does not restrict the mass spectral data used, as in both SIM and total ion current (TIC) methods. Significantly, the methodology allows for the classification of data with low resolution in the chromatographic dimension because of the added selectivity from the complete mass spectral dimension. This allows for the successful classification of data over significantly decreased chromatographic separation times, since high-speed separations can be employed. The methodology is demonstrated through the analysis of a set of four differing gasoline samples that serve as model complex samples. For comparison, the gasoline samples are analyzed by GC-MS over both 10-min and 10-s separation times. The successfully classified 10-min GC-MS TIC data served as the benchmark analysis to compare to the 10-s data. When only alignment and feature selection was applied to the 10-s gasoline separations using GC-MS TIC data, PCA failed. PCA was successful for 10-s gasoline separations when the methodology was applied with all the m/z information. With ANOVA feature selection, chromatographic regions with Fisher ratios greater than 1500 were retained in a new matrix and subjected to PCA yielding successful classification for the 10-s separations.

Analysis of Variance↗

Fisher ratio method applied to third-order separation data to identify significant chemical components of metabolite extracts.

This report is about applying a Fisher ratio method to entire four dimensional (4D) data sets from third-order instrumentation data. The Fisher ratio method uses a novel indexing scheme to discover the unknown chemical differences among known classes of complex samples. This is the first report of a Fisher ratio analysis procedure applied to entire 4D data sets of third-order separation data, which, in this case, is comprehensive two-dimensional gas chromatography coupled with time-of-flight mass spectrometry analyses of metabolite extracts using all of the collected mass channels. Current analysis methods for third-order separation data use only user-defined subsets of the 4D data set. First, in a validation study, the Fisher ratio method was demonstrated to objectively evaluate and determine the chemical differences between three controlled urine samples that differed by known spiked chemical components. It was determined that, out of more than 600 recognizable chemical components in a single sample, the six spiked components, along with only two other matrix components, differed most significantly in concentration among the control samples. In a second study, the Fisher ratio method was used in a novel application to discover the unknown chemical differences between urine metabolite samples from pregnant women and nonpregnant women. A brief list of the top 11 components that were most significantly different in concentration between the pregnant and nonpregnant samples was generated. Because the Fisher ratio calculation statistically differentiates regions of the chromatogram with large class-to-class variations from regions containing large within-class variations, the Fisher ratio method should generally be robust against biological diversity in a sample population. Indeed, application of principal component analysis in this second study failed due to biological diversity of the samples.

Algorithms↗

A comprehensive two-dimensional retention time alignment algorithm to enhance chemometric analysis of comprehensive two-dimensional separation data.

A comprehensive two-dimensional (2D) retention time alignment algorithm was developed using a novel indexing scheme. The algorithm is termed comprehensive because it functions to correct the entire chromatogram in both dimensions and it preserves the separation information in both dimensions. Although the algorithm is demonstrated by correcting comprehensive two-dimensional gas chromatography (GC x GC) data, the algorithm is designed to correct shifting in all forms of 2D separations, such as LC x LC, LC x CE, CE x CE, and LC x GC. This 2D alignment algorithm was applied to three different data sets composed of replicate GC x GC separations of (1) three 22-component control mixtures, (2) three gasoline samples, and (3) three diesel samples. The three data sets were collected using slightly different temperature or pressure programs to engender significant retention time shifting in the raw data and then demonstrate subsequent corrections of that shifting upon comprehensive 2D alignment of the data sets. Thirty 12-min GC x GC separations from three 22-component control mixtures were used to evaluate the 2D alignment performance (10 runs/mixture). The average standard deviation of first column retention time improved 5-fold from 0.020 min (before alignment) to 0.004 min (after alignment). Concurrently, the average standard deviation of second column retention time improved 4-fold from 3.5 ms (before alignment) to 0.8 ms (after alignment). Alignment of the 30 control mixture chromatograms took 20 min. The quantitative integrity of the GC x GC data following 2D alignment was also investigated. The mean integrated signal was determined for all components in the three 22-component mixtures for all 30 replicates. The average percent difference in the integrated signal for each component before and after alignment was 2.6%. Singular value decomposition (SVD) was applied to the 22-component control mixture data before and after alignment to show the restoration of trilinearity to the data, since trilinearity benefits chemometric analysis. By applying comprehensive 2D retention time alignment to all three data sets (control mixtures, gasoline samples, and diesel samples), classification by principal component analysis (PCA) substantially improved, resulting in 100% accurate scores clustering.

Journal Article↗

Classification of gasoline data obtained by gas chromatography using a piecewise alignment algorithm combined with feature selection and principal component analysis.

A fast and objective chemometric classification method is developed and applied to the analysis of gas chromatography (GC) data from five commercial gasoline samples. The gasoline samples serve as model mixtures, whereas the focus is on the development and demonstration of the classification method. The method is based on objective retention time alignment (referred to as piecewise alignment) coupled with analysis of variance (ANOVA) feature selection prior to classification by principal component analysis (PCA) using optimal parameters. The degree-of-class-separation is used as a metric to objectively optimize the alignment and feature selection parameters using a suitable training set thereby reducing user subjectivity, as well as to indicate the success of the PCA clustering and classification. The degree-of-class-separation is calculated using Euclidean distances between the PCA scores of a subset of the replicate runs from two of the five fuel types, i.e., the training set. The unaligned training set that was directly submitted to PCA had a low degree-of-class-separation (0.4), and the PCA scores plot for the raw training set combined with the raw test set failed to correctly cluster the five sample types. After submitting the training set to piecewise alignment, the degree-of-class-separation increased (1.2), but when the same alignment parameters were applied to the training set combined with the test set, the scores plot clustering still did not yield five distinct groups. Applying feature selection to the unaligned training set increased the degree-of-class-separation (4.8), but chemical variations were still obscured by retention time variation and when the same feature selection conditions were used for the training set combined with the test set, only one of the five fuels was clustered correctly. However, piecewise alignment coupled with feature selection yielded a reasonably optimal degree-of-class-separation for the training set (9.2), and when the same alignment and ANOVA parameters were applied to the training set combined with the test set, the PCA scores plot correctly classified the gasoline fingerprints into five distinct clusters.

Algorithms↗

High-speed peak matching algorithm for retention time alignment of gas chromatographic data for chemometric analysis.

A rapid retention time alignment algorithm was developed as a preprocessing utility to be used prior to chemometric analysis of large datasets of diesel fuel profiles obtained using gas chromatography (GC). Retention time variation from chromatogram-to-chromatogram has been a significant impediment against the use of chemometric techniques in the analysis of chromatographic data due to the inability of current chemometric techniques to correctly model information that shifts from variable to variable within a dataset. The alignment algorithm developed is shown to increase the efficacy of pattern recognition methods applied to diesel fuel chromatograms by retaining chemical selectivity while reducing chromatogram-to-chromatogram retention time variations and to do so on a time scale that makes analysis of large sets of chromatographic data practical. Two sets of diesel fuel gas chromatograms were studied using the novel alignment algorithm followed by principal component analysis (PCA). In the first study, retention times for corresponding chromatographic peaks in 60 chromatograms varied by as much as 300 ms between chromatograms before alignment. In the second study of 42 chromatograms, the retention time shifting exhibited was on the order of 10 s between corresponding chromatographic peaks, and required a coarse retention time correction prior to alignment with the algorithm. In both cases, an increase in retention time precision afforded by the algorithm was clearly visible in plots of overlaid chromatograms before and then after applying the retention time alignment algorithm. Using the alignment algorithm, the standard deviation for corresponding peak retention times following alignment was 17 ms throughout a given chromatogram, corresponding to a relative standard deviation of 0.003% at an average retention time of 8 min. This level of retention time precision is a 5-fold improvement over the retention time precision initially provided by a state-of-the-art GC instrument equipped with electronic pressure control and was critical to the performance of the chemometric analysis. This increase in retention time precision does not come at the expense of chemical selectivity, since the PCA results suggest that essentially all of the chemical selectivity is preserved. Cluster resolution between dissimilar groups of diesel fuel chromatograms in a two-dimensional scores space generated with PCA is shown to substantially increase after alignment. The alignment method is robust against missing or extra peaks relative to a target chromatogram used in the alignment, and operates at high speed, requiring roughly 1 s of computation time per GC chromatogram.

Algorithms↗