PubMed Health⌕ Search

Biomedical subjects

M Daszykowski

Publications and source records attributed to M Daszykowski.

10 recordsLinked to original sources

A comparison of three algorithms for chromatograms alignment.

In this paper the performance of three alignment algorithms, correlation optimized warping, parametric time warping and semi-parametric time warping, is compared on real chromatograms. Among these, parametric time warping is the simplest and fastest; generally less than 1s is required to align two chromatograms. It does not require the optimization of input parameters and allows the alignment of peak shifts in only one direction, or non-complex peak shifts in both directions. With correlation optimized warping and semi-parametric time warping complex peak shifts in both directions can be corrected but at the expense of the optimization of two input parameters. Semi-parametric time warping requires the selection of the proper number of B-splines in the warping function and, if necessary, the optimization of the penalty parameter. Often the default values can be used to obtain aligned signals. The optimization of the input parameters for correlation optimized warping (section length, slack) is not easy and time-consuming. Moreover, dependent on the input parameters, the computation time of the correlation optimized warping algorithm can be twice as long as for semi-parametric time warping for which computation times up to 23 s are required. However, the performance of both algorithms is equally good considering the improvement of the precision of the peak retention times and correlation coefficients between the chromatograms, after alignment. For the data aligned in this study, the average retention time precision and the lowest correlation before warping were 14 and 0.17, and were improved to three and 0.83, and six and 0.87 after warping, with correlation optimized warping and semi-parametric time warping, respectively.

Algorithms↗

Retention prediction of peptides based on uninformative variable elimination by partial least squares.

A quantitative structure-retention relationship analysis was performed on the chromatographic retention data of 90 peptides, measured by gradient elution reversed-phase liquid chromatography, and a large set of molecular descriptors computed for each peptide. Such approach may be useful in proteomics research in order to improve the correct identification of peptides. A principal component analysis on the set of 1726 molecular descriptors reveals a high information overlap in the descriptor space. Since variable selection is advisable, the retention of the peptides is modeled with uninformative variable elimination partial least squares, besides classic partial least squares regression. The Kennard and Stone algorithm was used to select a calibration set (63 peptides) from the available samples. This set was used to build the quantitative structure-retention relationship models. The remaining 27 peptides were used as independent external test set to evaluate the predictive power of the constructed models. The UVE-PLS model consists of 5 components only (compared to 7 components in the best PLS model), and has the best predictive properties, i.e., the average error on the retention time is less than 30 s. When compared also to stepwise regression and an empirical model, the obtained UVE-PLS model leads to better and much better predictions, respectively.

Algorithms↗

Evaluation of chemometric techniques to select orthogonal chromatographic systems.

Several chemometric techniques were compared for their performance to determine the orthogonality and similarity between chromatographic systems. Pearson's correlation coefficient (r) based color maps earlier were used to indicate selectivity differences between systems. These maps, in which the systems were ranked according to decreasing or increasing dissimilarities observed in the weighted-average-linkage dendrogram, were now applied as reference method. A number of chemometric techniques were evaluated as potential alternative (visualization) methods for the same purpose. They include hierarchical clustering techniques (single, complete, unweighted-average-linkage, centroid and Ward's method), the Kennard and Stone algorithm, auto-associative multivariate regression trees (AAMRT), and the generalized pairwise correlation method (GPCM) with McNemar's statistical test. After all, the reference method remained our preferred technique to select orthogonal and identify similar systems.

Algorithms↗

Prediction of total green tea antioxidant capacity from chromatograms by multivariate modeling.

In this paper, a fast strategy for determining the total antioxidant capacity of Chinese green tea extracts is developed. This strategy includes the use of experimental techniques, such as fast high-performance liquid chromatography (HPLC) on monolithic columns and a spectrophotometric approach to determine the total antioxidant capacity of the extracts. To extract the chemically relevant information from the obtained data, chemometrical approaches are used. Among them there are correlation optimized warping (COW) to align the chromatograms, robust principal component analysis (robust PCA) to detect outliers, and partial least squares (PLS) and uninformative variable elimination partial least squares (UVE-PLS) to construct a reliable multivariate regression model to predict the total antioxidant capacity from the fast chromatograms.

Antioxidants↗

Chemometrical exploration of the wet precipitation chemistry from the Austrian Monitoring Network (1988-1999).

The present paper deals with the application of different chemometric methods to an environmental data set derived from the monitoring of wet precipitation in Austria (1988-1999). These methods are: principal component analysis (PCA); projection pursuit (PP); density-based spatial clustering of application with noise (DBSCAN); ordering points to identify the clustering structures (OPTICS); self-organizing maps (SOM), also called the Kohonen network; and the neural gas (NG) network. The aim of the study is to introduce some new approaches into environmetrics and to compare their usefulness with already existing techniques for the classification and interpretation of environmental data. The density-based approaches give information about the occurrence of natural clusters in the studied data set, which, however, do not occur in the case presented here; information about high-density zones (very similar samples) and extreme samples is also obtained. The partitioning techniques (clustering, but also neural gas and Kohonen networks) offer an opportunity to classify the objects of interest into several defined groups, the patterns of ionic concentration of which can be studied in detail. The visual aids, such as the color map and the Kohonen map, for each site are very helpful in understanding the relationships between samples and between samples and variables. All methods, and in particular projection pursuit, give information about samples with extreme characteristics.

Austria↗

Determining orthogonal chromatographic systems prior to the development of methods to characterise impurities in drug substances.

To define starting conditions for the development of methods to separate impurities from the active substance and from each other in drugs with an unknown impurity profile, the parallel application of generic orthogonal chromatographic systems could be useful. The possibilities to define orthogonal chromatographic systems were examined by calculation of the correlation coefficients between retention factors k for a set of 68 drugs on 11 systems, by visual evaluation of the selectivity differences, by using principal component analysis, by drawing color maps and evaluating dendrograms. A zirconia-based stationary phase coated with a polybutadiene (PBD) polymer and three silica-based phases (base-deactivated, polar-embedded and monolithic) were used. Besides the stationary phase, the influence of pH and of organic modifier, on the selectivity of a system were evaluated. The dendrograms of hierarchical clusters were found good aids to assess orthogonality of chromatographic systems. The PBD-zirconia phase/methanol/pH 2.5 system is found most orthogonal towards several silica-based systems, e.g. a base-deactivated C16 -amide silica/methanol/pH 2.5 system. The orthogonality was validated using cross-validation, and two other validation sets, i.e. a set of non-ionizable solutes and a mixture of a drug and its impurities.

Chromatography, High Pressure Liquid↗

Looking for natural patterns in analytical data. 2. Tracing local density with OPTICS.

The main principles and the algorithm of a density-based clustering approach, OPTICS, are described, and its unique properties, such as the ability to reveal clusters of arbitrary shapes and different densities, are illustrated on simulated and real spectral and chromatographic data sets. A "reachability plot" visualizing density fluctuations of data in multivariate space and a "color map" relating the original and/or descriptive features with data clustering allow a deeper insight into the data structure and its interpretation in chemical terms.

Journal Article↗

On the optimal partitioning of data with K-means, growing K-means, neural gas, and growing neural gas.

In this paper, the performance of new clustering methods such as Neural Gas (NG) and Growing Neural Gas (GNG) is compared with the K-means method for real and simulated data sets. Moreover, a new algorithm called growing K-means, GK, is introduced as the alternative to Neural Gas and Growing Neural Gas. It has small input requirements and is conceptually very simple. The GK leads to nearly optimal values of the cost function, and, contrary to K-means, it is independent of the initial data set partition. The incremental property of GK additionally helps to estimate the number of "natural" clusters in data, i.e., the well-separated groups of objects in the data space.

Journal Article↗

Classification and regression trees--studies of HIV reverse transcriptase inhibitors.

In this paper, the application of Classification And Regression Trees (CART) is presented for the analysis of biological activity of Non-Nucleoside Reverse Transcriptase Inhibitors (NNRTIs). The data consist of the biological activities, expressed as pIC50, of 208 NNRTIs against wild-type HIV virus (HIV-1) and four mutant strains (181C, 103N, 100I, 188L) and the computed interaction energies with the Reverse Transcriptase (RT) binding pocket. CART explains the observed biological activity of NNRTIs in terms of interactions with individual amino acids in the RT binding pocket, i.e., the original data variables.

Algorithms↗