PubMed Health⌕ Search

Biomedical subjects

R Put

Publications and source records attributed to R Put.

7 recordsLinked to original sources

Retention prediction of peptides based on uninformative variable elimination by partial least squares.

A quantitative structure-retention relationship analysis was performed on the chromatographic retention data of 90 peptides, measured by gradient elution reversed-phase liquid chromatography, and a large set of molecular descriptors computed for each peptide. Such approach may be useful in proteomics research in order to improve the correct identification of peptides. A principal component analysis on the set of 1726 molecular descriptors reveals a high information overlap in the descriptor space. Since variable selection is advisable, the retention of the peptides is modeled with uninformative variable elimination partial least squares, besides classic partial least squares regression. The Kennard and Stone algorithm was used to select a calibration set (63 peptides) from the available samples. This set was used to build the quantitative structure-retention relationship models. The remaining 27 peptides were used as independent external test set to evaluate the predictive power of the constructed models. The UVE-PLS model consists of 5 components only (compared to 7 components in the best PLS model), and has the best predictive properties, i.e., the average error on the retention time is less than 30 s. When compared also to stepwise regression and an empirical model, the obtained UVE-PLS model leads to better and much better predictions, respectively.

Algorithms↗

Evaluation of chemometric techniques to select orthogonal chromatographic systems.

Several chemometric techniques were compared for their performance to determine the orthogonality and similarity between chromatographic systems. Pearson's correlation coefficient (r) based color maps earlier were used to indicate selectivity differences between systems. These maps, in which the systems were ranked according to decreasing or increasing dissimilarities observed in the weighted-average-linkage dendrogram, were now applied as reference method. A number of chemometric techniques were evaluated as potential alternative (visualization) methods for the same purpose. They include hierarchical clustering techniques (single, complete, unweighted-average-linkage, centroid and Ward's method), the Kennard and Stone algorithm, auto-associative multivariate regression trees (AAMRT), and the generalized pairwise correlation method (GPCM) with McNemar's statistical test. After all, the reference method remained our preferred technique to select orthogonal and identify similar systems.

Algorithms↗

Selection of orthogonal reversed-phase HPLC systems by univariate and auto-associative multivariate regression trees.

In order to select chromatographic starting conditions to be optimized during further method development of the separation of a given mixture, so-called generic orthogonal chromatographic systems could be explored in parallel. In this paper the use of univariate and multivariate regression trees (MRT) was studied to define the most orthogonal subset from a given set of chromatographic systems. Two data sets were considered, which contain the retention data of 68 structurally diversive drugs on sets of 32 and 38 chromatographic systems, respectively. For both the univariate and multivariate approaches no other data but the measured retention factors are needed to build the decision trees. Since multivariate regression trees are used in an unsupervised way, they are called auto-associative multivariate regression trees (AAMRT). For all decision trees used, a variable importance list of the predictor variables can be derived. It was concluded that based on these ranked lists, both for univariate and multivariate regression trees, a selection of the most orthogonal systems from a given set of systems can be obtained in a user-friendly and fast way.

Chromatography, High Pressure Liquid↗

Prediction of gastro-intestinal absorption using multivariate adaptive regression splines.

Multivariate adaptive regression splines (MARS) and a derived method two-step MARS (TMARS) were used for modelling the gastro-intestinal absorption of 140 drug-like molecules. The published absorption values for these molecules were used as response variable and calculated molecular descriptors as potential explanatory variables. Both methods were compared and their potential use in quantitative structure-activity relationship (QSAR) context evaluated. The predictive abilities of the models were studied using different sequences of Monte Carlo cross validation (MCCV). It was shown that both types of models had good predictive abilities and that for the data used, MARS gave better results than TMARS. It could be concluded that both methods could be valuable for QSAR modelling.

Algorithms↗

Multivariate adaptive regression splines (MARS) in chromatographic quantitative structure-retention relationship studies.

The multivariate adaptive regression splines (MARS) methodology was applied to build quantitative structure-retention relationships (QSRRs). The response (dependent variable) in the MARS models consisted of the logarithms of the extrapolated retention factors (log k(w)) of 83 structurally diverse drugs on a Unisphere PBD column, using isocratic elutions at pH 11.7. A set of 266 molecular descriptors was used as predictor (independent) variables in the MARS model building. The optimal MARS model uses 34 basis functions to describe the retention and has acceptable predictive properties for new objects. The molecular descriptors included in the model describe hydrophobicity, molecular size, complexity, shape and polarisability. Some additional MARS models were created using alternative strategies. These include models with log P as the single predictor and models obtained with only the three most important molecular descriptors. The use of classification and regression trees (CART) as feature selection technique for predictor variables used in the MARS model was also investigated. Further, it is also studied whether allowing quadratic terms instead of interaction terms might lead to better MARS models.

Chromatography, Liquid↗

Classification and regression tree analysis for molecular descriptor selection and retention prediction in chromatographic quantitative structure-retention relationship studies.

The use of the classification and regression tree (CART) methodology was studied in a quantitative structure-retention relationship (QSRR) context on a data set consisting of the retentions of 83 structurally diverse drugs on a Unisphere PBD column, using isocratic elutions at pH 11.7. The response (dependent variable) in the tree models consisted of the predicted rention factor (log kw) of the solutes, while a set of 266 molecular descriptors was used as explanatory variables in the tree building. Molecular descriptors related to the hydrophobicity (log P and Hy) and the size (TPC) of the molecules were selected out of these 266 descriptors in order to describe and predict retention. Besides the above mentioned, CART was also able to select hydrogen-bonding and molecular complexity descriptors. Since these variables are expected from QSRR knowledge, it demonstrates the potential of CART as a methodology to understand retention in chromatographic systems. The potential of CART to predict retention and thus occasionally to select an appropriate system for a given mixture was also evaluated. Reasonably good prediction, i.e. only 9% serious misclassification, was observed. Moreover, some of the misclassifications probably are inherent to the data set applied.

Chromatography↗