PubMed Health⌕ Search

Biomedical subjects

Peter C Jurs

Publications and source records attributed to Peter C Jurs.

At least 19 recordsLinked to original sources

Assessing the reliability of a QSAR model's predictions.

Quantitative structure activity relationships (QSAR) are one of the well-developed areas in computational chemistry. In this field, many successful predictive models have been developed for various property, activity or toxicity predictions. However, the predictive power of models for new query compounds is often not well characterized. The breadth of applicability of models is often not characterized. In other words, with a given QSAR model and a specific query compound to be predicted, can the model be used reliably for the desired prediction? In this study, we assessed the reliability of QSAR models' prediction on query compounds. Our approach, employing hierarchical clustering, was developed and tested using a test dataset containing 322 organic compounds with fathead minnow acute aquatic toxicity as the activity of interest. The hypothesis of the approach was that if a query compound is more similar to the compounds used to generate the QSAR model, it should be predicted more accurately. Thus, the core of the approach is to determine the relationship between the similarity of query compounds to the training set compounds of the QSAR model and the prediction accuracy given by that model. This relationship determination was achieved by comparing the results given by the two major components of the approach: objects clustering and activity prediction. With the resultant information from the two steps, a direct relationship was shown.

Models, Chemical↗

Probabilistic neural network multiple classifier system for predicting the genotoxicity of quinolone and quinoline derivatives.

Quinolone and quinoline are known to be liver carcinogens in rodents, and a number of their derivatives have been shown to exhibit mutagenicity in the Ames test, using Salmonella typhimurium strain TA 100 in the presence of S9. Both the carcinogenicity and the mutagenicity of quinolone and quinoline derivatives, as determined by SAS, can be attributed to their genotoxicity potential. This potential, which is measured by genotoxicity tests, is a good indication of carcinogenicity and mutagenicity because compounds that are positive in these tests have the potential to be human carcinogens and/or mutagens. In this study, a collection of quinolone and quinoline derivatives' carcinogenicity is determined by qualitatively predicting their genotoxicity potential with predictive PNN (probabilistic neural network) classification models. In addition, a multiple classifier system is also developed to improve the predictability of genotoxicity. Superior results are seen with the multiple classifier system over the individual PNN classification models. With the multiple classifier system, 89.4% of the quinolone derivatives were predicted correctly, and higher predictability is seen with the quinoline derivatives at 92.2% correct. The multiple classifier system not only is able to accurately predict the genotoxicity but also provides an insight about the main determinants of genotoxicity of the quinolone and quinoline derivatives. Thus, the PNN multiple classifier system generated in this study is a beneficial contributor toward predictive toxicology in the design of less carcinogenic bioactive compounds.

Animals↗

Generation of QSAR sets with a self-organizing map.

A Kohonen self-organizing map (SOM) is used to classify a data set consisting of dihydrofolate reductase inhibitors with the help of an external set of Dragon descriptors. The resultant classification is used to generate training, cross-validation (CV) and prediction sets for QSAR modeling using the ADAPT methodology. The results are compared to those of QSAR models generated using sets created by activity binning and a sphere exclusion method. The results indicate that the SOM is able to generate QSAR sets that are representative of the composition of the overall data set in terms of similarity. The resulting QSAR models are half the size of those published and have comparable RMS errors. Furthermore, the RMS errors of the QSAR sets are consistent, indicating good predictive capabilities as well as generalizability.

Algorithms↗

QSAR and classification of murine and human soluble epoxide hydrolase inhibition by urea-like compounds.

A data set of 348 urea-like compounds that inhibit the soluble epoxide hydrolase enzyme in mice and humans is examined. Compounds having IC(50) values ranging from 0.06 to >500 microM (murine) and 0.10 to >500 microM (human) are categorized as active or inactive for classification, while quantitation is performed on smaller compound subsets ranging from 0.07 to 431 microM (murine) and 0.11 to 490 microM (human). Each compound is represented by calculated structural descriptors that encode topological, geometrical, electronic, and polar surface features. Multiple linear regression (MLR) and computational neural networks (CNNs) are employed for quantitative models. Three classification algorithms, k-nearest neighbor (kNN), linear discriminant analysis (LDA), and radial basis function neural networks (RBFNN), are used to categorize compounds as active or inactive based on selected data split points. Quantitative modeling of human enzyme inhibition results in a nonlinear, five-descriptor model with root-mean-square errors (log units of IC(50) [microM]) of 0.616 (r(2) = 0.66), 0.674 (r(2) = 0.61), and 0.914 (r(2) = 0.33) for training, cross-validation, and prediction sets, respectively. The best classification results for human and murine enzyme inhibition are found using kNN. Human classification rates using a seven-descriptor model for training and prediction sets are 89.1% and 91.4%, respectively. Murine classification rates using a five-descriptor model for training and prediction sets are 91.5% and 88.6%, respectively.

Animals↗

Prediction of dihydrofolate reductase inhibition and selectivity using computational neural networks and linear discriminant analysis.

A data set of 345 dihydrofolate reductase inhibitors was used to build QSAR models that correlate chemical structure and inhibition potency for three types of dihydrofolate reductase (DHFR): rat liver (rl), Pneumocystis carinii (pc), and Toxoplasma gondii (tg). Quantitative models were built using subsets of molecular structure descriptors being analyzed by computational neural networks. Neural network models were able to accurately predict log IC(50) values for the three types of DHFR to within +/-0.65 log units (data sets ranged approximately 5.5 log units) of the experimentally determined values. Classification models were also constructed using linear discriminant analysis to identify compounds as selective or nonselective inhibitors of bacterial DHFR (pcDHFR and tgDHFR) relative to mammalian DHFR (rlDHFR). A leave-N-out training procedure was used to add robustness to the models and to prove that consistent results could be obtained using different training and prediction set splits. The best linear discriminant analysis (LDA) models were able to correctly predict DHFR selectivity for approximately 70% of the external prediction set compounds. A set of new nitrogen and oxygen-specific descriptors were developed especially for this data set to better encode structural features, which are believed to directly influence DHFR inhibition and selectivity.

Animals↗

Predicting the genotoxicity of thiophene derivatives from molecular structure.

We report several binary classification models that directly link the genetic toxicity of a series of 140 thiophene derivatives with information derived from the compounds' molecular structure. Genetic toxicity was measured using an SOS Chromotest. IMAX (maximal SOS induction factor) values were recorded for each of the 140 compounds both in the presence and in the absence of S9 rat liver homogenate. Compounds were classified as genotoxic if IMAX >or= 1.5 in either test or nongenotoxic if IMAX < 1.5 for both tests. The molecular structures were represented by numerical descriptors that encoded the topological, geometric, electronic, and polar surface area properties of the thiophene derivatives. The classification models used were linear discriminant analysis (LDA), k-nearest neighbor classification (k-NN), and the probabilistic neural network (PNN). These were used in conjunction with either a genetic algorithm or a generalized simulated annealing to find optimal subsets of descriptors for each classifier. The quality of the resulting models was determined by the number of misclassified compounds, with preference given to models that produced fewer false negative classifications. Model sizes ranged from seven descriptors for LDA to three descriptors for k-NN and PNN. Very good classification results were obtained with all three classifiers. Classification rates for the LDA, k-NN, and PNN models were 80, 85, and 85%, respectively, for the prediction set compounds. Additionally, a consensus model was generated that incorporated all three of the basic model types. This consensus model correctly predicted the genotoxicity of 95% of the prediction set compounds.

DNA Damage↗

Predicting the genotoxicity of polycyclic aromatic compounds from molecular structure with different classifiers.

Classification models were developed to provide accurate prediction of genotoxicity of 277 polycyclic aromatic compounds (PACs) directly from their molecular structures. Numerical descriptors encoding the topological, geometric, electronic, and polar surface area properties of the compounds were calculated to represent the structural information. Each compound's genotoxicity was represented with IMAX (maximal SOS induction factor) values measured by the SOS Chromotest in the presence and absence of S9 rat liver homogenate. The compounds' class identity was determined by a cutoff IMAX value of 1.25-compounds with IMAX > 1.25 in either test were classified as genotoxic, and the ones with IMAX < or = 1.25 were nongenotoxic. Several binary classification models were generated to predict genotoxicity: k-nearest neighbor (k-NN), linear discriminant analysis, and probabilistic neural network. The study showed k-NN to provide the highest predictive ability among the three classifiers with a training set classification rate of 93.5%. A consensus model was also developed that incorporated the three classifiers and correctly predicted 81.2% of the 277 compounds. It also provided a higher prediction rate on the genotoxic class than any other single model.

Animals↗

Prediction of peptide ion collision cross sections from topological molecular structure and amino acid parameters.

Quantitative structure-property relationships (QSPRs) have been developed to predict the ion mobility spectrometry (IMS) collision cross sections of singly protonated lysine-terminated peptides using information derived from topological molecular structure and various amino acid parameters. The primary amino acid sequence alone is sufficient to accurately predict the collision cross section. The models were built using multiple linear regression (MLR) and computational neural networks (CNNs). The best MLR model found contains six descriptors and predicts 94 of 113 peptides (83%) to within 2% of their experimentally determined values. The best CNN model using the same six descriptors predicts 105 of the 113 peptides (93%) to within 2% of their experimentally determined values. The best overall CNN model, using a different set of six descriptors, predicts 109 of the 113 peptides (96%) to within 2% of their experimentally determined values. In addition, this model can discriminate among peptides having identical amino acid composition, but differing in primary amino acid sequence. This represents a capability not found in previously described models. The descriptors used in the models presented may provide some insight into the nature of peptide ion folding in the gas phase.

Amino Acid Sequence↗

Prediction of glass transition temperatures from monomer and repeat unit structure using computational neural networks.

Quantitative structure-property relationships (QSPR) are developed to correlate glass transition temperatures and chemical structure. Both monomer and repeat unit structures are used to build several QSPR models for Parts 1 and 2 of this study, respectively. Models are developed using numerical descriptors, which encode important information about chemical structure (topological, electronic, and geometric). Multiple linear regression analysis (MLRA) and computational neural networks (CNNs) are used to generate the models after descriptor generation. Optimization routines (simulated annealing and genetic algorithm) are utilized to find information-rich subsets of descriptors for prediction. A 10-descriptor CNN model was found to be optimal in predicting T(g) values using the monomer structure (Part 1) for 165 polymers. A committee of 10 CNNs produced a training set rms error of 10.1K (r2 = 0.98) and a prediction set rms error of 21.7 K (r2 = 0.92). An 11-descriptor CNN model was developed for 251 polymers using the repeat unit structure (Part 2). A committee of CNNs produced a training set rms error of 21.1K (r2 = 0.96) and a prediction set rms error of 21.9 K (r2 = 0.96).

Journal Article↗

Development of quantitative structure-activity relationship and classification models for a set of carbonic anhydrase inhibitors.

Mathematical models are developed to find quantitative structure-activity relationships that correlate chemical structure and inhibition toward three carbonic anhydrase (CA) isozymes: CA I, II, and IV. Numerical descriptors are generated to encode important topological, geometric, and electronic features of molecular structure. After descriptor generation, multiple linear regression, and computational neural network (CNN) analyses are performed on various descriptor subsets to find superior models for prediction. Committees of five CNNs were utilized to average final predicted values for the 142-compound data set. For inhibitors of CA I, an 8-5-1 CNN committee produced a training set rms error of 0.105 log K(i) (r(2) = 0.994) and prediction set rms error of 0.208 log K(i) (r(2) = 0.980). Training and prediction set rms errors of 0.140 log K(i) (r(2) = 0.992) and 0.231 log K(i) (r(2) = 0.971), respectively, were produced by a 9-5-1 CNN committee for inhibitors of CA II. For prediction of CA IV inhibitors, an 8-5-1 CNN committee produced training and prediction set rms errors of 0.147 log K(i) (r(2) = 0.992) and 0.211 log K(i) (r(2) = 0.991), respectively. In addition, classification models were built using k-nearest neighbor (kNN) analysis to solve two- and three-class problems for inhibitors of CA IV. A three-descriptor classification model proved superior in labeling compounds as active or inactive inhibitors for the two-class problem. Training and prediction set percent classification rates of 100% and 87.1%, respectively, were obtained. For the three-class (active/moderate/inactive) problem, a five-descriptor model was deemed optimal producing a training set percent classification rate of 98.8% and prediction set rate of 79.0%.

Carbonic Anhydrase Inhibitors↗

QSAR/QSPR studies using probabilistic neural networks and generalized regression neural networks.

The Probabilistic Neural Network (PNN) and its close relative, the Generalized Regression Neural Network (GRNN), are presented as simple yet powerful neural network techniques for use in Quantitative Structure-Activity Relationship (QSAR) and Quantitative Structure-Property Relationship (QSPR) studies. The PNN methodology is applicable to classification problems, and the GRNN is applicable to continuous function mapping problems. The basic underlying theory behind these probability-based methods is presented along with two applications of the PNN/GRNN methodology. The PNN model presented identifies molecules as potential soluble epoxide hydrolase inhibitors using a binary classification scheme. The GRNN model presented predicts the aqueous solubility of nitrogen- and oxygen-containing small organic molecules. For each application, the network inputs consist of a small set of descriptors that encode structural features at the molecular level. Each of these studies has also been previously addressed in this research group using more traditional techniques such as k-nearest neighbor classification, multiple linear regression, and multilayer feed-forward neural networks. In each case, the predictive power of the PNN and GRNN models was found to be comparable to that of the more traditional techniques but requiring significantly fewer input descriptors.

Bayes Theorem↗

Predicting the genotoxicity of secondary and aromatic amines using data subsetting to generate a model ensemble.

Binary quantitative structure-activity relationship (QSAR) models are developed to classify a data set of 334 aromatic and secondary amine compounds as genotoxic or nongenotoxic based on information calculated solely from chemical structure. Genotoxic endpoints for each compound were determined using the SOS Chromotest in both the presence and absence of an S9 rat liver homogenate. Compounds were considered genotoxic if assay results indicated a positive genotoxicity hit for either the S9 inactivated or S9 activated assay. Each compound in the data set was encoded through the calculation of numerical descriptors that describe various aspects of chemical structure (e.g. topological, geometric, electronic, polar surface area). Furthermore, five additional descriptors that focused on the secondary and aromatic nitrogen atoms in each molecule were calculated specifically for this study. Descriptor subsets were examined using a genetic algorithm search engine interfaced with a k-Nearest Neighbor fitness evaluator to find the most information-rich subsets, which ultimately served as the final predictive models. Models were chosen for their ability to minimize the total number of misclassifications, with special attention given to those models that possessed fewer occurrences of positive toxicity hits being misclassified as nontoxic (false negatives). In addition, a subsetting procedure was used to form an ensemble of models using different combinations of compounds in the training and prediction sets. This was done to ensure that consistent results could be obtained regardless of training set composition. The procedure also allowed for each compound to be externally validated three times by different training set data with the resultant predictions being used in a "majority rules" voting scheme to produce a consensus prediction for each member of the data set. The individual models produced an average training set classification rate of 71.6% and an average prediction set classification rate of 67.7%. However, the model ensemble was able to correctly classify the genotoxicity of 72.2% of all prediction set compounds.

Algorithms↗

Classification of diverse organic compounds that induce chromosomal aberrations in Chinese hamster cells.

A data set of 297 diverse organic compounds that cause varying degrees of chromosomal aberrations in Chinese hamster lung cells is examined. Responses of an assay are categorized as clastogenic (>10% aberrant cells) and nonclastogenic (<5% aberrant cells). Each of the compounds is represented by calculated structural descriptors that encode topological, geometric, electronic, and polar surface features. A genetic algorithm (GA) employing a k-nearest neighbor (kNN) fitness evaluator is used to iteratively search a reduced descriptor space to find small, information-rich subsets of descriptors that maximize the classification rates for clastogenic and nonclastogenic responses. To further improve modeling, a similarity measure using atom-pair descriptors is employed to create more homogeneous data subsets. Three different data sets are examined. Results for a set of 297 compounds using the GA-kNN method were 86.5% and 80.0% correct classification in the training set and prediction set, respectively. Results for a subset of 279 compounds in model 2 are 85.7% and 85.7% for the training and prediction sets, respectively. Results for a subset of 182 compounds in model 3 are 91.5% and 94.4% for the training and prediction sets, respectively. Creating smaller, more topologically similar data sets result in improved classification rates.

Algorithms↗

Development and use of hydrophobic surface area (HSA) descriptors for computer-assisted quantitative structure-activity and structure-property relationship studies.

A new series of 25 whole-molecule molecular structure descriptors are proposed. The new descriptors are termed Hydrophobic Surface Area, or HSA descriptors, and are designed to capture information regarding the structural features responsible for hydrophobic and hydrophilic intermolecular interactions. The utility of the HSAs in capturing this type of information is demonstrated using two properties that have a known hydrophobic component. The first study involves the modeling of the inhibition of Gram-positive bacteria cell growth of a series of biarylamides. The second application involves the study of the blood-brain barrier penetration of a diverse series of drug molecules. In both cases, the HSAs are shown to effectively capture information related to the hydrophobic components of these two properties. Additional evaluation of the new class of descriptors shows them to be unique in their ability to measure hydrophobic features among a diverse set of conventional structural descriptors. The HSAs are evaluated regarding their sensitivity to conformational changes and are found to be similar in that regard to other widely used molecular descriptors.

Amides↗

Determining the validity of a QSAR model--a classification approach.

The determination of the validity of a QSAR model when applied to new compounds is an important concern in the field of QSAR and QSPR modeling. Various scoring techniques can be applied to specific types of models. We present a technique with which we can state whether a new compound will be well predicted by a previously built QSAR model. In this study we focus on linear regression models only, though the technique is general and could also be applied to other types of quantitative models. Our technique is based on a classification method that divides regression residuals from a previously generated model into a good class and bad class and then builds a classifier based on this division. The trained classifier is then used to determine the class of the residual for a new compound. We investigated the performance of a variety of classifiers, both linear and nonlinear. The technique was tested on two data sets from the literature and a hand built data set. The data sets selected covered both physical and biological properties and also presented the methodology with quantitative regression models of varying quality. The results indicate that this technique can determine whether a new compound will be well or poorly predicted with weighted success rates ranging from 73% to 94% for the best classifier.

Algorithms↗

Development of linear, ensemble, and nonlinear models for the prediction and interpretation of the biological activity of a set of PDGFR inhibitors.

A QSAR modeling study has been done with a set of 79 piperazyinylquinazoline analogues which exhibit PDGFR inhibition. Linear regression and nonlinear computational neural network models were developed. The regression model was developed with a focus on interpretative ability using a PLS technique. However, it also exhibits a good predictive ability after outlier removal. The nonlinear CNN model had superior predictive ability compared to the linear model with a training set error of 0.22 log(IC50) units (R2 = 0.93) and a prediction set error of 0.32 log(IC50) units (R2 = 0.61). A random forest model was also developed to provide an alternate measure of descriptor importance. This approach ranks descriptors, and its results confirm the importance of specific descriptors as characterized by the PLS technique. In addition the neural network model contains the two most important descriptors indicated by the random forest model.

Linear Models↗

Development of QSAR models to predict and interpret the biological activity of artemisinin analogues.

This work presents the development of Quantitative Structure-Activity Relationship (QSAR) models to predict the biological activity of 179 artemisinin analogues. The structures of the molecules are represented by chemical descriptors that encode topological, geometric, and electronic structure features. Both linear (multiple linear regression) and nonlinear (computational neural network) models are developed to link the structures to their reported biological activity. The best linear model was subjected to a PLS analysis to provide model interpretability. While the best linear model does not perform as well as the nonlinear model in terms of predictive ability, the application of PLS analysis allows for a sound physical interpretation of the structure-activity trend captured by the model. On the other hand, the best nonlinear model is superior in terms of pure predictive ability, having a training error of 0.47 log RA units (R2 = 0.96) and a prediction error of 0.76 log RA units (R2 = 0.88).

Antimalarials↗

Interpreting computational neural network QSAR models: a measure of descriptor importance.

We present a method to measure the relative importance of the descriptors present in a QSAR model developed with a computational neural network (CNN). The approach is based on a sensitivity analysis of the descriptors. We tested the method on three published data sets for which linear and CNN models were previously built. The original work reported interpretations for the linear models, and we compare the results of the new method to the importance of descriptors in the linear models as described by a PLS technique. The results indicate that the proposed method is able to rank descriptors such that important descriptors in the CNN model correspond to the important descriptors in the linear model.

Models, Molecular↗