PubMed Health⌕ Search

Biomedical subjects

John B O Mitchell

Publications and source records attributed to John B O Mitchell.

11 recordsLinked to original sources

MACiE (Mechanism, Annotation and Classification in Enzymes): novel tools for searching catalytic mechanisms.

MACiE (Mechanism, Annotation and Classification in Enzymes) is a database of enzyme reaction mechanisms, and is publicly available as a web-based data resource. This paper presents the first release of a web-based search tool to explore enzyme reaction mechanisms in MACiE. We also present Version 2 of MACiE, which doubles the dataset available (from Version 1). MACiE can be accessed from http://www.ebi.ac.uk/thornton-srv/databases/MACiE/

Catalysis↗

MACiE: a database of enzyme reaction mechanisms.

SUMMARY: MACiE (mechanism, annotation and classification in enzymes) is a publicly available web-based database, held in CMLReact (an XML application), that aims to help our understanding of the evolution of enzyme catalytic mechanisms and also to create a classification system which reflects the actual chemical mechanism (catalytic steps) of an enzyme reaction, not only the overall reaction. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/macie/.

Catalysis↗

Communication and re-use of chemical information in bioscience.

The current methods of publishing chemical information in bioscience articles are analysed. Using 3 papers as use-cases, it is shown that conventional methods using human procedures, including cut-and-paste are time-consuming and introduce errors. The meaning of chemical terms and the identity of compounds is often ambiguous. valuable experimental data such as spectra and computational results are almost always omitted. We describe an Open XML architecture at proof-of-concept which addresses these concerns. Compounds are identified through explicit connection tables or links to persistent Open resources such as PubChem. It is argued that if publishers adopt these tools and protocols, then the quality and quantity of chemical information available to bioscientists will increase and the authors, publishers and readers will find the process cost-effective.

Archives↗

Chemistry in bioinformatics.

Chemical information is now seen as critical for most areas of life sciences. But unlike Bioinformatics, where data is openly available and freely re-usable, most chemical information is closed and cannot be re-distributed without permission. This has led to a failure to adopt modern informatics and software techniques and therefore paucity of chemistry in bioinformatics. New technology, however, offers the hope of making chemical data (compounds and properties) free during the authoring process. We argue that the technology is already available; we require a collective agreement to enhance publication protocols.

Access to Information↗

A structure-odour relationship study using EVA descriptors and hierarchical clustering.

Structure-odour relationship analyses using hierarchical clustering were carried out on a diverse dataset of 47 molecules. These molecules were divided into seven odour categories: ambergris, bitter almond, camphoraceous, rose, jasmine, muguet, and musk. The alignment-independent descriptor EVA (EigenVAlue) was used as the molecular descriptor. The results were compared with those of another kind of descriptor, the UNITY 2D fingerprint. The dendrograms obtained with these descriptors were compared with the seven odour categories using the adjusted Rand index. The dendrograms produced by EVA consistently outperformed those from UNITY 2D in reproducing the experimental odour classifications of these 47 molecules.

Algorithms↗

Predicting protein-ligand binding affinities: a low scoring game?

We have investigated the performance of five well known scoring functions in predicting the binding affinities of a diverse set of 205 protein-ligand complexes with known experimental binding constants, and also on subsets of mutually similar complexes. We have found that the overall performance of the scoring functions on the diverse set is disappointing, with none of the functions achieving r(2) values above 0.32 on the whole dataset. Performance on the subsets was mixed, with four of the five functions predicting fairly well the binding affinities of 35 proteinases, but none of the functions producing any useful correlation on a set of 38 aspartic proteinases. We consider two algorithms for producing consensus scoring functions, one based on a linear combination of scores from the five individual functions and the other on averaging the rankings produced by the five functions. We find that both algorithms produce consensus functions that generally perform slightly better than the best individual scoring function on a given dataset.

Algorithms↗

Can we predict lattice energy from molecular structure?

By using simply the numbers of occurrences of different atom types as descriptors, a conceptually transparent and remarkably accurate model for the prediction of the enthalpies of sublimation of organic compounds has been generated. The atom types are defined on the basis of atomic number, hybridization state and bonded environment. Models of this kind were applied firstly to aliphatic hydrocarbons, secondly to both aliphatic and aromatic hydrocarbons, thirdly to a wide range of non-hydrogen-bonding molecules, and finally to a set of 226 organic compounds including 70 containing hydrogen-bond donors and acceptors. The final model gives squared correlation coefficients of 0.925 for the 226 compounds in the training set and 0.937 for an independent test set of 35 compounds. The success of such a simple model implies that the enthalpy of sublimation can be predicted accurately without knowledge of the crystal packing. This hypothesis is in turn consistent with the idea that, rather than being determined by the particular features of the lowest-energy packing, the lattice energy is similar for a number of hypothetical alternative crystal structures of a molecule.

Journal Article↗

L/D Protein Ligand Database (PLD): additional understanding of the nature and specificity of protein-ligand complexes.

SUMMARY: The Protein Ligand Database (PLD) is a publicly available web-based database that aims to provide further understanding of protein-ligand interactions. The PLD contains biomolecular data including calculated binding energies, Tanimoto ligand similarity scores and protein percentage sequence similarities. The database has potential for application as a tool in molecular design. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/pld/

Amino Acid Sequence↗

D-amino acid residues in peptides and proteins.

We have investigated the D-amino acid residues present in Protein Data Bank (PDB) entries, categorizing them into "real" D-residues and artifacts. In polypeptide chains of more than 20 residues, only a single instance of a "real" D-residue, other than those deliberately designed or engineered, was found. This example was the result of a slow chemical epimerization process. Another 12 designed D-residues were found in these longer polypeptide chains. Smaller peptides of 20 or fewer residues contained 479 "real" D-residues, the majority in various gramicidin, actinomycin, or cyclosporin structures. We found 148 PDB entries with "real" D-residues and a further 186, in which all apparent D-residues are artifacts. Investigating the (phi, psi) preferences of the "real" D-residues, we found that the region around (-60 degrees, -45 degrees ) was almost completely unoccupied, even though it is not formally disallowed. We link the low propensity to occupy this region with the alpha-helix destabilizing properties of D-residues.

Amino Acids↗

Chemoinformatics-based classification of prohibited substances employed for doping in sport.

Representative molecules from 10 classes of prohibited substances were taken from the World Anti-Doping Agency (WADA) list, augmented by molecules from corresponding activity classes found in the MDDR database. Together with some explicitly allowed compounds, these formed a set of 5245 molecules. Five types of fingerprints were calculated for these substances. The random forest classification method was used to predict membership of each prohibited class on the basis of each type of fingerprint, using 5-fold cross-validation. We also used a k-nearest neighbors (kNN) approach, which worked well for the smallest values of k. The most successful classifiers are based on Unity 2D fingerprints and give very similar Matthews correlation coefficients of 0.836 (kNN) and 0.829 (random forest). The kNN classifiers tend to give a higher recall of positives at the expense of lower precision. A naïve Bayesian classifier, however, lies much further toward the extreme of high recall and low precision. Our results suggest that it will be possible to produce a reliable and quantitative assignment of membership or otherwise of each class of prohibited substances. This should aid the fight against the use of bioactive novel compounds as doping agents, while also protecting athletes against unjust disqualification.

Algorithms↗

Melting point prediction employing k-nearest neighbor algorithms and genetic parameter optimization.

We have applied the k-nearest neighbor (kNN) modeling technique to the prediction of melting points. A data set of 4119 diverse organic molecules (data set 1) and an additional set of 277 drugs (data set 2) were used to compare performance in different regions of chemical space, and we investigated the influence of the number of nearest neighbors using different types of molecular descriptors. To compute the prediction on the basis of the melting temperatures of the nearest neighbors, we used four different methods (arithmetic and geometric average, inverse distance weighting, and exponential weighting), of which the exponential weighting scheme yielded the best results. We assessed our model via a 25-fold Monte Carlo cross-validation (with approximately 30% of the total data as a test set) and optimized it using a genetic algorithm. Predictions for drugs based on drugs (separate training and test sets each taken from data set 2) were found to be considerably better [root-mean-squared error (RMSE)=46.3 degrees C, r2=0.30] than those based on nondrugs (prediction of data set 2 based on the training set from data set 1, RMSE=50.3 degrees C, r2=0.20). The optimized model yields an average RMSE as low as 46.2 degrees C (r2=0.49) for data set 1, and an average RMSE of 42.2 degrees C (r2=0.42) for data set 2. It is shown that the kNN method inherently introduces a systematic error in melting point prediction. Much of the remaining error can be attributed to the lack of information about interactions in the liquid state, which are not well-captured by molecular descriptors.

Algorithms↗