PubMed Health⌕ Search

Biomedical subjects

João Aires-de-Sousa

Publications and source records attributed to João Aires-de-Sousa.

12 recordsLinked to original sources

QSAR analysis of phenolic antioxidants using MOLMAP descriptors of local properties.

Molecular maps of atom-level properties (MOLMAPs) were developed to represent the diversity of chemical bonds existing in a molecule. Chemical reactivity, being related to the ability for bond breaking and bond making, is primarily determined by the properties of bonds available in a molecule. In order to use physicochemical properties of individual bonds for an entire molecule, and at the same time having a fixed-length molecular representation, all the bonds of a molecule are mapped into a fixed-size 2D self-organizing map (MOLMAP). This article illustrates the application of MOLMAP descriptors to QSAR, with a study of the radical scavenging activity of 47 naturally occurring phenolic antioxidants. Counterpropagation neural networks (CPG NNs) were trained with MOLMAP descriptors selected using genetic algorithms to predict antioxidant activity. The model was subsequently validated by the leave-one-out (LOO) procedure obtaining a q(2) of 0.71. Random Forests were grown with the entire set of MOLMAP descriptors giving 70% of correct classifications as potent, active or inactive in a LOO experiment. Interpretations of both models in terms of discriminant variables were concordant and allowed identifying bonds and substructures that are mostly responsible for antioxidant activity. This work shows how MOLMAPs can be used for data mining of structural and biological activity data, leading to the extraction of relationships between local properties and activity.

Algorithms↗

Automatic assignment of absolute configuration from 1D NMR data.

[reaction: see text] Opposite enantiomers exhibit different NMR properties in the presence of an external common chiral element, and a chiral molecule exhibits different NMR properties in the presence of external enantiomeric chiral elements. Automatic prediction of such differences, and comparison with experimental values, leads to the assignment of the absolute configuration. Here two cases are reported, one using a dataset of 80 chiral secondary alcohols esterified with (R)-MTPA and the corresponding (1)H NMR chemical shifts and the other with 94 (13)C NMR chemical shifts of chiral secondary alcohols in two enantiomeric chiral solvents. For the first application, counterpropagation neural networks were trained to predict the sign of the difference between chemical shifts of opposite stereoisomers. The neural networks were trained to process the chirality code of the alcohol as the input, and to give the NMR property as the output. In the second application, similar neural networks were employed, but the property to predict was the difference of chemical shifts in the two enantiomeric solvents. For independent test sets of 20 objects, 100% correct predictions were obtained in both applications concerning the sign of the chemical shifts differences. Additionally, with the second dataset, the difference of chemical shifts in the two enantiomeric solvents was quantitatively predicted, yielding r(2) 0.936 for the test set between the predicted and experimental values.

Alcohols↗

Representation of DNA sequences with virtual potentials and their processing by (SEQREP) Kohonen self-organizing maps.

MOTIVATION: We propose representing individual positions in DNA sequences by virtual potentials generated by other bases of the same sequence. This is a compact representation of the neighbourhood of a base. The distribution of the virtual potentials over the whole sequence can be used as a representation of the entire sequence (SEQREP code). It is a flexible code, with a length independent of the sequence size, does not require previous alignment, and is convenient for processing by neural networks or statistical techniques. RESULTS: To evaluate its biological significance, the SEQREP code was used for training Kohonen self-organizing maps (SOMs) in two applications: (a) detection of Alu sequences, and (b) classification of sequences encoding for HIV-1 envelope glycoprotein (env) into subtypes A-G. It was demonstrated that SOMs clustered sequences belonging to different classes into distinct regions. For independent test sets, very high rates of correct predictions were obtained (97% in the first application, 91% in the second). Possible areas of application of SEQREP codes include functional genomics, phylogenetic analysis, detection of repetitions, database retrieval, and automatic alignment. AVAILABILITY: Software for representing sequences by SEQREP code, and for training Kohonen SOMs is made freely available from http://www.dq.fct.unl.pt/qoa/jas/seqrep. SUPPLEMENTARY INFORMATION: Supplementary material is available at http://www.dq.fct.unl.pt/qoa/jas/seqrep/bioinf2002

Algorithms↗

Prediction of 1H NMR chemical shifts using neural networks.

Counterpropagation neural networks were applied to the fast prediction of 1H NMR chemical shifts of CHn groups in organic compounds. The training set consisted of 744 examples of protons that were represented by physicochemical, topological, and geometric descriptors. The selection of descriptors was performed by genetic algorithms, and the models obtained were compared to those containing all the descriptors. The best models yielded very good predictions for an independent prediction set of 259 cases (mean absolute error for whole set, 0.25 ppm; mean absolute error for 90% of cases, 0.19 ppm) and for application cases consisting of four natural products recently described. Some stereochemical effects could be correctly predicted. A useful feature of the system resides in its ability to be retrained with a specific data set of compounds if improved predictions for related structures are required.

Artificial Intelligence↗

Prediction of enantiomeric selectivity in chromatography. Application of conformation-dependent and conformation-independent descriptors of molecular chirality.

In order to process molecular chirality by computational methods and to obtain predictions for properties that are influenced by chirality, a fixed-length conformation-dependent chirality code is introduced. The code consists of a set of molecular descriptors representing the chirality of a 3D molecular structure. It includes information about molecular geometry and atomic properties, and can distinguish between enantiomers, even if chirality does not result from chiral centers. The new molecular transform was applied to two datasets of chiral compounds, each of them containing pairs of enantiomers that had been separated by chiral chromatography. The elution order within each pair of isomers was predicted by means of Kohonen neural networks (NN) using the chirality codes as input. A previously described conformation-independent chirality code was also applied and the results were compared. In both applications clustering of the two classes of enantiomers (first eluted and last eluted enantiomers) could be successfully achieved by NN and accurate predictions could be obtained for independent test sets. The chirality code described here has a potential for a broad range of applications from stereoselective reactions to analytical chemistry and to the study of biological activity of chiral compounds.

Amino Acids↗

Prediction of enantiomeric excess in a combinatorial library of catalytic enantioselective reactions.

A quantitative structure-enantioselectivity relationship was established for a combinatorial library of enantioselective reactions performed by addition of diethyl zinc to benzaldehyde. Chiral catalysts and additives were encoded by their chirality codes and presented as input to neural networks. The networks were trained to predict the enantiomeric excess. With independent test sets, predictions of enantiomeric excess could be made with an average error as low as 6% ee. Multilinear regression, perceptrons, and support vector machines were also evaluated as modeling tools. The method is of interest for the computer-aided design of combinatorial libraries involving chiral compounds or enantioselective reactions. This is the first example of a quantitative structure-property relationship based on chirality codes.

Benzaldehydes↗

Chirality codes and molecular structure.

Some time ago a structure-descriptor, named "chirality code", was put forward [J. Chem. Inf. Comput. Sci. 2001, 41, 369-375], aimed at distinguishing between enantiomers. The chirality code is a sequence of (typically 100) numbers, being equal to the value of a certain "chirality function" at equidistant points within a chosen interval. For molecules of moderate size the chirality function has thousands of peaks (maxima and minima), one for each quartet of atoms. Therefore it looks as if the chirality code cannot provide a faithful representation of the chirality function and thus a faithful representation of the molecular structure. We now show that functional groups present in the molecule result in clusters of near-lying and partially overlapping peaks, whose position in the chirality code is characteristic for the particular functional group. This enables a sound structural interpretation of the chirality code.

Journal Article↗

Structure-based predictions of 1H NMR chemical shifts using feed-forward neural networks.

Feed-forward neural networks were trained for the general prediction of 1H NMR chemical shifts of CH(n) protons in organic compounds in CDCl3. The training set consisted of 744 1H NMR chemical shifts from 120 molecular structures. The method was optimized in terms of selected proton descriptors (selection of variables), the number of hidden neurons, and integration of different networks in ensembles. Predictions were obtained for an independent test set of 952 cases with a mean average error of 0.29 ppm (0.20 ppm for 90% of the cases). The results were significantly better than those obtained with counterpropagation neural networks.

Hydrogen↗

The impact of available experimental data on the prediction of 1H NMR chemical shifts by neural networks.

Two different ways were explored to incorporate new available experimental data into previously trained ensembles of feed-forward neural networks, for the structure-based prediction of (1)H NMR chemical shifts of organic compounds. One approach used the new data as the memory of an associative neural network (ASNN) system. For an independent prediction set of 952 cases, a mean average error of 0.19 ppm was achieved (0.13 ppm for 90% of the cases). This approach advantageously avoids retraining the networks, and the predictions compared favorably with those obtained by available commercial software packages. Excellent predictions could also be achieved by retraining the networks with the new data, but only if the training sets were selected so as to be balanced or if the retraining started with the weights of the previously trained networks.

Magnetic Resonance Spectroscopy↗

Structure-based classification of chemical reactions without assignment of reaction centers.

The automatic classification of chemical reactions is of high importance for the analysis of reaction databases, reaction retrieval, reaction prediction, or synthesis planning. In this work, the classification of photochemical reactions was investigated with no explicit assignment of the reacting centers. Classifications were explored with Random Forests or Kohonen neural networks in three different situations, using different levels of information: (a) pairs of reactants were classified according to the type of reaction they produce, (b) products were classified according to the type of reaction from which they can be synthesized, and (c) reactions were classified from the difference between the descriptors of the product and the descriptors of the reactants. In all cases molecular maps of atom-level properties (MOLMAPs) were used as descriptors. They are generated by a self-organizing map and encode physicochemical properties of the bonds available in a molecule. Correct classification could be achieved for approximately 90% of the 78 reactions in an independent test set.

Journal Article↗

Physicochemical stereodescriptors of atomic chiral centers.

Physicochemical atomic stereodescriptors (PAS) were implemented that represent the chirality of an atomic chiral center on the basis of empirical physicochemical properties of the ligands. The ligands are ranked according to a specific property, and the chiral center takes an S/R-like descriptor relative to that property. The procedure is performed for a series of properties, yielding a chirality profile. Application of the PAS descriptors to the prediction of enantioselectivity in chemical reactions, from the molecular structures, is illustrated here. The relationship between the molecular structures, represented by the PAS descriptors, and the enantioselectivity was learned by neural networks, decision trees, or random forests. In a first application, a data set was employed with chiral amino alcohols that enantioselectively catalyze the addition of diethylzinc to benzaldehyde. Prediction of the major enantiomer obtained in the reaction, from the molecular structure of the catalyst, was achieved with accuracy up to 90%. The second application investigated the enantiopreference of Pseudomonas cepacia lipase (PCL) toward primary alcohols. The learned models could make correct predictions about the preferred enantiomer, from the molecular structure of the substrate, in up to 93% of the cases. These included substrates with and without O-atoms bonded to the chiral center. The properties automatically selected to build the models can give indications on the relevant factors guiding the observed chemical behavior.

Alcohols↗