PubMed Health⌕ Search

Biomedical subjects

A Tropsha

Publications and source records attributed to A Tropsha.

At least 19 recordsLinked to original sources

Accurate classification of protein structural families using coherent subgraph analysis.

Protein structural annotation and classification is an important problem in bioinformatics. We report on the development of an efficient subgraph mining technique and its application to finding characteristic substructural patterns within protein structural families. In our method, protein structures are represented by graphs where the nodes are residues and the edges connect residues found within certain distance from each other. Application of subgraph mining to proteins is challenging for a number reasons: (1) protein graphs are large and complex, (2) current protein databases are large and continue to grow rapidly, and (3) only a small fraction of the frequent subgraphs among the huge pool of all possible subgraphs could be significant in the context of protein classification. To address these challenges, we have developed an information theoretic model called coherent subgraph mining. From information theory, the entropy of a random variable X measures the information content carried by X and the Mutual Information (MI) between two random variables X and Y measures the correlation between X and Y. We define a subgraph X as coherent if it is strongly correlated with every sufficiently large sub-subgraph Y embedded in it. Based on the MI metric, we have designed a search scheme that only reports coherent subgraphs. To determine the significance of coherent protein subgraphs, we have conducted an experimental study in which all coherent subgraphs were identified in several protein structural families annotated in the SCOP database (Murzin et al, 1995). The Support Vector Machine algorithm was used to classify proteins from different families under the binary classification scheme. We find that this approach identifies spatial motifs unique to individual SCOP families and affords excellent discrimination between families.

Algorithms↗

Four-body potentials reveal protein-specific correlations to stability changes caused by hydrophobic core mutations.

Mutational experiments show how changes in the hydrophobic cores of proteins affect their stabilities. Here, we estimate these effects computationally, using four-body likelihood potentials obtained by simplicial neighborhood analysis of protein packing (SNAPP). In this procedure, the volume of a known protein structure is tiled with tetrahedra having the center of mass of one amino acid side-chain at each vertex. Log-likelihoods are computed for the 8855 possible tetrahedra with equivalent compositions from structural databases and amino acid frequencies. The sum of these four-body potentials for tetrahedra present in a given protein yields the SNAPP score. Mutations change this sum by changing the compositions of tetrahedra containing the mutated residue and their related potentials. Linear correlation coefficients between experimental mutational stability changes, Delta(DeltaG(unfold)), and those based on SNAPP scoring range from 0.70 to 0.94 for hydrophobic core mutations in five different proteins. Accurate predictions for the effects of hydrophobic core mutations can therefore be obtained by virtual mutagenesis, based on changes to the total SNAPP likelihood potential. Significantly, slopes of the relation between Delta(DeltaG(unfold)) and DeltaSNAPP for different proteins are statistically distinct, and we show that these protein-specific effects can be estimated using the average SNAPP score per residue, which is readily derived from the analysis itself. This result enhances the predictive value of statistical potentials and supports previous suggestions that "comparable" mutations in different proteins may lead to different Delta(DeltaG(unfold)) values because of differences in their flexibility and/or conformational entropy.

Amino Acid Substitution↗

Lattice protein folding with two and four-body statistical potentials.

The cooperative folding of proteins implies a description by multibody potentials. Such multibody potentials can be generalized from common two-body statistical potentials through a relation to probability distributions of residue clusters via the Boltzmann condition. In this exploratory study, we compare a four-body statistical potential, defined by the Delaunay tessellation of protein structures, to the Miyazawa-Jernigan (MJ) potential for protein structure prediction, using a lattice chain growth algorithm. We use the four-body potential as a discriminatory function for conformational ensembles generated with the MJ potential and examine performance on a set of 22 proteins of 30-76 residues in length. We find that the four-body potential yields comparable results to the two-body MJ potential, namely, an average coordinate root-mean-square deviation (cRMSD) value of 8 A for the lowest energy configurations of all-alpha proteins, and somewhat poorer cRMSD values for other protein classes. For both two and four-body potentials, superpositions of some predicted and native structures show a rough overall agreement. Formulating the four-body potential using larger data sets and direct, but costly, generation of conformational ensembles with multibody potentials may offer further improvements. Proteins 2001;43:161-174.

Algorithms↗

Accurate prediction of the bound conformation of galanthamine in the active site of Torpedo californica acetylcholinesterase using molecular docking.

The alkaloid (-)-galanthamine is known to produce significant improvement of cognitive performances in patients with the Alzheimer's disease. Its mechanism of action involves competitive and reversible inhibition of acetylcholinesterase (AChE). Herein, we correctly predict the orientation and conformation of the galanthamine molecule in the active site of AChE from Torpedo californica (TcAChE) using a combination of rigid docking and flexible geometry optimization with a molecular mechanics force field. The quality of the predicted model is remarkable, as indicated by the value of the RMS deviation of approximately 0.5A when compared with the crystal structure of the TcAChE-galanthamine complex. A molecular model of the complex between TcAChE and a galanthamine derivative, SPH1107, with a long chain substituent on the nitrogen has been generated as well. The side chain of this ligand is predicted to extend along the enzyme active site gorge from the anionic subsite, at the bottom, to the peripheral anionic site, at the top. The docking procedure described in this paper can be applied to produce models of ligand-receptor complexes for AChE and other macromolecular targets of drug design.

Acetylcholinesterase↗

Identification of the descriptor pharmacophores using variable selection QSAR: applications to database mining.

The pharmacophore concept is central to the rational drug design and discovery process. Traditionally, a pharmacophore is defined as a specific three-dimensional (3D) arrangement of chemical functional groups found in active molecules, which are characteristic of a certain pharmacological class of compounds. Herein, by analogy with 3D pharmacophores, a more general concept of descriptor pharmacophore is introduced. The descriptor pharmacophores are defined by the means of variable selection QSAR as a subset of molecular descriptors that afford the most statistically significant structure-activity correlation. The two variable selection QSAR methods developed in this laboratory are discussed; these include Genetic Algorithms--Partial Least Squares (GA-PLS) and K-Nearest Neighbors (KNN). Both methods employ multiple topological descriptors of chemical structures such as molecular connectivity indices or atom pairs (AP), and stochastic optimization algorithms to achieve a robust QSAR model, which is characterized by the highest value of cross-validated R2 (q2). By default, the descriptor pharmacophore represents an invariant selection of descriptor types however, descriptor values are generally different for different molecules. We demonstrate that chemical similarity searches using descriptor pharmacophores as opposed to using all descriptors afford more efficient mining of chemical databases or virtual libraries to discover compounds with a desired biological activity.

Algorithms↗

Antitumor agents. 199. Three-dimensional quantitative structure-activity relationship study of the colchicine binding site ligands using comparative molecular field analysis.

Inhibitors of tubulin polymerization interacting at the colchicine binding site are potential anticancer agents. We have been involved in the synthesis of a number of colchicine site agents, such as thiocolchicinoids and allocolchicinoids, which are colchicine analogues, and 2-phenyl-quinolones and 2-aryl-naphthyridinones, which are the amino analogues of cytotoxic antimitotic flavonoids. The most cytotoxic of the latter compounds strongly inhibit binding of radiolabeled colchicine to tubulin, and these agents therefore probably bind in the colchicine site of tubulin. We have applied conventional CoMFA and q(2)-GRS CoMFA to identify the essential structural requirements for increasing the ability of these compounds to form tubulin complexes. The CoMFA model for the training set of 51 compounds yielded cross-validated R(2) (q(2)) values of 0.637 for conventional CoMFA and 0.692 for q(2)-GRS CoMFA. The predictive power of this model was confirmed by successful activity prediction for a test set of 53 compounds with known potencies as inhibitors of tubulin polymerization. The activities of 88% of the compounds were predicted with absolute value of residuals of less than 0.5. The predictive q(2) values were 0.546 for conventional CoMFA and 0.426 for q(2)-GRS CoMFA. The conventional CoMFA model with the highest predictive q(2) (0.546) was analyzed in detail in terms of underlying structure-activity relationships.

Antineoplastic Agents↗

Novel variable selection quantitative structure--property relationship approach based on the k-nearest-neighbor principle

A novel automated variable selection quantitative structure-activity relationship (QSAR) method, based on the kappa-nearest neighbor principle (kNN-QSAR) has been developed. The kNN-QSAR method explores formally the active analogue approach, which implies that similar compounds display similar profiles of pharmacological activities. The activity of each compound is predicted as the average activity of K most chemically similar compounds from the data set. The robustness of a QSAR model is characterized by the value of cross-validated R2 (q2) using the leave-one-out cross-validation method. The chemical structures are characterized by multiple topological descriptors such as molecular connectivity indices or atom pairs. The chemical similarity is evaluated by Euclidean distances between compounds in multidimensional descriptor space, and the optimal subset of descriptors is selected using simulated annealing as a stochastic optimization algorithm. The application of the kNN-QSAR method to 58 estrogen receptor ligands as well as to several other groups of pharmacologically active compounds yielded QSAR models with q2 values of 0.6 or higher. Due to its relative simplicity, high degree of automation, nonlinear nature, and computational efficiency, this method could be applied routinely to a large variety of experimental data sets.

Journal Article↗

The "random-coil" state of proteins: comparison of database statistics and molecular simulations.

This study presents a comparison of two models of the random-coil state, one based on statistical distributions from the structural database and the other based on molecular dynamics simulations. The database model relies on the assumption that the random- or statistical-coil state of a particular residue can be described by its conformational distribution in a sufficiently diverse subset of protein structures. The molecular dynamics model is based on distributions from molecular simulations carried out on "dipeptide" models (single residues with N-terminal acetyl and C-terminal N'-methyl amide blocking groups). A comparison of the two models for the residues Ala, Asn, Asp, Gly, and Val indicates that the database distributions are greatly influenced by long-range interactions and dominated by specific recognizable elements of protein structure. In contrast, the limited structural scope of the dipeptide models presents the extreme case of a peptide under the influence of only short-range interactions. The models were evaluated by a comparison of scalar coupling constants calculated from the conformational distributions and compared with experimentally values determined for unstructured peptides. Although the models gave different distributions, there was similar agreement with experiment. This comparison emphasizes the differences and limitations in each model and highlights the difficulty in presenting an accurate picture of the random-coil state. Proteins 1999;36:407- 418.

Computer Simulation↗

Quantitative structure-activity relationship modeling of dopamine D(1) antagonists using comparative molecular field analysis, genetic algorithms-partial least-squares, and K nearest neighbor methods.

Several quantitative structure-activity relationship (QSAR) methods were applied to 29 chemically diverse D(1) dopamine antagonists. In addition to conventional 3D comparative molecular field analysis (CoMFA), cross-validated R(2) guided region selection (q(2)-GRS) CoMFA (see ref 1) was employed, as were two novel variable selection QSAR methods recently developed in one of our laboratories. These latter methods included genetic algorithm-partial least squares (GA-PLS) and K nearest neighbor (KNN) procedures (see refs 2-4), which utilize 2D topological descriptors of chemical structures. Each QSAR approach resulted in a highly predictive model, with cross-validated R(2) (q(2)) values of 0.57 for CoMFA, 0.54 for q(2)-GRS, 0.73 for GA-PLS, and 0.79 for KNN. The success of all of the QSAR methods indicates the presence of an intrinsic structure-activity relationship in this group of compounds and affords more robust design and prediction of biological activities of novel D(1) ligands.

Algorithms↗

Synthesis, evaluation, and comparative molecular field analysis of 1-phenyl-3-amino-1,2,3,4-tetrahydronaphthalenes as ligands for histamine H(1) receptors.

A series of 1-phenyl-3-amino-1,2,3,4-tetrahydronaphthalenes (1-phenyl-3-aminotetralins, PATs) previously was found to modulate tyrosine hydroxylase activity and dopamine synthesis in rodent forebrain through interaction with a binding site labeled by [(3)H]-(-)-(1R,3S)-trans-H(2)-PAT. Recently, we have discovered that PATs also bind with high affinity to the [(3)H]mepyramine-labeled H(1) receptor in rat and guinea pig brain. Here, we report the synthesis and biological evaluation of additional PAT analogues in order to identify differences in binding at these two sites. Further molecular modifications involve the pendant phenyl ring as well as quaternary amine compounds. Comparison of about 38 PAT analogues, 10 structurally diverse H(1) ligands, and several other CNS-active compounds revealed no significant differences in affinity at [(3)H]-(-)-trans-H(2)-PAT sites versus [(3)H]mepyramine-labeled H(1) receptors. These results, together with previous autoradiographic brain receptor-mapping studies that indicate similar distribution of [(3)H]-(-)-trans-H(2)-PAT sites and [(3)H]mepyramine-labeled H(1) receptors, suggest that both radioligands label the same histamine H(1) receptors in rodent brain. We also report a revision of our previous comparative molecular field analysis (CoMFA) study of the PAT ligands that yields a highly predictive model for 66 compounds with a cross-validated R(2) (q(2)) value of 0.67. This model will be useful for the prediction of high-affinity ligands at radiolabeled H(1) receptors in mammalian brain.

Animals↗

Molecular cloning and characterization of an invertebrate cellular retinoic acid binding protein.

We have cloned a cDNA and gene from the tobacco hornworm, Manduca sexta, which is related to the vertebrate cellular retinoic acid binding proteins (CRABPs). CRABPs are members of the superfamily of lipid binding proteins (LBPs) and are thought to mediate the effects of retinoic acid (RA) on morphogenesis, differentiation, and homeostasis. This discovery of a Manduca sexta CRABP (msCRABP) demonstrates the presence of a CRABP in invertebrates. Compared with bovine/murine CRABP I, the deduced amino acid sequence of msCRABP is 71% homologous overall and 88% homologous for the ligand binding pocket. The genomic organization of msCRABP is conserved with other CRABP family members and the larger LBP superfamily. Importantly, the promoter region contains a motif that resembles an RA response element characteristic of the promoter region of most CRABPs analyzed. Three-dimensional molecular modeling based on postulated structural homology with bovine/murine CRABP I shows msCRABP has a ligand binding pocket that can accommodate RA. The existence of an invertebrate CRABP has significant evolutionary implications, suggesting CRABPs appeared during the evolution of the LBP superfamily well before vertebrate/invertebrate divergence, instead of much later in evolution in selected vertebrates.

Amino Acid Sequence↗

Focus-2D: a new approach to the design of targeted combinatorial chemical libraries.

A strategy for rational design of targeted combinatorial libraries is described. The aim of this approach is to select a subset of available building blocks for the library synthesis that are most likely to be present in the active compounds. Building blocks that are used in the underlying combinatorial chemical reaction are randomly assembled to produce virtual combinatorial library compounds, which are represented by various chemical descriptors. Stochastic algorithms (simulated annealing, genetic algorithms, neural net methods) are used to search the potentially large structural space of virtual chemical libraries in order to identify compounds similar to lead compound(-s). The selection of a virtual molecule as a candidate for the targeted library is based either on its chemical similarity to a biologically active probe or on its biological activity predicted from a pre-constructed QSAR equation. Frequency analysis of building block composition of the selected virtual compounds identifies building blocks that can be used in combinatorial synthesis of chemical libraries with high similarity to the lead compound(-s). This method is applied to rational design of the library with bradykinin potentiating activity. Twenty eight bradykinin potentiating pentapeptides were used as a training set for the development of a QSAR equation, and, alternatively, two active pentapeptides, VEWAK and VKWAP, were used as probe molecules. In each case, the frequency distribution of amino acids in the top 100 peptides suggested by the method resembles the frequency distribution of amino acids found in the active peptides. The results obtained after GA optimization also compared favorably with those obtained by the exhaustive analysis of all possible 3.2 millions pentapeptides.

Algorithms↗

A new approach to protein fold recognition based on Delaunay tessellation of protein structure.

We propose new algorithms for sequence-structure compatibility (fold recognition) searches in multi-dimensional sequence-structure space. Individual amino acid residues in protein structures are represented by their C alpha atoms; thus each protein is described as a collection of points in three-dimensional space. Delaunay tessellation of a protein generates an aggregate of space-filling, irregular tetrahedra, or Delaunay simplices. Statistical analysis of quadruplet residue compositions of all Delaunay simplices in a representative dataset of protein structures leads to a novel four body contact residue potential expressed as log likelihood factor q. The q factors are calculated for native 20 letter amino acid alphabet and several reduced alphabets. Two sequence-structure compatibility functions are computed as (i) the sum of q factors for all Delaunay simplices in a given protein, or (ii) 3D-1D Delaunay tessellation profiles where the individual residue profile value is calculated as the sum of q factors for all simplices that share this vertex residue. Both threading functions have been implemented in structure-recognizes-sequence and sequence-recognizes-structure protocols for protein fold recognition. We find that both profile and total score based threading functions can distinguish both the native fold from incorrect folds for a sequence, and the native sequence from non-native sequences for a fold.

Amino Acid Sequence↗

Structure-based alignment and comparative molecular field analysis of acetylcholinesterase inhibitors.

The method of comparative molecular field analysis (CoMFA) was used to develop quantitative structure-activity relationships for physostigmine, 9-amino-1,2,3,4-tetrahydroacridine (THA), edrophonium (EDR), and other structurally diverse inhibitors of acetylcholinesterase (AChE). The availability of the crystal structures of enzyme/inhibitor complexes (EDR/AChE, THA/AChE, and decamethonium (DCM)/AChE) (Harel, M.; et al. Quaternary ligand binding to aromatic residues in the active-site gorge of acetylcholinesterase. Proc. Natl. Acad. Sci. U.S.A. 1993, 90, 9031-9035) provided information regarding not only the active conformation of the inhibitors but also the relative mutual orientation of the inhibitors in the active site of the enzyme. Crystallographic conformations of EDR and THA were used as templates onto which additional inhibitors were superimposed. The application of cross-validated R2 guided region selection method, recently developed in this laboratory (Cho, S.J.; Tropsha, A. Cross-Validated R2 Guided Region Selection for Comparative Molecular Field Analysis (CoMFA): A Simple Method to Achieve Consistent Results. J. Med. Chem. 1995, 38, 1060-1066), to 60 AChE inhibitors led to a highly predictive CoMFA model with the q2 of 0.734.

Acetylcholinesterase↗

Molecular simulations of beta-sheet twisting.

Twisted conformations of two- and three-stranded antiparallel beta-sheet models containing alanine, glycine and valine with three or five residues per strand have been studied by molecular dynamics simulations. Free molecular dynamics and free energy simulations have been carried out to characterize the dynamics and energetics of the conformational change from a flat sheet to a twisted sheet. By altering the charges on the model in the free energy simulations, we have been able to analyze the contributions to the twist from electrostatic and van der Waals interactions. We have found that alanine and valine beta-sheets prefer conformations with a right-handed twist. In contrast, model glycine sheets do not have a pronounced preference to twist. Single beta-strands are found to be easily twisted, but to not have a strong preference for twisted conformations. Hence, the driving forces for the right-handed twist of beta-sheets must come principally from interactions between strands. These results disagree with several previous theoretical studies and constitute a different paradigm of the origin of beta-sheet twist observed in proteins.

Alanine↗

Antitumor agents. 163. Three-dimensional quantitative structure-activity relationship study of 4'-O-demethylepipodophyllotoxin analogs using the modified CoMFA/q2-GRS approach.

Analogs of 4'O-demethylepipodophyllotoxin are considered as potential anticancer agents. We have applied comparative molecular field analysis (CoMFA) and a novel CoMFA/q2-GRS technique recently developed in our group to identify the essential structural requirements for increasing the ability of these compounds to form cellular protein-DNA complex. In addition, a new method to incorporate different types of probe atoms as part of q2-GRS routine has been developed. The best final model with 101 compounds using a combination of four different sets of probe atoms and charges [C (sp3, +1), C (sp3, 0), H (+1), and O (sp3, -1)] yielded a q 2 of 0.584 and the standard error of prediction of 0.660 at 5 principal components. The steric and electrostatic contour plots of the final model were compared with the DNA phosphate backbone environment of the DNA-4'O-demethylepipodophylltoxin analog complex, which was generated using the X-ray structure of the DNA-nogalamycin complex. The comparison reveals that the CoMFA steric and electrostatic fields are compatible with stereochemical properties of the DNA backbone. The results obtained from this study shall guide our future synthetic efforts.

Antineoplastic Agents↗

Conformational analysis of D1 dopamine receptor agonists: pharmacophore assessment and receptor mapping.

Compute-aided conformational analysis was used to characterize the agonist pharmacophore for D1 dopamine receptor recognition and activation. Dihydrexidine (DHX), a high-affinity full agonist with limited conformational flexibility, served as a structural template that aided in determining a molecular geometry that would be common for other more flexible, biologically active agonists. The intrinsic activity of the drugs at D1 receptors was assessed by their ability to stimulate adenylate cyclase activity in rat striatal homogenates (the accepted measure of D1 receptor activation). In addition, affinity data on 12 agonists including six purported full agonists (dopamine, dihydrexidine, SKF89626, SKF82958, A70108, and A77636), as well as six less efficacious structural analogs, were obtained from D1 dopamine radioreceptor-binding assays. The active analog approach to pharmacophore building was applied as implemented in the SYBYL software package. Conformational analysis and molecular mechanics calculations were used to determine the lowest energy conformation of the active analogs (i.e., full agonists), as well as the conformations of each compound that displayed a common pharmacophoric geometry. It is hypothesized that DHX and other full agonists may share a D1 pharmacophore made up of two hydroxy groups, the nitrogen atom (ca. 7 A from the oxygen of m-hydroxyl) and the accessory ring system characterized by the angle between its plane and that of the catechol ring (except for dopamine and A77636). For all full agonists (DHX, SKF89626, SKF82958, A70108, A77636, and dopamine), the energy difference between the lowest energy conformer and those that displayed a common pharmacophore geometry was relatively small (< 5 kcal/mol). The pharmacophoric conformations of the full agonists were also used to infer the shape of the receptor binding site. Based on the union of the van der Waals density maps of the active analogs, the excluded receptor volume was calculated. Various inactive analogs (partial agonists with D1 K0.5 > 300 nM) subsequently were used to define the receptor essential volume (i.e., sterically intolerable receptor regions). These volumes, together with the pharmacophore results, were integrated into a three-dimensional model estimating the D1 receptor active site topography.

Adenylyl Cyclases↗