PubMed Health⌕ Search

Biomedical subjects

Tingjun Hou

Publications and source records attributed to Tingjun Hou.

13 recordsLinked to original sources

Computational analysis and prediction of the binding motif and protein interacting partners of the Abl SH3 domain.

Protein-protein interactions, particularly weak and transient ones, are often mediated by peptide recognition domains, such as Src Homology 2 and 3 (SH2 and SH3) domains, which bind to specific sequence and structural motifs. It is important but challenging to determine the binding specificity of these domains accurately and to predict their physiological interacting partners. In this study, the interactions between 35 peptide ligands (15 binders and 20 non-binders) and the Abl SH3 domain were analyzed using molecular dynamics simulation and the Molecular Mechanics/Poisson-Boltzmann Solvent Area method. The calculated binding free energies correlated well with the rank order of the binding peptides and clearly distinguished binders from non-binders. Free energy component analysis revealed that the van der Waals interactions dictate the binding strength of peptides, whereas the binding specificity is determined by the electrostatic interaction and the polar contribution of desolvation. The binding motif of the Abl SH3 domain was then determined by a virtual mutagenesis method, which mutates the residue at each position of the template peptide relative to all other 19 amino acids and calculates the binding free energy difference between the template and the mutated peptides using the Molecular Mechanics/Poisson-Boltzmann Solvent Area method. A single position mutation free energy profile was thus established and used as a scoring matrix to search peptides recognized by the Abl SH3 domain in the human genome. Our approach successfully picked ten out of 13 experimentally determined binding partners of the Abl SH3 domain among the top 600 candidates from the 218,540 decapeptides with the PXXP motif in the SWISS-PROT database. We expect that this physical-principle based method can be applied to other protein domains as well.

Amino Acid Motifs↗

Prediction of binding sites of peptide recognition domains: an application on Grb2 and SAP SH2 domains.

Determination of the binding motif and identification of interaction partners of the modular domains such as SH2 domains can enhance our understanding of the regulatory mechanism of protein-protein interactions. We propose here a new computational method to achieve this goal by integrating the orthogonal information obtained from binding free energy estimation and peptide sequence analysis. We performed a proof-of-concept study on the SH2 domains of SAP and Grb2 proteins. The method involves the following steps: (1) estimating the binding free energy of a set of randomly selected peptides along with a sample of known binders; (2) clustering all these peptides using sequence and energy characteristics; (3) extracting a sequence motif, which is represented by a hidden Markov model (HMM), from the cluster of peptides containing the sample of known binders; and (4) scanning the human proteome to identify binding sites of the domain. The binding motifs of the SAP and Grb2 SH2 domains derived by the method agree well with those determined through experimental studies. Using the derived binding motifs, we have predicted new possible interaction partners for the Grb2 and SAP SH2 domains as well as possible interaction sites for interaction partners already known. We also suggested novel roles for the proteins by reviewing their top interaction candidates.

Amino Acid Sequence↗

Prediction of binding affinities between the human amphiphysin-1 SH3 domain and its peptide ligands using homology modeling, molecular dynamics and molecular field analysis.

The SH3 domain of the human protein amphiphysin-1, which plays important roles in clathrin-mediated endocytosis, actin function and signaling transduction, can recognize peptide motif PXRPXR (X is any amino acid) with high affinity and specificity. We have constructed a complex structure of the amphiphysin-1 SH3 domain and a high-affinity peptide ligand PLPRRPPRA using homology modeling and molecular docking, which was optimized by molecular dynamics (MD). Three-dimensional quantitative structure-affinity relationship (3D-QSAR) analyses on the 200 peptides with known binding affinities to the amphiphysin-1 SH3 domain was then performed using comparative molecular field analysis (CoMFA) and comparative molecular similarity indices analysis (CoMSIA). The best CoMSIA model showed promising predictive power, giving good predictions for about 95% of the peptides in the test set (absolute prediction errors less than 1.0). It was used to validate peptide-SH3 binding structure and provide insight into the structural requirements for binding of peptides to SH3 domains. Finally, MD simulations were performed to analyze the interaction between the SH3 domain and another peptide GFPRRPPPRG that contains with the PXRPXsR (s represents residues with small side chains) motif. MD simulations demonstrated that the binding conformation of GFPRRPPPRG is quite different from that of PLPRRPPRAA especially the four residues at the C terminal, which may explain why the CoMSIA model cannot give good predictions on the peptides of the PXRPXsR motif. Because of its efficiency and predictive power, the 3D-QSAR model can be used as a scoring filter for predicting peptide sequences bound to SH3 domains.

Amino Acid Sequence↗

Recent advances in computational prediction of drug absorption and permeability in drug discovery.

Approximately 40%-60% of developing drugs failed during the clinical trials because of ADME/Tox deficiencies. Virtual screening should not be restricted to optimize binding affinity and improve selectivity; and the pharmacokinetic properties should also be included as important filters in virtual screening. Here, the current development in theoretical models to predict drug absorption-related properties, such as intestinal absorption, Caco-2 permeability, and blood-brain partitioning are reviewed. The important physicochemical properties used in the prediction of drug absorption, and the relevance of predictive models in the evaluation of passive drug absorption are discussed. Recent developments in the prediction of drug absorption, especially with the application of new machine learning methods and newly developed software are also discussed. Future directions for research are outlined.

Computer Simulation↗

Calculation of the Maxwell stress tensor and the Poisson-Boltzmann force on a solvated molecular surface using hypersingular boundary integrals.

The electrostatic interaction among molecules solvated in ionic solution is governed by the Poisson-Boltzmann equation (PBE). Here the hypersingular integral technique is used in a boundary element method (BEM) for the three-dimensional (3D) linear PBE to calculate the Maxwell stress tensor on the solvated molecular surface, and then the PB forces and torques can be obtained from the stress tensor. Compared with the variational method (also in a BEM frame) that we proposed recently, this method provides an even more efficient way to calculate the full intermolecular electrostatic interaction force, especially for macromolecular systems. Thus, it may be more suitable for the application of Brownian dynamics methods to study the dynamics of protein/protein docking as well as the assembly of large 3D architectures involving many diffusing subunits. The method has been tested on two simple cases to demonstrate its reliability and efficiency, and also compared with our previous variational method used in BEM.

Algorithms↗

An extended aqueous solvation model based on atom-weighted solvent accessible surface areas: SAWSA v2.0 model.

A new method is proposed for calculating aqueous solvation free energy based on atom-weighted solvent accessible surface areas. The method, SAWSA v2.0, gives the aqueous solvation free energy by summing the contributions of component atoms and a correction factor. We applied two different sets of atom typing rules and fitting processes for small organic molecules and proteins, respectively. For small organic molecules, the model classified the atoms in organic molecules into 65 basic types and additionally. For small organic molecules we proposed a correction factor of "hydrophobic carbon" to account for the aggregation of hydrocarbons and compounds with long hydrophobic aliphatic chains. The contributions for each atom type and correction factor were derived by multivariate regression analysis of 379 neutral molecules and 39 ions with known experimental aqueous solvation free energies. Based on the new atom typing rules, the correlation coefficient (r) for fitting the whole neutral organic molecules is 0.984, and the absolute mean error is 0.40 kcal mol(-1), which is much better than those of the model proposed by Wang et al. and the SAWSA model previously proposed by us. Furthermore, the SAWSA v2.0 model was compared with the simple atom-additive model based on the number of atom types (NA). The calculated results show that for small organic molecules, the predictions from the SAWSA v2.0 model are slightly better than those from the atom-additive model based on NA. However, for macromolecules such as proteins, due to the connection between their molecular conformation and their molecular surface area, the atom-additive model based on the number of atom types has little predictive power. In order to investigate the predictive power of our model, a systematic comparison was performed on seven solvation models including SAWSA v2.0, GB/SA_1, GB/SA_2, PB/SA_1, PB/SA_2, AM1/SM5.2R and SM5.0R. The results showed that for organic molecules the SAWSA v2.0 model is better than the other six solvation models. For proteins, the model classified the atoms into 20 basic types and the predicted aqueous free energies of solvation by PB/SA were used for fitting. The solvation model based on the new parameters was employed to predict the solvation free energies of 38 proteins. The predicted values from our model were in good agreement with those from the PB/SA model and were much better than those given by the other four models developed for proteins.

Chemistry, Organic↗

Recent development and application of virtual screening in drug discovery: an overview.

Virtual screening, especially the structure-based virtual screening, has emerged as a reliable, cost-effective and time-saving technique for the discovery of lead compounds. Here, the basic ideas and computational tools for virtual screening have been briefly introduced, and emphasis is placed on aspects of recent development of docking-based virtual screening, scoring functions in molecular docking and ADME/Tox-based virtual screening in the past three years (2000 to 2003). Moreover, successful examples are provided to further demonstrate the effectiveness of virtual screening in drug discovery.

Combinatorial Chemistry Techniques↗

ADME evaluation in drug discovery. 1. Applications of genetic algorithms to the prediction of blood-brain partitioning of a large set of drugs.

In this study, the relationships between the brain-blood concentration ratio of 96 structurally diverse compounds with a large number of structurally derived descriptors were investigated. The linear models were based on molecular descriptors that can be calculated for any compound simply from a knowledge of its molecular structure. The linear correlation coefficients of the models were optimized by genetic algorithms (GAs), and the descriptors used in the linear models were automatically selected from 27 structurally derived descriptors. The GA optimizations resulted in a group of linear models with three or four molecular descriptors with good statistical significance. The change of descriptor use as the evolution proceeds demonstrates that the octane/water partition coefficient and the partial negative solvent-accessible surface area multiplied by the negative charge are crucial to brain-blood barrier permeability. Moreover, we found that the predictions using multiple QSPR models from GA optimization gave quite good results in spite of the diversity of structures, which was better than the predictions using the best single model. The predictions for the two external sets with 37 diverse compounds using multiple QSPR models indicate that the best linear models with four descriptors are sufficiently effective for predictive use. Considering the ease of computation of the descriptors, the linear models may be used as general utilities to screen the blood-brain barrier partitioning of drugs in a high-throughput fashion.

Algorithms↗

Molecular docking studies of a group of hydroxamate inhibitors with gelatinase-A by molecular dynamics.

We have performed docking and molecular dynamics simulations of hydroxamates complexed with human gelatinase-A (MMP-2) to gain insight into the structural and energetic preferences of these inhibitors. The study was conducted on a selected set of eleven compounds with variation in structure and activity. Molecular dynamics simulations were performed at 300 K for 100 ps with equilibration for 50 ps. The structural analyses of the trajectories indicate that the coordinate bond interactions, the hydrogen bond interactions, the van der Waals interactions as well as the hydrophobic interactions between ligand and receptor are responsible simultaneously for the preference of inhibition and potency. The ligand hydroxamate group is coordinated to the catalytic zinc ion and form stable hydrogen bonds with the carbonyl oxygen of Gly 162. The P1' group makes extensive van der Waals and hydrophobic contacts with the nonpolar side chains of several residues in the S1' subsite, including Leu 197, Val 198, Leu 218 and Tyr 223. Moreover, four to eight hydrogen bonds between hydroxamates and MMP-2 are formed to stabilize the inhibitors in the active site. Compared with the P2' and P3' groups, the P1' groups of inhibitors are oriented regularly, which is produced by the restrain of the S1' subsite. From the relationship between the length of the nonpolar P1' group and the biological activity, we confirm that MMP-2 has a pocket-like S1' subsite, not a channel-like S1' subsite proposed by Kiyama (Kiyama, R. et al., J. Med. Chem. 42 (1999), 1723). The energetic analyses show that the experimental binding free energies can be well correlated with the interactions between the inhibitors and their environments, which could be used as a simple score function to evaluate the binding affinities for other similar hydroxamates. The validity of the force field parameters and the MD simulations can be fully testified by the satisfactory agreements between the experimental structure-activity relationship and the information from the structural and energetic analyses. The information generated from the predicted complexes should be useful for further work in the area of structure-based design of new compounds.

Binding Sites↗

A 3D structure database of components from Chinese traditional medicinal herbs.

This article described a 3D structure database of components extracted from Chinese Traditional Medicinal (CTM) herbs. It offers not only basic molecular properties and optimized 3D structure of the compounds but also detailed information on their herbal origin, including basic herbal category (e.g. English name, Latin name, and family), effective parts, and clinical effects. An easy to use, interactive GUI browser allows users to perform various searches via complex logical query builder. Combined with the latest network database engine (MySQL), it can achieve excellent performance under both a local network and an Internet environment. We have tested it on the design of inhibitors of NS3-NS4A protease. Results show that the structure database of components extracted from Chinese medicinal herbs can be a rich source in searching the lead compound.

Databases, Factual↗

Mapping the binding site of a large set of quinazoline type EGF-R inhibitors using molecular field analyses and molecular docking studies.

In the current work, three-dimensional QSAR studies for one large set of quinazoline type epidermal growth factor receptor (EGF-R) inhibitors were conducted using two types of molecular field analysis techniques: comparative molecular field analysis (CoMFA) and comparative molecular similarity indices analysis (CoMSIA). These compounds belonging to six different structural classes were randomly divided into a training set of 122 compounds and a test set of 13 compounds. The statistical results showed that the 3D-QSAR models derived from CoMFA were superior to those generated from CoMSIA. The most optimal CoMFA model after region focusing bears significant cross-validated r(2)(cv) of 0.60 and conventional r(2) of 0.92. The predictive power of the best CoMFA model was further validated by the accurate estimation to these compounds in the external test set, and the mean agreement of experimental and predicted log(IC(50)) values of the inhibitors is 0.6 log unit. Separate CoMFA models were conducted to evaluate the influence of different partial charges (Gasteiger-Marsili, Gasteiger-Hückel, MMFF94, ESP-AM1, and MPA-AM1) on the statistical quality of the models. The resulting CoMFA field map provides information on the geometry of the binding site cavity and the relative weights of various properties in different site pockets for each of the substrates considered. Moreover, in the current work, we applied MD simulations combined with MM/PBSA (Molecular mechanics/Possion-Boltzmann Surface Area) to determine the correct binding mode of the best inhibitor for which no ligand-protein crystal structure was present. To proceed, we define the following procedure: three hundred picosecond molecular dynamics simulations were first performed for the four binding modes suggested by DOCK 4.0 and manual docking, and then MM/PBSA was carried out for the collected snapshots. The most favorable binding mode identified by MM/PBSA has a binding free energy about 10 kcal/mol more favorable than the second best one. The most favorable binding mode identified by MM/PBSA can give satisfactory explanation of the SAR data of the studied molecules and is in good agreement with the contour maps of CoMFA. The most favorable binding mode suggests that with the quinazoline-based inhibitor, the N3 atom is hydrogen-bonded to a water molecule which, in turn, interacts with Thr 766, not Thr 830 as proposed by Wissner et al. (J. Med. Chem. 2000, 43, 3244). The predicted complex structure of quinazoline type inhibitor with EGF-R as well as the pharmacophore mapping from CoMFA can interpret the structure activities of the inhibitors well and afford us important information for structure-based drug design.

Binding Sites↗

New Born radii deriving method for Generalized Born model.

Here we report a method to calculate Born radii, an important parameter used in a Generalized Born model. Traditional methods to derive Born radii are mostly based on a complicated formula, while our method is easier and more direct. Atoms are classified according to their atom type, and the Born radii of each type are obtained by fitting to experimental solvation free energy. The SMARTS language is used for the exact definition of atoms types, and Ullmann's subgraph isomorphism algorithm is used to deduce the environment. A generic algorithm is used for the parameter fitting because of its efficiency in searching a huge phase space, and its results are then optimized by using the conjugate gradient method. The final parameter set is fitting from a training set containing 357 molecules and is tested using a test set of 44 small organic molecules, and the average error is 0.58 kcal/mol for 36 neutral molecules and is 1.67 kcal/mol for 8 ions. The model is further tested under organic molecules, biopolymers, and a protein-inhibitor complex and yields reliable results in all these cases. This method can be used to accelerate molecular docking calculations.

Journal Article↗

Some basic data structures and algorithms for chemical generic programming.

Here, we report a template library used for molecular operation, the Molecular Handling Template Library (MHTL). The library includes some generic data structures and generic algorithms, and the two parts are associated with each other by two concepts: Properties and Molecule. The concept Properties describes the interface to access objects' properties, and the concept Molecule describes the minimum requirement for a molecular class. Data structures include seven models of Properties, each using a different method to access properties, and two models of molecular classes. Algorithms include molecular file manipulation subroutines, SMARTS language interpreter and matcher functions, and molecular OpenGL rendering functions.

Journal Article↗