PubMed Health⌕ Search

Biomedical subjects

Richard M Jackson

Publications and source records attributed to Richard M Jackson.

18 recordsLinked to original sources

Identification of the REST regulon reveals extensive transposable element-mediated binding site duplication.

The genome-wide mapping of gene-regulatory motifs remains a major goal that will facilitate the modelling of gene-regulatory networks and their evolution. The repressor element 1 is a long, conserved transcription factor-binding site which recruits the transcriptional repressor REST to numerous neuron-specific target genes. REST plays important roles in multiple biological processes and disease states. To map RE1 sites and target genes, we created a position specific scoring matrix representing the RE1 and used it to search the human and mouse genomes. We identified 1301 and 997 RE1s inhuman and mouse genomes, respectively, of which >40% are novel. By employing an ontological analysis we show that REST target genes are significantly enriched in a number of functional classes. Taking the novel REST target gene CACNA1A as an experimental model, we show that it can be regulated by multiple RE1s of different binding affinities, which are only partially conserved between human and mouse. A novel BLAST methodology indicated that many RE1s belong to closely related families. Most of these sequences are associated with transposable elements, leading us to propose that transposon-mediated duplication and insertion of RE1s has led to the acquisition of novel target genes by REST during evolution.

Animals↗

Delineation and modelling of a nucleolar retention signal in the coronavirus nucleocapsid protein.

Unlike nuclear localization signals, there is no obvious consensus sequence for the targeting of proteins to the nucleolus. The nucleolus is a dynamic subnuclear structure which is crucial to the normal operation of the eukaryotic cell. Studying nucleolar trafficking signals is problematic as many nucleolar retention signals (NoRSs) are part of classical nuclear localization signals (NLSs). In addition, there is no known consensus signal with which to inform a study. The avian infectious bronchitis virus (IBV), coronavirus nucleocapsid (N) protein, localizes to the cytoplasm and the nucleolus. Mutagenesis was used to delineate a novel eight amino acid motif that was necessary and sufficient for nucleolar retention of N protein and colocalize with nucleolin and fibrillarin. Additionally, a classical nuclear export signal (NES) functioned to direct N protein to the cytoplasm. Comparison of the coronavirus NoRSs with known cellular and other viral NoRSs revealed that these motifs have conserved arginine residues. Molecular modelling, using the solution structure of severe acute respiratory (SARS) coronavirus N-protein, revealed that this motif is available for interaction with cellular factors which may mediate nucleolar localization. We hypothesise that the N-protein uses these signals to traffic to and from the nucleolus and the cytoplasm.

Active Transport, Cell Nucleus↗

Predicting protein interaction sites: binding hot-spots in protein-protein and protein-ligand interfaces.

MOTIVATION: Protein assemblies are currently poorly represented in structural databases and their structural elucidation is a key goal in biology. Here we analyse clefts in protein surfaces, likely to correspond to binding 'hot-spots', and rank them according to sequence conservation and simple measures of physical properties including hydrophobicity, desolvation, electrostatic and van der Waals potentials, to predict which are involved in binding in the native complex. RESULTS: The resulting differences between predicting binding-sites at protein-protein and protein-ligand interfaces are striking. There is a high level of prediction accuracy (< or =93%) for protein-ligand interactions, based on the following attributes: van der Waals potential, electrostatic potential, desolvation and surface conservation. Generally, the prediction accuracy for protein-protein interactions is lower, with the exception of enzymes. Our results show that the ease of cleft desolvation is strongly predictive of interfaces and strongly maintained across all classes of protein-binding interface.

Binding Sites↗

SitesBase: a database for structure-based protein-ligand binding site comparisons.

There are many components which govern the function of a protein within a cell. Here, we focus on the molecular recognition of small molecules and the prediction of common recognition by similarity between protein-ligand binding sites. SitesBase is an easily accessible database which is simple to use and holds information about structural similarities between known ligand binding sites found in the Protein Data Bank. These similarities are presented to the wider community enabling full analysis of molecular recognition and potentially protein structure-function relationships. SitesBase is accessible at http://www.bioinformatics.leeds.ac.uk/sb.

Binding Sites↗

Methods for the prediction of protein-ligand binding sites for structure-based drug design and virtual ligand screening.

Structure Based Drug Design (SBDD) is a computational approach to lead discovery that uses the three-dimensional structure of a protein to fit drug-like molecules into a ligand binding site to modulate function. Identifying the location of the binding site is therefore a vital first step in this process, restricting the search space for SBDD or virtual screening studies. The detection and characterisation of functional sites on proteins has increasingly become an area of interest. Structural genomics projects are increasingly yielding protein structures with unknown functions and binding sites. Binding site prediction was pioneered by pocket detection, since the binding site is often found in the largest pocket. More recent methods involve phylogenetic analysis, identifying structural similarity with proteins of known function and identifying regions on the protein surface with a potential for high binding affinity. Binding site prediction has been used in several SBDD projects and has been incorporated into several docking tools. We discuss different methods of ligand binding site prediction, their strengths and weaknesses, and how they have been used in SBDD.

Animals↗

Fold independent structural comparisons of protein-ligand binding sites for exploring functional relationships.

The rapid growth in protein structural data and the emergence of structural genomics projects have increased the need for automatic structure analysis and tools for function prediction. Small molecule recognition is critical to the function of many proteins; therefore, determination of ligand binding site similarity is important for understanding ligand interactions and may allow their functional classification. Here, we present a binding sites database (SitesBase) that given a known protein-ligand binding site allows rapid retrieval of other binding sites with similar structure independent of overall sequence or fold similarity. However, each match is also annotated with sequence similarity and fold information to aid interpretation of structure and functional similarity. Similarity in ligand binding sites can indicate common binding modes and recognition of similar molecules, allowing potential inference of function for an uncharacterised protein or providing additional evidence of common function where sequence or fold similarity is already known. Alternatively, the resource can provide valuable information for detailed studies of molecular recognition including structure-based ligand design and in understanding ligand cross-reactivity. Here, we show examples of atomic similarity between superfamily or more distant fold relatives as well as between seemingly unrelated proteins. Assignment of unclassified proteins to structural superfamiles is also undertaken and in most cases substantiates assignments made using sequence similarity. Correct assignment is also possible where sequence similarity fails to find significant matches, illustrating the potential use of binding site comparisons for newly determined proteins.

Animals↗

Comparison of the ATP binding sites of protein kinases using conformationally diverse bisindolylmaleimides.

The conformation of a bisindolylmaleimide may be controlled by the size of a macrocyclic ring in which it is constrained. A range of techniques were used to demonstrate that the tether controls both the ratio of the two limiting conformers (syn and anti) in solution and the extent of conjugation between the maleimide and indole rings. Screening the conformationally diverse bisindolylmaleimides against a panel of protein kinases allowed their ATP binding sites to be compared using a chemical approach which, like sequence alignment, does not require detailed structural information. This approach lead to the conclusion that several AGC group protein kinases (including PKCalpha, PKCbeta, MSK1, p70 S6K, PDK-1, and MAPKAP-K1alpha) may be best inhibited by bisindolylmaleimides which adopt a compressed approximately C2-symmetric anti conformation; in constrast, GSK3beta may be best inhibited by bisindolylmaleimides whose ground state is a distorted syn conformation. It is concluded that PDK-1, whose structure has been determined by X-ray crystallography, and its mutants, may serve as particularly useful surrogates for the study of PKC inhibitors.

Adenosine Triphosphate↗

Prediction of protein-protein interactions using distant conservation of sequence patterns and structure relationships.

MOTIVATION: Given that association and dissociation of protein molecules is crucial in most biological processes several in silico methods have been recently developed to predict protein-protein interactions. Structural evidence has shown that usually interacting pairs of close homologs (interologs) physically interact in the same way. Moreover, conservation of an interaction depends on the conservation of the interface between interacting partners. In this article we make use of both, structural similarities among domains of known interacting proteins found in the Database of Interacting Proteins (DIP) and conservation of pairs of sequence patches involved in protein-protein interfaces to predict putative protein interaction pairs. RESULTS: We have obtained a large amount of putative protein-protein interaction (approximately 130,000). The list is independent from other techniques both experimental and theoretical. We separated the list of predictions into three sets according to their relationship with known interacting proteins found in DIP. For each set, only a small fraction of the predicted protein pairs could be independently validated by cross checking with the Human Protein Reference Database (HPRD). The fraction of validated protein pairs was always larger than that expected by using random protein pairs. Furthermore, a correlation map of interacting protein pairs was calculated with respect to molecular function, as defined in the Gene Ontology database. It shows good consistency of the predicted interactions with data in the HPRD database. The intersection between the lists of interactions of other methods and ours produces a network of potentially high-confidence interactions.

Algorithms↗

Q-SiteFinder: an energy-based method for the prediction of protein-ligand binding sites.

MOTIVATION: Identifying the location of ligand binding sites on a protein is of fundamental importance for a range of applications including molecular docking, de novo drug design and structural identification and comparison of functional sites. Here, we describe a new method of ligand binding site prediction called Q-SiteFinder. It uses the interaction energy between the protein and a simple van der Waals probe to locate energetically favourable binding sites. Energetically favourable probe sites are clustered according to their spatial proximity and clusters are then ranked according to the sum of interaction energies for sites within each cluster. RESULTS: There is at least one successful prediction in the top three predicted sites in 90% of proteins tested when using Q-SiteFinder. This success rate is higher than that of a commonly used pocket detection algorithm (Pocket-Finder) which uses geometric criteria. Additionally, Q-SiteFinder is twice as effective as Pocket-Finder in generating predicted sites that map accurately onto ligand coordinates. It also generates predicted sites with the lowest average volumes of the methods examined in this study. Unlike pocket detection, the volumes of the predicted sites appear to show relatively low dependence on protein volume and are similar in volume to the ligands they contain. Restricting the size of the pocket is important for reducing the search space required for docking and de novo drug design or site comparison. The method can be applied in structural genomics studies where protein binding sites remain uncharacterized since the 86% success rate for unbound proteins appears to be only slightly lower than that of ligand-bound proteins. AVAILABILITY: Both Q-SiteFinder and Pocket-Finder have been made available online at http://www.bioinformatics.leeds.ac.uk/qsitefinder and http://www.bioinformatics.leeds.ac.uk/pocketfinder

Algorithms↗

Further studies on hepatitis C virus NS5A-SH3 domain interactions: identification of residues critical for binding and implications for viral RNA replication and modulation of cell signalling.

The NS5A protein of hepatitis C virus has been shown to interact with a subset of Src homology 3 (SH3) domain-containing proteins. The molecular mechanisms underlying these observations have not been fully characterized, therefore a previous analysis of NS5A-SH3 domain interactions was extended. By using a semi-quantitative ELISA assay, a hierarchy of binding between various SH3 domains for NS5A was demonstrated. Molecular modelling of a polyproline motif within NS5A (termed PP2.2) bound to the FynSH3 domain predicted that the specificity-determining RT-loop region within the SH3 domain did not interact directly with the PP2.2 motif. However, it was demonstrated that the RT loop did contribute to the specificity of binding, implicating the involvement of other intermolecular contacts between NS5A and SH3 domains. The modelling analysis also predicted a critical role for a conserved arginine located at the C terminus of the PP2.2 motif; this was confirmed experimentally. Finally, it was demonstrated that, in comparison with wild-type replicon cells, inhibition of the transcription factor AP-1, a function previously assigned to NS5A, was not observed in cells harbouring a subgenomic replicon containing a mutation within the PP2.2 motif. However, the ability of the mutated replicon to establish itself within Huh-7 cells was unaffected. The highly conserved nature of the PP2.2 motif within NS5A suggests that functions involving this motif are of importance, but are unlikely to play a role in replication of the viral RNA genome. It is more likely that they play a role in altering the cellular environment to favour viral persistence.

Amino Acid Sequence↗

Identification of critical active-site residues in angiotensin-converting enzyme-2 (ACE2) by site-directed mutagenesis.

Angiotensin-converting enzyme-2 (ACE2) may play an important role in cardiorenal disease and it has also been implicated as a cellular receptor for the severe acute respiratory syndrome (SARS) virus. The ACE2 active-site model and its crystal structure, which was solved recently, highlighted key differences between ACE2 and its counterpart angiotensin-converting enzyme (ACE), which are responsible for their differing substrate and inhibitor sensitivities. In this study the role of ACE2 active-site residues was explored by site-directed mutagenesis. Arg273 was found to be critical for substrate binding such that its replacement causes enzyme activity to be abolished. Although both His505 and His345 are involved in catalysis, it is His345 and not His505 that acts as the hydrogen bond donor/acceptor in the formation of the tetrahedral peptide intermediate. The difference in chloride sensitivity between ACE2 and ACE was investigated, and the absence of a second chloride-binding site (CL2) in ACE2 confirmed. Thus ACE2 has only one chloride-binding site (CL1) whereas ACE has two sites. This is the first study to address the differences that exist between ACE2 and ACE at the molecular level. The results can be applied to future studies aimed at unravelling the role of ACE2, relative to ACE, in vivo.

Angiotensin-Converting Enzyme 2↗

Towards a structural classification of phosphate binding sites in protein-nucleotide complexes: an automated all-against-all structural comparison using geometric matching.

A method is described for the rapid comparison of protein binding sites using geometric matching to detect similar three-dimensional structure. The geometric matching detects common atomic features through identification of the maximum common sub-graph or clique. These features are not necessarily evident from sequence or from global structural similarity giving additional insight into molecular recognition not evident from current sequence or structural classification schemes. Here we use the method to produce an all-against-all comparison of phosphate binding sites in a number of different nucleotide phosphate-binding proteins. The similarity search is combined with clustering of similar sites to allow a preliminary structural classification. Clustering by site similarity produces a classification of binding sites for the 476 representative local environments producing ten main clusters representing half of the representative environments. The similarities make sense in terms of both structural and functional classification schemes. The ten main clusters represent a very limited number of unique structural binding motifs for phosphate. These are the structural P-loop, di-nucleotide binding motif [FAD/NAD(P)-binding and Rossman-like fold] and FAD-binding motif. Similar classification schemes for nucleotide binding proteins have also been arrived at independently by others using different methods.

Algorithms↗

Mutations in LRP5 or FZD4 underlie the common familial exudative vitreoretinopathy locus on chromosome 11q.

Familial exudative vitreoretinopathy (FEVR) is an inherited blinding disorder of the retinal vascular system. Autosomal dominant FEVR is genetically heterogeneous, but its principal locus, EVR1, is on chromosome 11q13-q23. The gene encoding the Wnt receptor frizzled-4 (FZD4) was recently reported to be the EVR1 gene, but our mutation screen revealed fewer patients harboring mutations than expected. Here, we describe mutations in a second gene at the EVR1 locus, low-density-lipoprotein receptor-related protein 5 (LRP5), a Wnt coreceptor. This finding further underlines the significance of Wnt signaling in the vascularization of the eye and highlights the potential dangers of using multiple families to refine genetic intervals in gene-identification studies.

Amino Acid Sequence↗

Angiotensin-converting enzyme-2 (ACE2): comparative modeling of the active site, specificity requirements, and chloride dependence.

Angiotensin-converting enzyme 2 (ACE2), a homologue of ACE, represents a new and potentially important target in cardio-renal disease. A model of the active site of ACE2, based on the crystal structure of testicular ACE, has been developed and indicates that the catalytic mechanism of ACE2 resembles that of ACE. Structural differences exist between the active site of ACE (dipeptidyl carboxypeptidase) and ACE2 (carboxypeptidase) that are responsible for the differences in specificity. The main differences occur in the ligand-binding pockets, particularly at the S2' subsite and in the binding of the peptide carboxy-terminus. The model explains why the classical ACE inhibitor lisinopril is unable to bind to ACE2. On the basis of the ability of ACE2 to cleave a variety of biologically active peptides, a consensus sequence of Pro-X-Pro-hydrophobic/basic for the protease specificity of ACE2 has been defined that is supported by the ACE2 model. The dipeptide, Pro-Phe, completely inhibits ACE2 activity at 180 microM with angiotensin II as the substrate. As with ACE, the chloride dependence of ACE2 is substrate-specific such that the hydrolysis of angiotensin I and the synthetic peptide substrate, Mca-APK(Dnp), are activated in the presence of chloride ions, whereas the cleavage of angiotensin II is inhibited. The ACE2 model is also suggestive of a possible mechanism for chloride activation. The structural insights provided by these analyses for the differences in inhibition pattern and substrate specificity among ACE and its homologue ACE2 and for the chloride dependence of ACE/ACE2 activity are valuable in understanding the function and regulation of ACE2.

Amino Acid Sequence↗

Ligand binding: functional site location, similarity and docking.

Computational methods for the detection and characterisation of protein ligand-binding sites have increasingly become an area of interest now that large amounts of protein structural information are becoming available prior to any knowledge of protein function. There have been particularly interesting recent developments in the following areas: first, functional site detection, whereby protein evolutionary information has been used to locate binding sites on the protein surface; second, functional site similarity, whereby structural similarity and three-dimensional templates can be used to compare and classify and potentially locate new binding sites; and third, ligand docking, which is being used to find and validate functional sites, in addition to having more conventional uses in small-molecule lead discovery.

Amino Acid Sequence↗

Q-fit: a probabilistic method for docking molecular fragments by sampling low energy conformational space.

A new method is presented that docks molecular fragments to a rigid protein receptor. It uses a probabilistic procedure based on statistical thermodynamic principles to place ligand atom triplets at the lowest energy sites. The probabilistic method ranks receptor binding modes so that the lowest energy ones are sampled first. This allows constraints to be introduced to limit the depth of the search leading to a computationally efficient method of sampling low energy conformational space. This is combined with energy minimization of the initial fragment placement to arrive at a low energy conformation for the molecular fragment. Two different search methods are tested involving (i) geometric hashing and (ii) pose clustering methods. Ten molecular fragments were docked that have commonly been used to test docking methods. The success rate was 8/10 and 10/10 for generating a close solution ranked first using the two different sampling procedures. In general, all five of the top ranked solutions reproduce the observed binding mode, which increases confidence in the predictions. A set of ten molecular fragments that have previously been identified as problematic were docked. Success was achieved in 3/10 and 4/10 using the two different methods. Again there is a high level of agreement between the two methods and again in the successful cases the top ranked solutions are correct whilst in the case of the failures none are. The geometric hashing and pose clustering methods are fast averaging approximately 13 and approximately 11 s per placement respectively using conservative parameters. The results are very encouraging and will facilitate the process of finding novel small molecule lead compounds by virtual screening of chemical databases.

Algorithms↗

A searchable database for comparing protein-ligand binding sites for the analysis of structure-function relationships.

The rapid expansion of structural information for protein-ligand binding sites is potentially an important source of information in structure-based drug design and in understanding ligand cross reactivity and toxicity. We have developed a large database of ligand binding sites extracted automatically from the Protein Data Bank. This has been combined with a method for calculating binding site similarity based on geometric hashing to create a relational database for the retrieval of site similarity and binding site superposition. It contains an all-against-all comparison of binding sites and holds known protein-ligand binding sites, which are made accessible to data mining. Here we demonstrate its utility in two structure-based applications: in determining site similarity and in aiding the derivation of a receptor-based pharmacophore model. The database is available from http://www.bioinformatics.leeds.ac.uk/sb/.

Binding Sites↗

Structure-based pharmacophore design and virtual screening for novel angiotensin converting enzyme 2 inhibitors.

The metallopeptidase Angiotensin Converting Enzyme (ACE) is an important drug target for the treatment of hypertension, heart, kidney, and lung disease. Recently, a close and unique human ACE homologue termed ACE2 has been identified and found to be an interesting new cardiorenal disease target. With the recently resolved inhibitor-bound ACE2 crystal structure available, we have attempted a structure-based approach to identify novel potent and selective inhibitors. Computational approaches focus on pharmacophore-based virtual screening of large compound databases. Selectivity was ensured by initial screening for ACE inhibitors within an internal database and the Derwent World Drug Index, which could be reduced to zero false positives and 0.1% hit rate, respectively. An average hit reduction of 0.44% was achieved with a five feature hypothesis, searching approximately 3.8 million compounds from various commercial databases. Seventeen compounds were selected based on high fit values as well as diverse structure and subjected to experimental validation in a bioassay. We show that all compounds displayed an inhibitory effect on ACE2 activity, the six most promising candidates exhibiting IC50 values in the range of 62-179 microM.

Angiotensin-Converting Enzyme 2↗