PubMed Health⌕ Search

Biomedical subjects

Janet M Thornton

Publications and source records attributed to Janet M Thornton.

At least 19 recordsLinked to original sources

An algorithm for constraint-based structural template matching: application to 3D templates with statistical analysis.

MOTIVATION: Structural templates consisting of a few atoms in a specific geometric conformation provide a powerful tool for studying the relationship between protein structure and function. Current methods for template searching constrain template syntax and semantics by their design. Hence there is a need for a more flexible core algorithm upon which to build more sophisticated tools. Statistical analysis of structural similarity is still in its infancy when compared with its analogue in sequence alignment. In the context of template matching, there is an urgent need for normalization of scores so that results from templates with differing sensitivity may be compared directly. RESULTS: We introduce Jess, a fast and flexible algorithm for searching protein structures for small groups of atoms under arbitrary constraints on geometry and chemistry. We apply the algorithm to a set of manually derived enzyme active site templates, and derive an empirical measure for estimating the relative significance of hits encountered using differing templates.

Algorithms↗

Catalysing new reactions during evolution: economy of residues and mechanism.

The diversity of function in some enzyme superfamilies shows that during evolution, enzymes have evolved to catalyse different reactions on the same structure scaffold. In this analysis, we examine in detail how enzymes can modify their chemistry, through a comparison of the catalytic residues and mechanisms in 27 pairs of homologous enzymes of totally different functions. We find that evolution is very economical. Enzymes retain structurally conserved residues to aid catalysis, including residues that bind catalytic metal ions and modulate cofactor chemistry. We examine the conservation of residue type and residue function in these structurally conserved residue pairs. Additionally, enzymes often retain common mechanistic steps catalyzed by structurally conserved residues. We have examined these steps in the context of their overall reactions.

Binding Sites↗

A template search reveals mechanistic similarities and differences in beta-ketoacyl synthases (KAS) and related enzymes.

A detailed comparison of the active sites in beta-ketoacyl synthases (KAS) and related enzymes has been made. Using three-dimensional templates of the three catalytic residues to scan the protein structural database reveals differences in both the geometry and the catalytic role of equivalent residues in different members of the family. The template based on the catalytic cysteine and two histidines in the KAS I and II is totally specific for this family, with no false hits. However, the role of the histidines in catalysis is different between KAS I/II and thiolase on the one hand and KAS III/chalcone synthase on the other. In contrast, a template comprising only cysteine and one histidine is not specific with many hits including members of the KAS family, metal binding sites, other active sites in nonhomologous proteins, and some "random" nonactive sites.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

Using a neural network and spatial clustering to predict the location of active sites in enzymes.

Structural genomics projects aim to provide a sharp increase in the number of structures of functionally unannotated, and largely unstudied, proteins. Algorithms and tools capable of deriving information about the nature, and location, of functional sites within a structure are increasingly useful therefore. Here, a neural network is trained to identify the catalytic residues found in enzymes, based on an analysis of the structure and sequence. The neural network output, and spatial clustering of the highly scoring residues are then used to predict the location of the active site.A comparison of the performance of differently trained neural networks is presented that shows how information from sequence and structure come together to improve the prediction accuracy of the network. Spatial clustering of the network results provides a reliable way of finding likely active sites. In over 69% of the test cases the active site is correctly predicted, and a further 25% are partially correctly predicted. The failures are generally due to the poor quality of the automatically generated sequence alignments. We also present predictions identifying the active site, and potential functional residues in five recently solved enzyme structures, not used in developing the method. The method correctly identifies the putative active site in each case. In most cases the likely functional residues are identified correctly, as well as some potentially novel functional groups.

Algorithms↗

Diversity of protein-protein interactions.

In this review, we discuss the structural and functional diversity of protein-protein interactions (PPIs) based primarily on protein families for which three-dimensional structural data are available. PPIs play diverse roles in biology and differ based on the composition, affinity and whether the association is permanent or transient. In vivo, the protomer's localization, concentration and local environment can affect the interaction between protomers and are vital to control the composition and oligomeric state of protein complexes. Since a change in quaternary state is often coupled with biological function or activity, transient PPIs are important biological regulators. Structural characteristics of different types of PPIs are discussed and related to their physiological function, specificity and evolution.

Animals↗

Using structural motif templates to identify proteins with DNA binding function.

This work describes a method for predicting DNA binding function from structure using 3-dimensional templates. Proteins that bind DNA using small contiguous helix-turn-helix (HTH) motifs comprise a significant number of all DNA-binding proteins. A structural template library of seven HTH motifs has been created from non-homologous DNA-binding proteins in the Protein Data Bank. The templates were used to scan complete protein structures using an algorithm that calculated the root mean squared deviation (rmsd) for the optimal superposition of each template on each structure, based on C(alpha) backbone coordinates. Distributions of rmsd values for known HTH-containing proteins (true hits) and non-HTH proteins (false hits) were calculated. A threshold value of 1.6 A rmsd was selected that gave a true hit rate of 88.4% and a false positive rate of 0.7%. The false positive rate was further reduced to 0.5% by introducing an accessible surface area threshold value of 990 A2 per HTH motif. The template library and the validated thresholds were used to make predictions for target proteins from a structural genomics project.

Algorithms↗

Integrating structure, bioinformatics, and enzymology to discover function: BioH, a new carboxylesterase from Escherichia coli.

Structural proteomics projects are generating three-dimensional structures of novel, uncharacterized proteins at an increasing rate. However, structure alone is often insufficient to deduce the specific biochemical function of a protein. Here we determined the function for a protein using a strategy that integrates structural and bioinformatics data with parallel experimental screening for enzymatic activity. BioH is involved in biotin biosynthesis in Escherichia coli and had no previously known biochemical function. The crystal structure of BioH was determined at 1.7 A resolution. An automated procedure was used to compare the structure of BioH with structural templates from a variety of different enzyme active sites. This screen identified a catalytic triad (Ser82, His235, and Asp207) with a configuration similar to that of the catalytic triad of hydrolases. Analysis of BioH with a panel of hydrolase assays revealed a carboxylesterase activity with a preference for short acyl chain substrates. The combined use of structural bioinformatics with experimental screens for detecting enzyme activity could greatly enhance the rate at which function is determined from structure.

Biotin↗

A novel approach to the recognition of protein architecture from sequence using Fourier analysis and neural networks.

A novel method is presented for the prediction of protein architecture from sequence using neural networks. The method involves the preprocessing of protein sequence data by numerically encoding it and then applying a Fourier transform. The encoded and transformed data are then used to train a neural network to recognize a number of different protein architectures. The method proved significantly better than comparable alternative strategies such as percentage dipeptide frequency, but is still limited by the size of the data set and the input demands of a neural network. Its main potential is as a complement to existing fold recognition techniques, with its ability to identify global symmetries within protein structures its greatest strength.

Algorithms↗

Structural characterisation and functional significance of transient protein-protein interactions.

Protein-protein complexes that dissociate and associate readily, often depending on the physiological condition or environment, play an important role in many biological processes. In order to characterise these "transient" protein-protein interactions, two sets of complexes were collected and analysed. The first set consists of 16 experimentally validated "weak" transient homodimers, which are known to exist as monomers and dimers at physiological concentration, with dissociation constants in the micromolar range. A set of 23 functionally validated transient (i.e. intracellular signalling) heterodimers comprise the second set. This set includes complexes that are more stable, with nanomolar binding affinities, and require a molecular trigger to form and break the interaction. In comparison to more stable homodimeric complexes, the weak homodimers demonstrate smaller contact areas between protomers and the interfaces are more planar and polar on average. The physicochemical and geometrical properties of these weak homodimers more closely resemble those of non-obligate hetero-oligomeric complexes, whose components can exist either as monomers or as complexes in vivo. In contrast to the weak transient dimers, "strong" transient dimers often undergo large conformational changes upon association/dissociation and are characterised with larger, less planar and sometimes more hydrophobic interfaces. From sequence alignments we find that the interface residues of the weak transient homodimers are generally more conserved than surface residues, consistent with being constrained to maintain the protein-protein interaction during evolution. Protein families that include members with different oligomeric states or structures are identified, and found to exhibit a lower sequence conservation at the interface. The results are discussed in terms of the physiological function and evolution of protein-protein interactions.

Animals↗

Gene3D: structural assignments for the biologist and bioinformaticist alike.

The Gene3D database (http://www.biochem.ucl.ac.uk/bsm/cath_new/Gene3D/) provides structural assignments for genes within complete genomes. These are available via the internet from either the World Wide Web or FTP. Assignments are made using PSI-BLAST and subsequently processed using the DRange protocol. The DRange protocol is an empirically benchmarked method for assessing the validity of structural assignments made using sequence searching methods where appropriate assignment statistics are collected and made available. Gene3D links assignments to their appropriate entries in relevent structural and classification resources (PDBsum, CATH database and the Dictionary of Homologous Superfamilies). Release 2.0 of Gene3D includes 62 genomes, 2 eukaryotes, 10 archaea and 40 bacteria. Currently, structural assignments can be made for between 30 and 40 percent of any given genome. In any genome, around half of those genes assigned a structural domain are assigned a single domain and the other half of the genes are assigned multiple structural domains. Gene3D is linked to the CATH database and is updated with each new update of CATH.

Animals↗

Analysis of metabolic networks using a pathway distance metric through linear programming.

The solution of the shortest path problem in biochemical systems constitutes an important step for studies of their evolution. In this paper, a linear programming (LP) algorithm for calculating minimal pathway distances in metabolic networks is studied. Minimal pathway distances are identified as the smallest number of metabolic steps separating two enzymes in metabolic pathways. The algorithm deals effectively with circularity and reaction directionality. The applicability of the algorithm is illustrated by calculating the minimal pathway distances for Escherichia coli small molecule metabolism enzymes, and then considering their correlations with genome distance (distance separating two genes on a chromosome) and enzyme function (as characterised by enzyme commission number). The results illustrate the effectiveness of the LP model. In addition, the data confirm that propinquity of genes on the genome implies similarity in function (as determined by co-involvement in the same region of the metabolic network), but suggest that no correlation exists between pathway distance and enzyme function. These findings offer insight into the probable mechanism of pathway evolution.

Algorithms↗

Analysis of catalytic residues in enzyme active sites.

We present an analysis of the residues directly involved in catalysis in 178 enzyme active sites. Specific criteria were derived to define a catalytic residue, and used to create a catalytic residue dataset, which was then analysed in terms of properties including secondary structure, solvent accessibility, flexibility, conservation, quaternary structure and function. The results indicate the dominance of a small set of amino acid residues in catalysis and give a picture of a general active site environment. It is hoped that this information will provide a better understanding of the molecular mechanisms involved in catalysis and a heuristic basis for predicting catalytic residues in enzymes of unknown function.

Amino Acid Sequence↗

One fold with many functions: the evolutionary relationships between TIM barrel families based on their sequences, structures and functions.

The eightfold (betaalpha) barrel structure, first observed in triose-phosphate isomerase, occurs ubiquitously in nature. It is nearly always an enzyme and most often involved in molecular or energy metabolism within the cell. In this review we bring together data on the sequence, structure and function of the proteins known to adopt this fold. We highlight the sequence and functional diversity in the 21 homologous superfamilies, which include 76 different sequence families. In many structures, the barrels are "mixed and matched" with other domains generating additional variety. Global and local structure-based alignments are used to explore the distribution of the associated functional residues on this common structural scaffold. Many of the substrates/co-factors include a phosphate moiety, which is usually but not always bound towards the C-terminal end of the sequence. Some, but not all, of these structures, exhibit a structurally conserved "phosphate binding motif". In contrast metal-ligating residues and catalytic residues are distributed along the sequence. However, we also found striking structural superposition of some of these residues. Lastly we consider the possible evolutionary relationships between these proteins, whose sequences are so diverse that even the most powerful approaches find few relationships, yet whose active sites all cluster at one end of the barrel. This extreme example of the "one fold-many functions" paradigm illustrates the difficulty of assigning function through a structural genomics approach for some folds.

Aldose-Ketose Isomerases↗

Toward predicting protein topology: an approach to identifying beta hairpins.

Although secondary structure prediction methods have recently improved, progress from secondary to tertiary structure prediction has been limited. A promising but largely unexplored route to this goal is to predict structure motifs from secondary structure knowledge. Here we present a novel method for the recognition of beta hairpins that combines secondary structure predictions and threading methods by using a database search and a neural network approach. The method successfully predicts 48 and 77%, respectively, of all of hairpin and nonhairpin beta-coil-beta motifs in a protein database. We find that the main contributors to motif recognition are predicted accessibility and turn propensities.

Databases, Factual↗

Prediction of strand pairing in antiparallel and parallel beta-sheets using information theory.

An information theory approach was developed to predict the alignment of interacting antiparallel and parallel beta-strands. Information scores were derived for the preference of a residue on a beta-strand to be opposite a sequence of residues on an adjacent beta-strand. These scores were used to predict the interstrand register of interacting beta-strands from 10 alternative offset positions either side of the experimentally observed beta-sheet register. The amino acid sequence of an internal beta-strand can be correctly aligned with two beta-strands in a fixed position either side of the strand in 45% of antiparallel and 48% of parallel arrangements. For comparison, when another beta-strand from a nonhomologous protein substitutes the internal beta-strand, the same register is predicted for only 24 and 36% of antiparallel and parallel arrangements. As expected, alignment of a single fixed strand with just a second beta-strand sequence was more difficult, and gave a correct register in 31 and 37% of antiparallel and parallel beta-pairs, respectively. These scores are 10% higher than for two randomly selected beta-strand sequences. In general, prediction accuracy was not improved by information tables that distinguished hydrogen-bonding patterns or beta-strand order. These results will contribute to predicting the arrangement of beta-strands in beta-pleated sheets and protein topology.

Amino Acid Sequence↗