PubMed Health⌕ Search

Biomedical subjects

Ruth V Spriggs

Publications and source records attributed to Ruth V Spriggs.

6 recordsLinked to original sources

Effective function annotation through catalytic residue conservation.

Because of the extreme impact of genome sequencing projects, protein sequences without accompanying experimental data now dominate public databases. Homology searches, by providing an opportunity to transfer functional information between related proteins, have become the de facto way to address this. Although a single, well annotated, close relationship will often facilitate sufficient annotation, this situation is not always the case, particularly if mutations are present in important functional residues. When only distant relationships are available, the transfer of function information is more tenuous, and the likelihood of encountering several well annotated proteins with different functions is increased. The consequence for a researcher is a range of candidate functions with little way of knowing which, if any, are correct. Here, we address the problem directly by introducing a computational approach to accurately identify and segregate related proteins into those with a functional similarity and those where function differs. This approach should find a wide range of applications, including the interpretation of genomics/proteomics data and the prioritization of targets for high-throughput structure determination. The method is generic, but here we concentrate on enzymes and apply high-quality catalytic site data. In addition to providing a series of comprehensive benchmarks to show the overall performance of our approach, we illustrate its utility with specific examples that include the correct identification of haptoglobin as a nonenzymatic relative of trypsin, discrimination of acid-d-amino acid ligases from a much larger ligase pool, and the successful annotation of BioH, a structural genomics target.

Amino Acid Sequence↗

Identification of a motif that mediates polypyrimidine tract-binding protein-dependent internal ribosome entry.

We have identified a novel motif which consists of the sequence (CCU)(n) as part of a polypyrimidine-rich tract and permits internal ribosome entry. A number of constructs containing variations of this motif were generated and these were found to function as artificial internal ribosome entry segments (AIRESs) in vivo and in vitro in the presence of polypyrimidine tract-binding protein (PTB). The data show that for these sequences to function as IRESs the RNA must be present as a double-stranded stem and, in agreement with this, rather surprisingly, we show that PTB binds strongly to double-stranded RNA. All the cellular 5' untranslated regions (UTRs) tested that harbor this sequence were shown to contain internal ribosome entry segments that are dependent upon PTB for function in vivo and in vitro. This therefore raises the possibility that PTB or its interacting protein partners could provide a bridge between the IRES-RNA and the ribosome. Given the number of putative cellular IRESs that could be dependent on PTB for function, these data strongly suggest that PTB-1 is a universal IRES-trans-acting factor.

5' Untranslated Regions↗

The complement of enzymatic sets in different species.

We present here a comprehensive analysis of the complement of enzymes in a large variety of species. As enzymes are a relatively conserved group there are several classification systems available that are common to all species and link a protein sequence to an enzymatic function. Enzymes are therefore an ideal functional group to study the relationship between sequence expansion, functional divergence and phenotypic changes. By using information retrieved from the well annotated SWISS-PROT database together with sequence information from a variety of fully sequenced genomes and information from the EC functional scheme we have aimed here to estimate the fraction of enzymes in genomes, to determine the extent of their functional redundancy in different domains of life and to identify functional innovations and lineage specific expansions in the metazoa lineage. We found that prokaryote and eukaryote species differ both in the fraction of enzymes in their genomes and in the pattern of expansion of their enzymatic sets. We observe an increase in functional redundancy accompanying an increase in species complexity. A quantitative assessment was performed in order to determine the degree of functional redundancy in different species. Finally, we report a massive expansion in the number of mammalian enzymes involved in signalling and degradation.

Animals↗

A ligand-centric analysis of the diversity and evolution of protein-ligand relationships in E.coli.

As enzymes evolve and diverge from common ancestor sequences, they often keep their overall reaction chemistry but specialize in the binding of different cognate ligands. This study borrows methods for the computational assessment of 2D similarity of small molecules from the field of chemoinformatics, to examine the extent of structure conservation of cognate ligands binding to similar proteins. Proteins from 87 structural superfamilies from Escherichia coli form the core dataset, which is extended using homologues with functional assignments from any organism. We find that correlation of the substrate similarity with protein similarity (measured by either sequence-based or structure-based scores) can only be clearly established for very similar proteins. At low sequence identities, the superfamily to which a protein belongs can give helpful clues to its function, and more importantly, the confidence attached to such clues is superfamily-dependent. Our data indicate that only a few superfamilies show great substrate diversity, and that most exhibit conservation of at least part of the structural scaffold of the substrate.

Databases, Protein↗

SCOPEC: a database of protein catalytic domains.

MOTIVATION: Domains are the units of protein structure, function and evolution. It is therefore essential to utilize knowledge of domains when studying the evolution of function, or when assigning function to genome sequence data. For this purpose, we have developed a database of catalytic domains, SCOPEC, by combining structural domain information from SCOP, full-length sequence information from Swiss-Prot, and verified functional information from the Enzyme Classification (EC) database. Two major problems need to be overcome to create a database of domain-function relationships; (1) for sequences, EC numbers are typically assigned to whole sequences rather than the functional unit, and (2) The Protein Data Bank (PDB) structures elucidated from a larger multi-domain protein will often have EC annotation although the relevant catalytic domain may lie elsewhere. RESULTS: SCOPEC entries have high quality enzyme assignments; having passed both computational and manual checks. SCOPEC currently contains entries for 75% of all EC annotations in the PDB. Overall, EC number is fairly well conserved within a superfamily, even when the proteins are distantly related. Initial analysis is encouraging; suggesting that there is a 50:50 chance of conserved function in distant homologues first detected by a third iteration PSI-BLAST search. Therefore, we envisage that a knowledge-based approach to function assignment using the domain-EC relationships in SCOPEC will gain a marked improvement over this base line. AVAILABILITY: The SCOPEC database is a valuable resource in the analysis and prediction of protein structure and function. It can be obtained or queried at our website http://www.enzome.com

Catalysis↗

Searching for patterns of amino acids in 3D protein structures.

This paper describes the program ASSAM, which has been developed to search for patterns of amino acid side-chains in the 3D structures in the Protein Data Bank. ASSAM represents an amino acid by a vector drawn from the main chain towards the functional part of the amino acid and then computes a graph representation of a protein in which the individual side-chain vectors are the nodes and the intervector distances are the edges. The presence of a query pattern in a Protein Data Bank structure can then be searched for by means of a subgraph isomorphism algorithm. Recent enhancements to ASSAM allow searches to include the following: the main-chain structure in addition to the side-chains; the secondary structure and solvent accessibility of side-chains; allowable distances from a known binding-site; disulfide bridges; and improved generic and wild-card queries. The effectiveness of these approaches is demonstrated by extensive searches of the Protein Data Bank for typical 3D query patterns.

Algorithms↗