PubMed Health⌕ Search

Biomedical subjects

Alex Bateman

Publications and source records attributed to Alex Bateman.

16 recordsLinked to original sources

A comparison of Pfam and MEROPS: two databases, one comprehensive, and one specialised.

BACKGROUND: We wished to compare two databases based on sequence similarity: one that aims to be comprehensive in its coverage of known sequences, and one that specialises in a relatively small subset of known sequences. One of the motivations behind this study was quality control. Pfam is a comprehensive collection of alignments and hidden Markov models representing families of proteins and domains. MEROPS is a catalogue and classification of enzymes with proteolytic activity (peptidases or proteases). These secondary databases are used by researchers worldwide, yet their contents are not peer reviewed. Therefore, we hoped that a systematic comparison of the contents of Pfam and MEROPS would highlight missing members and false-positives leading to improvements in quality of both databases. An additional reason for carrying out this study was to explore the extent of consensus in the definition of a protein family. RESULTS: About half (89 out of 174) of the peptidase families in MEROPS overlapped single Pfam families. A further 32 MEROPS families overlapped multiple Pfam families. Where possible, new Pfam families were built to represent most of the MEROPS families that did not overlap Pfam. When comparing the numbers of sequences found in the overlap between a MEROPS family and its corresponding Pfam family, in most cases the overlap was substantial (52 pairs of MEROPS and Pfam families had an intersection size of greater than 75% of the union) but there were some differences in the sets of sequences included in the MEROPS families versus the overlapping Pfam families. CONCLUSIONS: A number of the discrepancies between MEROPS families and their corresponding Pfam families arose from differences in the aims and philosophies of the two databases. Examination of some of the discrepancies highlighted additional members of families, which have subsequently been added in both Pfam and MEROPS. This has led to improvements in the quality of both databases. Overall there was a great deal of consensus between the databases in definitions of a protein family.

Animals↗

Enhanced protein domain discovery by using language modeling techniques from speech recognition.

Most modern speech recognition uses probabilistic models to interpret a sequence of sounds. Hidden Markov models, in particular, are used to recognize words. The same techniques have been adapted to find domains in protein sequences of amino acids. To increase word accuracy in speech recognition, language models are used to capture the information that certain word combinations are more likely than others, thus improving detection based on context. However, to date, these context techniques have not been applied to protein domain discovery. Here we show that the application of statistical language modeling methods can significantly enhance domain recognition in protein sequences. As an example, we discover an unannotated Tf_Otx Pfam domain on the cone rod homeobox protein, which suggests a possible mechanism for how the V242M mutation on this protein causes cone-rod dystrophy.

Algorithms↗

New knowledge from old: in silico discovery of novel protein domains in Streptomyces coelicolor.

BACKGROUND: Streptomyces coelicolor has long been considered a remarkable bacterium with a complex life-cycle, ubiquitous environmental distribution, linear chromosomes and plasmids, and a huge range of pharmaceutically useful secondary metabolites. Completion of the genome sequence demonstrated that this diversity carried through to the genetic level, with over 7000 genes identified. We sought to expand our understanding of this organism at the molecular level through identification and annotation of novel protein domains. Protein domains are the evolutionary conserved units from which proteins are formed. RESULTS: Two automated methods were employed to rapidly generate an optimised set of targets, which were subsequently analysed manually. A final set of 37 domains or structural repeats, represented 204 times in the genome, was developed. Using these families enabled us to correlate items of information from many different resources. Several immediately enhance our understanding both of S. coelicolor and also general bacterial molecular mechanisms, including cell wall biosynthesis regulation and streptomycete telomere maintenance. DISCUSSION: Delineation of protein domain families enables detailed analysis of protein function, as well as identification of likely regions or residues of particular interest. Hence this kind of prior approach can increase the rate of discovery in the laboratory. Furthermore we demonstrate that using this type of in silico method it is possible to fairly rapidly generate new biological information from previously uncorrelated data.

Amino Acid Motifs↗

Rfam: an RNA family database.

Rfam is a collection of multiple sequence alignments and covariance models representing non-coding RNA families. Rfam is available on the web in the UK at http://www.sanger.ac.uk/Software/Rfam/ and in the US at http://rfam.wustl.edu/. These websites allow the user to search a query sequence against a library of covariance models, and view multiple sequence alignments and family annotation. The database can also be downloaded in flatfile form and searched locally using the INFERNAL package (http://infernal.wustl.edu/). The first release of Rfam (1.0) contains 25 families, which annotate over 50 000 non-coding RNA genes in the taxonomic divisions of the EMBL nucleotide database.

Animals↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

The CHAP domain: a large family of amidases including GSP amidase and peptidoglycan hydrolases.

Cleavage of peptidoglycan plays an important role in bacterial cell division, cell growth and cell lysis. Here, we reveal that several known peptidoglycan amidases fall into a family, which includes many proteins of previously unknown function. The family includes two different peptidoglycan cleavage activities: L-muramoyl-L-alanine amidase and D-alanyl-glycyl endopeptidase activity. The family includes the amidase portion of the bifunctional glutathionylspermidine synthase/amidase enzyme from bacteria and pathogenic trypanosomes. The glutathionylspermidine synthase is thought to be a key component of the alternative pathway in trypanosomes for protection from oxygen-radical damage and has been proposed as a potential drug target. The CHAP (cysteine, histidine-dependent amidohydrolases/peptidases) domain is often found in association with other domains that cleave peptidoglycan. The large number of multifunctional hydrolases suggests that they might act in a cooperative manner to cleave specialized substrates.

Amidohydrolases↗

Membrane-bound progesterone receptors contain a cytochrome b5-like ligand-binding domain.

BACKGROUND: Membrane-associated progesterone receptors (MAPRs) are thought to mediate a number of rapid cellular effects not involving changes in gene expression. They do not show sequence similarity to any of the classical steroid receptors. We were interested in identifying distant homologs of MAPR better to understand their biological roles. RESULTS: We have identified MAPRs as distant homologs of cytochrome b5. We have also found regions homologous to cytochrome b5 in the mammalian HERC2 ubiquitin transferase proteins and a number of fungal chitin synthases. CONCLUSIONS: In view of these findings, we propose that the heme-binding cytochrome b5 domain served as a template for the evolution of membrane-associated binding pockets for non-heme ligands.

Animals↗

The SGS3 protein involved in PTGS finds a family.

BACKGROUND: Post transcriptional gene silencing (PTGS) is a recently discovered phenomenon that is an area of intense research interest. Components of the PTGS machinery are being discovered by genetic and bioinformatics approaches, but the picture is not yet complete. RESULTS: The gene for the PTGS impaired Arabidopsis mutant sgs3 was recently cloned and was not found to have similarity to any other known protein. By a detailed analysis of the sequence of SGS3 we have defined three new protein domains: the XH domain, the XS domain and the zf-XS domain, that are shared with a large family of uncharacterised plant proteins. This work implicates these plant proteins in PTGS. CONCLUSION: The enigmatic SGS3 protein has been found to contain two predicted domains in common with a family of plant proteins. The other members of this family have been predicted to be transcription factors, however this function seems unlikely based on this analysis. A bioinformatics approach has implicated a new family of plant proteins related to SGS3 as potential candidates for PTGS related functions.

Amino Acid Sequence↗

The ENTH domain.

The epsin NH2-terminal homology (ENTH) domain is a membrane interacting module composed by a superhelix of alpha-helices. It is present at the NH2-terminus of proteins that often contain consensus sequences for binding to clathrin coat components and their accessory factors, and therefore function as endocytic adaptors. ENTH domain containing proteins have additional roles in signaling and actin regulation and may have yet other actions in the nucleus. The ENTH domain is structurally similar to the VHS domain. These domains define two families of adaptor proteins which function in membrane traffic and whose interaction with membranes is regulated, in part, by phosphoinositides.

Amino Acid Sequence↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.

Animals↗

The PASTA domain: a beta-lactam-binding domain.

The PASTA domain (for penicillin-binding protein and serine/threonine kinase associated domain) is found in the high molecular weight penicillin-binding proteins and eukaryotic-like serine/threonine kinases of a range of pathogens. We describe this previously uncharacterized domain and infer that it binds beta-lactam antibiotics and their peptidoglycan analogues. We postulate that PknB-like kinases are key regulators of cell-wall biosynthesis. The essential function of these enzymes suggests an additional pathway for the action of beta-lactam antibiotics.

Amino Acid Sequence↗

InterPro: an integrated documentation resource for protein families, domains and functional sites.

The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.

Algorithms↗

HMM-based databases in InterPro.

Protein family databases are an important resource for protein annotation and understanding protein evolution and function. In recent years hidden Markov models (HMMs) have become one of the key technologies used for detection of members of these families. This paper reviews the Pfam, TIGRFAMs and SMART databases that use the profile-HMMs provided by the HMMER package.

Computational Biology↗

QuickTree: building huge Neighbour-Joining trees of protein sequences.

We have written a fast implementation of the popular Neighbor-Joining tree building algorithm. QuickTree allows the reconstruction of phylogenies for very large protein families (including the largest Pfam alignment containing 27000 HIV GP120 glycoprotein sequences) that would be infeasible using other popular methods.

Algorithms↗

The use of structure information to increase alignment accuracy does not aid homologue detection with profile HMMs.

MOTIVATION: The best quality multiple sequence alignments are generally considered to derive from structural superposition. However, no previous work has studied the relative performance of profile hidden Markov models (HMMs) derived from such alignments. Therefore several alignment methods have been used to generate multiple sequence alignments from 348 structurally aligned families in the HOMSTRAD database. The performance of profile HMMs derived from the structural and sequence-based alignments has been assessed for homologue detection. RESULTS: The best alignment methods studied here correctly align nearly 80% of residues with respect to structure alignments. Alignment quality and model sensitivity are found to be dependent on average number, length, and identity of sequences in the alignment. The striking conclusion is that, although structural data may improve the quality of multiple sequence alignments, this does not add to the ability of the derived profile HMMs to find sequence homologues. SUPPLEMENTARY INFORMATION: A list of HOMSTRAD families used in this study and the corresponding Pfam families is available at http://www.sanger.ac.uk/Users/sgj/alignments/map.html CONTACT: sgj@sanger.ac.uk

Amino Acid Sequence↗