PubMed Health⌕ Search

Biomedical subjects

David J Studholme

Publications and source records attributed to David J Studholme.

17 recordsLinked to original sources

G8: a novel domain associated with polycystic kidney disease and non-syndromic hearing loss.

UNLABELLED: We report a novel protein domain-G8-which contains five repeated beta-strand pairs and is present in some disease-related proteins such as PKHD1, KIAA1199, TMEM2 as well as other uncharacterized proteins. Most G8-containing proteins are predicted to be membrane-integral or secreted. The G8 domain may be involved in extracellular ligand binding and catalysis. It has been reported that mis-sense mutations in the two G8 domains of human PKHD1 protein resulted in a less stable protein and are associated with autosomal-recessive polycystic kidney disease, indicating the importance of the domain structure. G8 is also present in the N-terminus of some non-syndromic hearing loss disease-related proteins such as KIAA1109 and TMEM2. Discovery of G8 domain will be important for the research of the structure/function of related proteins and beneficial for the development of novel therapeutics. CONTACT: liangsp@hunnu.edu.cn

Amino Acid Sequence↗

Protein domains and architectural innovation in plant-associated Proteobacteria.

BACKGROUND: Evolution of new complex biological behaviour tends to arise by novel combinations of existing building blocks. The functional and evolutionary building blocks of the proteome are protein domains, the function of a protein being dependent on its constituent domains. We clustered completely-sequenced proteomes of prokaryotes on the basis of their protein domain content, as defined by Pfam (release 16.0). This revealed that, although there was a correlation between phylogeny and domain content, other factors also have an influence. This observation motivated an investigation of the relationship between an organism's lifestyle and the complement of domains and domain architectures found within its proteome. RESULTS: We took a census of all protein domains and domain combinations (architectures) encoded in the completely-sequenced proteobacterial genomes. Nine protein domain families were identified that are found in phylogenetically disparate plant-associated bacteria but are absent from non-plant-associated bacteria. Most of these are known to play a role in the plant-associated lifestyle, but they also included domain of unknown function DUF1427, which is found in plant symbionts and pathogens of the alpha-, beta- and gamma-Proteobacteria, but not known in any other organism. Further, several domains were identified as being restricted to phytobacteria and Eukaryotes. One example is the RolB/RolC glucosidase family, which is found only in Agrobacterium species and in plants. We identified the 0.5% of Pfam protein domain families that were most significantly over-represented in the plant-associated Proteobacteria with respect to the background frequencies in the whole set of available proteobacterial proteomes. These included guanylate cyclase, domains implicated in aromatic catabolism, cellulase and several domains of unknown function. We identified 459 unique domain architectures found in phylogenetically diverse plant pathogens and symbionts that were absent from non-pathogenic and non-symbiotic relatives. The vast majority of these were restricted to a single species or several closely related species and so their distributions could be better explained by phylogeny than by lifestyle. However, several architectures were found in two or more very distantly related phytobacteria but absent from non-plant-associated bacteria. Many of the proteins with these unique architectures are predicted to be secreted. In Pseudomonas syringae pathovar tomato, those genes encoding genes with novel domain architectures tended to have atypical GC contents and were adjacent to insertion sequence elements and phage-like sequences, suggesting acquisition by horizontal transfer. CONCLUSIONS: By identifying domains and architectures unique to plant pathogens and symbionts, we highlighted candidate proteins for involvement in plant-associated bacterial lifestyles. Given that characterisation of novel gene products in vivo and in vitro is time-consuming and expensive, this computational approach may be useful for reducing experimental search space. Furthermore we discuss the biological significance of novel proteins highlighted by this study in the context of plant-associated lifestyles.

Cluster Analysis↗

InterPro, progress and status in 2005.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created to integrate the major protein signature databases. Currently, it includes PROSITE, Pfam, PRINTS, ProDom, SMART, TIGRFAMs, PIRSF and SUPERFAMILY. Signatures are manually integrated into InterPro entries that are curated to provide biological and functional information. Annotation is provided in an abstract, Gene Ontology mapping and links to specialized databases. New features of InterPro include extended protein match views, taxonomic range information and protein 3D structure data. One of the new match views is the InterPro Domain Architecture view, which shows the domain composition of protein matches. Two new entry types were introduced to better describe InterPro entries: these are active site and binding site. PIRSF and the structure-based SUPERFAMILY are the latest member databases to join InterPro, and CATH and PANTHER are soon to be integrated. InterPro release 8.0 contains 11 007 entries, representing 2573 domains, 8166 families, 201 repeats, 26 active sites, 21 binding sites and 20 post-translational modification sites. InterPro covers over 78% of all proteins in the Swiss-Prot and TrEMBL components of UniProt. The database is available for text- and sequence-based searches via a webserver (http://www.ebi.ac.uk/interpro), and for download by anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Databases, Protein↗

Profiling the secretomes of plant pathogenic Proteobacteria.

Secreted proteins are central to the success of plant pathogenic bacteria. They are used by plant pathogens to adhere to and degrade plant cell walls, to suppress plant defence responses, and to deliver bacterial DNA and proteins into the cytoplasm of plant cells. However, experimental investigations into the identity and role of secreted proteins in plant pathogenesis have been hindered by the fact that many of these proteins are only expressed or secreted in planta, that knockout mutations of individual proteins frequently have little or no obvious phenotype, and that some obligate and fastidious plant pathogens remain recalcitrant to genetic manipulation. The availability of genome sequence data for a large number of agriculturally and scientifically important plant pathogens enables us to predict and compare the complete secretomes of these bacteria. In this paper we outline strategies that are currently being used to identify secretion systems and secreted proteins in Proteobacterial plant pathogens and discuss the implications of these analyses for future investigations into the molecular mechanisms of plant pathogenesis.

Bacterial Proteins↗

Novel protein domains and motifs in the marine planctomycete Rhodopirellula baltica.

The planctomycetes are a phylum of bacteria that have a unique cell compartmentalisation and yeast-like budding cell division and peptidoglycan-less proteinaceous cell walls. We wished to further our understanding of these unique organisms at the molecular level by searching for conserved amino acid sequence motifs and domains in the proteins encoded by Rhodopirellula baltica. Using BLAST and single-linkage clustering, we have discovered several new protein domains and sequence motifs in this planctomycete. R. baltica has multiple members of the newly discovered GEFGR protein family and the ASPIC C-terminal domain family, whilst most other organisms for which whole genome sequence is available have no more than one. Many of the domains and motifs appear to be restricted to the planctomycetes. It is possible that these protein domains and motifs may have been lost or replaced in other phyla, or they may have undergone multiple duplication events in the planctomycete lineage. One of the novel motifs probably represents a novel N-terminal export signal peptide. With their unique cell biology, it may be that the planctomycete cell compartmentalisation plan in particular needs special membrane transport mechanisms. The discovery of these new domains and motifs, many of which are associated with secretion and cell-surface functions, will help to stimulate experimental work and thus enhance further understanding of this fascinating group of organisms.

Amino Acid Motifs↗

Bioinformatic identification of novel regulatory DNA sequence motifs in Streptomyces coelicolor.

BACKGROUND: Streptomyces coelicolor is a bacterium with a vast repertoire of metabolic functions and complex systems of cellular development. Its genome sequence is rich in genes that encode regulatory proteins to control these processes in response to its changing environment. We wished to apply a recently published bioinformatic method for identifying novel regulatory sequence signals to gain new insights into regulation in S. coelicolor. RESULTS: The method involved production of position-specific weight matrices from alignments of over-represented words of DNA sequence. We generated 2497 weight matrices, each representing a candidate regulatory DNA sequence motif. We scanned the genome sequence of S. coelicolor against each of these matrices. A DNA sequence motif represented by one of the matrices was found preferentially in non-coding sequences immediately upstream of genes involved in polysaccharide degradation, including several that encode chitinases. This motif (TGGTCTAGACCA) was also found upstream of genes encoding components of the phosphoenolpyruvate phosphotransfer system (PTS). We hypothesise that this DNA sequence motif represents a regulatory element that is responsive to availability of carbon-sources. Other motifs of potential biological significance were found upstream of genes implicated in secondary metabolism (TTAGGTtAGgCTaACCTAA), sigma factors (TGACN19TGAC), DNA replication and repair (ttgtCAGTGN13TGGA), nucleotide conversions (CTACgcNCGTAG), and ArsR (TCAGN12TCAG). A motif found upstream of genes involved in chromosome replication (TGTCagtgcN7Tagg) was similar to a previously described motif found in UV-responsive promoters. CONCLUSIONS: We successfully applied a recently published in silico method to identify conserved sequence motifs in S. coelicolor that may be biologically significant as regulatory elements. Our data are broadly consistent with and further extend data from previously published studies. We invite experimental testing of our hypotheses in vitro and in vivo.

Base Sequence↗

In silico analysis of the sigma54-dependent enhancer-binding proteins in Pirellula species strain 1.

The planctomycetes are a phylogenetically distinct group of bacteria, widespread in aquatic and terrestrial environments. Their cell walls lack peptidoglycan and their compartmentalised cells undergo a yeast-like budding cell division process. Many bacteria regulate a subset of their genes by an enhancer-dependent mechanism involving the alternative sigma factor sigma54 (RpoN, sigmaN) in association with sigma54-dependent transcriptional activators known as enhancer-binding proteins (EBPs). The sigma54-dependent regulon has previously been studied in several groups of bacteria, but not in the planctomycetes. We wished to exploit the recently published complete genome sequence of Pirellula species strain 1 to predict and analyse the sigma54-dependent regulon in this interesting group of bacteria. The genome of Pirellula species strain 1 encodes one homologue of sigma54, and 16 sigma54-dependent EBPs, including 10 two-component response regulators and a homologue of Escherichia coli RtcR. Two EBPs contain forkhead-associated domains, representing a novel protein domain combination not previously observed in bacterial EBPs and suggesting a novel link between the enhancer-dependent regulon and 'eukaryotic-like' protein phosphorylation in bacterial signal transduction. We identified several potential sigma54-dependent promoters upstream of genes and operons including two homologues of csrA, which encodes the global regulator CsrA, and rtcBA, encoding a RNA 3'-terminal phosphate cyclase. Phylogenetic analysis of EBP sequences from a wide range of bacterial taxa suggested that planctomycete EBPs fall into several distinct clades. Also the phylogeny of the sigma54 factors is broadly consistent with that of the host organisms. These results are consistent with a very ancient origin of sigma54 within the bacterial lineage. The repertoire of functions predicted to be under the control of the sigma54-dependent regulon in Pirellula shares some similarities (e.g. rtcBA) as well as exhibiting differences with that in other taxonomic groups of bacteria, reinforcing the evolutionarily dynamic nature of this regulon.

Bacteria↗

The Pfam protein families database.

Pfam is a large collection of protein families and domains. Over the past 2 years the number of families in Pfam has doubled and now stands at 6190 (version 10.0). Methodology improvements for searching the Pfam collection locally as well as via the web are described. Other recent innovations include modelling of discontinuous domains allowing Pfam domain definitions to be closer to those found in structure databases. Pfam is available on the web in the UK (http://www.sanger.ac.uk/Software/Pfam/), the USA (http://pfam.wustl.edu/), France (http://pfam.jouy.inra.fr/) and Sweden (http://Pfam.cgb.ki.se/).

Animals↗

NCD3G: a novel nine-cysteine domain in family 3 GPCRs.

The NCD3G [for nine-cysteine domain of family 3 G-protein-coupled receptors (GPCRs)] domain is a novel protein domain that is conserved in family 3 GPCRs, including metabotropic glutamate receptors, calcium-sensing receptors, pheromone receptors and taste receptors, with the exception of GABA(B) receptors. The NCD3G domain contains nine highly conserved cysteine residues. Structural predictions suggest that NCD3G might possess four beta strands and three disulfide bridges. The structural model of NCD3G highlights the conserved residues co-segregated with certain familial diseases.

Amino Acid Sequence↗

DNA binding activity of the Escherichia coli nitric oxide sensor NorR suggests a conserved target sequence in diverse proteobacteria.

The Escherichia coli nitric oxide sensor NorR was shown to bind to the promoter region of the norVW transcription unit, forming at least two distinct complexes detectable by gel retardation. Three binding sites for NorR and two integration host factor binding sites were identified in the norR-norV intergenic region. The derived consensus sequence for NorR binding sites was used to search for novel members of the E. coli NorR regulon and to show that NorR binding sites are partially conserved in other members of the proteobacteria.

Amino Acid Sequence↗

A DNA element recognised by the molybdenum-responsive transcription factor ModE is conserved in Proteobacteria, green sulphur bacteria and Archaea.

BACKGROUND: The transition metal molybdenum is essential for life. Escherichia coli imports this metal into the cell in the form of molybdate ions, which are taken up via an ABC transport system. In E. coli and other Proteobacteria molybdenum metabolism and homeostasis are regulated by the molybdate-responsive transcription factor ModE. RESULTS: Orthologues of ModE are widespread amongst diverse prokaryotes, but not ubiquitous. We identified probable ModE-binding sites upstream of genes implicated in molybdenum metabolism in green sulphur bacteria and methanogenic Archaea as well as in Proteobacteria. We also present evidence of horizontal transfer of nitrogen fixation genes between green sulphur bacteria and methanogenic Archaea. CONCLUSIONS: Whereas most of the archaeal helix-turn-helix-containing transcription factors belong to families that are Archaea-specific, ModE is unusual in that it is found in both Archaea and Bacteria. Moreover, its cognate upstream DNA recognition sequence is also conserved between Archaea and Bacteria, despite the fundamental differences in their core transcription machinery. ModE is the third example of a transcriptional regulator with a binding signal that is conserved in Bacteria and Archaea.

Archaea↗

A comparison of Pfam and MEROPS: two databases, one comprehensive, and one specialised.

BACKGROUND: We wished to compare two databases based on sequence similarity: one that aims to be comprehensive in its coverage of known sequences, and one that specialises in a relatively small subset of known sequences. One of the motivations behind this study was quality control. Pfam is a comprehensive collection of alignments and hidden Markov models representing families of proteins and domains. MEROPS is a catalogue and classification of enzymes with proteolytic activity (peptidases or proteases). These secondary databases are used by researchers worldwide, yet their contents are not peer reviewed. Therefore, we hoped that a systematic comparison of the contents of Pfam and MEROPS would highlight missing members and false-positives leading to improvements in quality of both databases. An additional reason for carrying out this study was to explore the extent of consensus in the definition of a protein family. RESULTS: About half (89 out of 174) of the peptidase families in MEROPS overlapped single Pfam families. A further 32 MEROPS families overlapped multiple Pfam families. Where possible, new Pfam families were built to represent most of the MEROPS families that did not overlap Pfam. When comparing the numbers of sequences found in the overlap between a MEROPS family and its corresponding Pfam family, in most cases the overlap was substantial (52 pairs of MEROPS and Pfam families had an intersection size of greater than 75% of the union) but there were some differences in the sets of sequences included in the MEROPS families versus the overlapping Pfam families. CONCLUSIONS: A number of the discrepancies between MEROPS families and their corresponding Pfam families arose from differences in the aims and philosophies of the two databases. Examination of some of the discrepancies highlighted additional members of families, which have subsequently been added in both Pfam and MEROPS. This has led to improvements in the quality of both databases. Overall there was a great deal of consensus between the databases in definitions of a protein family.

Animals↗

Enhancer-dependent transcription in Salmonella enterica Typhimurium: new members of the sigmaN regulon inferred from protein sequence homology and predicted promoter sites.

DNA-looping mediated by regulatory proteins is a ubiquitous mode of transcriptional control that allows interactions between genetic elements separated over long distances in DNA. In prokaryotes, one of the best-studied examples of regulatory proteins that use DNA-looping is the NtrC family of enhancer-binding proteins (EBPs), which activate transcription from sigmaN (sigma-N, sigma-54) dependent promoters. The completely sequenced genome of food-borne pathogen Salmonella enterica serovar Typhimurium LT2 contains seven novel EBPs of unknown function. Four of these EBPs have a similar domain organisation to NtrC whilst surprisingly the remaining three resemble LevR in Bacillus subtilis. Probable transcriptional targets are identified for each of the EBPs, including novel homologues of phosphotransferase system Enzyme II (EII) and several virulence-associated functions. Comparisons are made with the related enteric bacteria Salmonella Typhi, Escherichia coli and Yersinia pestis.

Bacterial Proteins↗