PubMed Health⌕ Search

Biomedical subjects

James E Bray

Publications and source records attributed to James E Bray.

6 recordsLinked to original sources

Origin and Evolution of Bacterial Periplasmic Force Transducers.

In double-membraned bacteria, non-equilibrium processes that occur at the outer membrane are typically coupled to the chemiosmotically energized inner membrane. TolA and TonB are homologous proteins which energetically couple inner membrane motor proteins to the essential processes of outer membrane stabilization and substrate import, respectively. The evolutionary trajectories of these proteins have been difficult to elucidate due to low-sequence conservation, yet they are thought to transduce force similarly. Here, this problem was addressed using structural prediction approaches to identify and annotate force transduction operons to trace their distribution and evolutionary origins. In the process, we identify a novel outer membrane-tethering system and a previously unknown family of monomeric force transducers. This approach revealed putative tolA genes, and thus the core organizational principles of the tol-pal operon throughout diverse bacterial taxa. We discovered that the α-helical structure of the periplasm-spanning domain II of TolA previously thought its hallmark, is anomalous amongst most Tol-Pal systems. This structure is mainly prevalent in γ-proteobacteria, likely in adaptation to their lifestyle. Comparison of Tol-Pal and Ton system distribution suggests that TolA emerged from a TonB paralogue and co-emerged with Pal, the outer membrane-tethering lipoprotein that functionalizes the Tol-Pal system. We also determined that TolB, the Pal-mobilizing protein, likely emerged from a family of outer membrane proteins; and CpoB, a periplasmic factor that coordinates peptidoglycan remodeling with cell division, was originally a lipoprotein present in the ancestral Tol-Pal system. The extensive conservation of the Tol-Pal system throughout Gracilicutes highlights its significance in bacterial cell biology.

Evolution, Molecular↗

A practical and robust sequence search strategy for structural genomics target selection.

MOTIVATION: Target selection strategies for structural genomic projects must be able to prioritize gene regions on the basis of significant sequence similarity with proteins that have already been structurally determined. With the rapid development of protein comparison software a robust prioritization scheme should be independent of the choice of algorithm and be able to incorporate different sequence similarity thresholds. RESULTS: A robust target selection strategy has been developed that can assign a priority level to all genes in any genome. Structural assignments to genome sequences are calculated at two thresholds and six levels (1-6) describe the prioritization of all whole genes and partial gene regions. This simple two-threshold approach can be implemented with any fold recognition or homology detection algorithms. The results for 10 genomes are presented using the SSEARCH and PSI-BLAST programs. AVAILABILITY: Programs are available on request from the authors.

Algorithms↗

High-definition macromolecular composition of yeast RNA-processing complexes.

A remarkably large collection of evolutionarily conserved proteins has been implicated in processing of noncoding RNAs and biogenesis of ribonucleoproteins. To better define the physical and functional relationships among these proteins and their cognate RNAs, we performed 165 highly stringent affinity purifications of known or predicted RNA-related proteins from Saccharomyces cerevisiae. We systematically identified and estimated the relative abundance of stably associated polypeptides and RNA species using a combination of gel densitometry, protein mass spectrometry, and oligonucleotide microarray hybridization. Ninety-two discrete proteins or protein complexes were identified comprising 489 different polypeptides, many associated with one or more specific RNA molecules. Some of the pre-rRNA-processing complexes that were obtained are discrete sub-complexes of those previously described. Among these, we identified the IPI complex required for proper processing of the ITS2 region of the ribosomal RNA primary transcript. This study provides a high-resolution overview of the modular topology of noncoding RNA-processing machinery.

Amino Acid Sequence↗

Gene3D: structural assignments for the biologist and bioinformaticist alike.

The Gene3D database (http://www.biochem.ucl.ac.uk/bsm/cath_new/Gene3D/) provides structural assignments for genes within complete genomes. These are available via the internet from either the World Wide Web or FTP. Assignments are made using PSI-BLAST and subsequently processed using the DRange protocol. The DRange protocol is an empirically benchmarked method for assessing the validity of structural assignments made using sequence searching methods where appropriate assignment statistics are collected and made available. Gene3D links assignments to their appropriate entries in relevent structural and classification resources (PDBsum, CATH database and the Dictionary of Homologous Superfamilies). Release 2.0 of Gene3D includes 62 genomes, 2 eukaryotes, 10 archaea and 40 bacteria. Currently, structural assignments can be made for between 30 and 40 percent of any given genome. In any genome, around half of those genes assigned a structural domain are assigned a single domain and the other half of the genes are assigned multiple structural domains. Gene3D is linked to the CATH database and is updated with each new update of CATH.

Animals↗

The CATH extended protein-family database: providing structural annotations for genome sequences.

An automatic sequence search and analysis protocol (DomainFinder) based on PSI-BLAST and IMPALA, and using conservative thresholds, has been developed for reliably integrating gene sequences from GenBank into their respective structural families within the CATH domain database (http://www.biochem.ucl.ac.uk/bsm/cath_new). DomainFinder assigns a new gene sequence to a CATH homologous superfamily provided that PSI-BLAST identifies a clear relationship to at least one other Protein Data Bank sequence within that superfamily. This has resulted in an expansion of the CATH protein family database (CATH-PFDB v1.6) from 19,563 domain structures to 176,597 domain sequences. A further 50,000 putative homologous relationships can be identified using less stringent cut-offs and these relationships are maintained within neighbour tables in the CATH Oracle database, pending further evidence of their suggested evolutionary relationship. Analysis of the CATH-PFDB has shown that only 15% of the sequence families are close enough to a known structure for reliable homology modeling. IMPALA/PSI-BLAST profiles have been generated for each of the sequence families in the expanded CATH-PFDB and a web server has been provided so that new sequences may be scanned against the profile library and be assigned to a structure and homologous superfamily.

Algorithms↗

The CATH protein family database: a resource for structural and functional annotation of genomes.

Over the last decade, there have been huge increases in the numbers of protein sequences and structures determined. In parallel, many methods have been developed for recognising similarities between these proteins, arising from their common evolutionary background, and for clustering such relatives into protein families. Here we review some of the protein family resources available to the biologist and describe how these can be used to provide structural and functional annotations for newly determined sequences. In particular we describe recent developments to the CATH domain database of protein structural families which have facilitated genome annotation and which have also revealed important caveats that must be considered when transferring functional data between homologous proteins.

Databases, Protein↗