PubMed Health⌕ Search

Biomedical subjects

Nicholas O'Toole

Publications and source records attributed to Nicholas O'Toole.

5 recordsLinked to original sources

The structural genomics experimental pipeline: insights from global target lists.

Structural genomics (SG) initiatives are currently attempting to achieve the high-throughput determination of protein structures on a genome-wide scale. Here we analyze the SG target data that have been publicly released over a period of 16 months to assess the potential of the SG initiatives. We use statistical techniques most commonly applied in epidemiology to describe the dynamics of targets through the experimental SG pipeline. There is no clear bottleneck among the key stages of cloning, expression, purification and crystallization. An SG target will progress through each of these steps with a probability of approximately 45%. Around 80% of targets with diffraction data will yield a crystal structure, and 20% of targets with HSQC spectra will yield an NMR structure. We also find the overlaps among SG targets: 61% of SG protein sequences share at least 30% sequence identity with one or more other SG targets. There is no significant difference in average structure quality among SG structures and other structures in the PDB determined by "traditional" methods, but on average SG structures are deposited to the PDB twice as quickly after X-ray data collection.

Animals↗

A data management system for structural genomics.

BACKGROUND: Structural genomics (SG) projects aim to determine thousands of protein structures by the development of high-throughput techniques for all steps of the experimental structure determination pipeline. Crucial to the success of such endeavours is the careful tracking and archiving of experimental and external data on protein targets. RESULTS: We have developed a sophisticated data management system for structural genomics. Central to the system is an Oracle-based, SQL-interfaced database. The database schema deals with all facets of the structure determination process, from target selection to data deposition. Users access the database via any web browser. Experimental data is input by users with pre-defined web forms. Data can be displayed according to numerous criteria. A list of all current target proteins can be viewed, with links for each target to associated entries in external databases. To avoid unnecessary work on targets, our data management system matches protein sequences weekly using BLAST to entries in the Protein Data Bank and to targets of other SG centers worldwide. CONCLUSION: Our system is a working, effective and user-friendly data management tool for structural genomics projects. In this report we present a detailed summary of the various capabilities of the system, using real target data as examples, and indicate our plans for future enhancements.

Journal Article↗

The final player in the coenzyme A biosynthetic pathway.

The determination of the crystal structure of human phosphopantothenoylcysteine synthetase completes our knowledge of the enzyme structures involved in all steps of coenzyme A biosynthesis. This structure provides insight into the differences between bacterial and mammalian forms of the enzyme and may guide the structure-based development of novel antibacterial compounds.

Coenzyme A↗

Coverage of protein sequence space by current structural genomics targets.

By its purest definition the ultimate goal of structural genomics (SG) is the determination of the structures of all proteins encoded by genomes. Most of these will be obtained by homology modeling using the structures of a set of target proteins for experimental determination. Thanks to the open exchange of SG target information, we are able to analyze the sequences of the current target list to evaluate the extent of its coverage of protein sequence space. The presence of homologous sequences currently either in the Protein Data Bank (PDB) or among SG targets has been determined for each of the protein sequences in several organisms. In this way we are able to evaluate the coverage by existing or targeted structural data for the non-membranous parts of entire proteomes. For small bacterial proteomes such as that of H. influenzae almost all proteins have homologous sequences among SG targets or in the PDB. There is significantly lower coverage for more complex organisms, such as C. elegans. We have mapped the SG target list onto the ProtoMap clustering of protein sequences. Clusters occupied by SG targets represent over 150,000 protein sequences, which is approximately 44% of the total protein sequences classified by ProtoMap. The mapping of SG targets also enables an evaluation of the degree of overlap within the target list. An SG target typically occupies a ProtoMap cluster with more than six other homologous targets.

Algorithms↗

Crystal structure of a trimeric form of dephosphocoenzyme A kinase from Escherichia coli.

Coenzyme A (CoA) is an essential cofactor used in a wide variety of biochemical pathways. The final step in the biosynthesis of CoA is catalyzed by dephosphocoenzyme A kinase (DPCK, E.C. 2.7.1.24). Here we report the crystal structure of DPCK from Escherichia coli at 1.8 A resolution. This enzyme forms a tightly packed trimer in its crystal state, in contrast to its observed monomeric structure in solution and to the monomeric, homologous DPCK structure from Haemophilus influenzae. We have confirmed the existence of the trimeric form of the enzyme in solution using gel filtration chromatography measurements. Dephospho-CoA kinase is structurally similar to many nucleoside kinases and other P-loop-containing nucleotide triphosphate hydrolases, despite having negligible sequence similarity to these enzymes. Each monomer consists of five parallel beta-strands flanked by alpha-helices, with an ATP-binding site formed by a P-loop motif. Orthologs of the E. coli DPCK sequence exist in a wide range of organisms, including humans. Multiple alignment of orthologous DPCK sequences reveals a set of highly conserved residues in the vicinity of the nucleotide/CoA binding site.

Amino Acid Sequence↗