PubMed HealthSearch

Biomedical subjects

P D Karp

Publications and source records attributed to P D Karp.

3 recordsLinked to original sources

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

A knowledge base of the chemical compounds of intermediary metabolism.

This paper describes a publicly available knowledge base of the chemical compounds involved in intermediary metabolism. We consider the motivations for constructing a knowledge base of metabolic compounds, the methodology by which it was constructed, and the information that it currently contains. Currently the knowledge base describes 981 compounds, listing for each: synonyms for its name, a systematic name, CAS registry number, chemical formula, molecular weight, chemical structure and two-dimensional display coordinates for the structure. The Compound Knowledge Base (CompoundKB) illustrates several methodological principles that should guide the development of biological knowledge bases. I argue that biological datasets should be made available in multiple representations to increase their accessibility to end users, and I present multiple representations of the CompoundKB (knowledge base, relational data base and ASN. 1 representations). I also analyze the general characteristics of these representations to provide an understanding of their relative advantages and disadvantages. Another principle is that the error rate of biological data bases should be estimated and documented-this analysis is performed for the CompoundKB.

Artificial Intelligence

Artificial intelligence methods for theory representation and hypothesis formation.

This article describes artificial intelligence methods for representing theories in molecular biology, and for improving the predictive power of these theories using experimental data. A program called GENSIM provides a framework for representing theories that includes descriptions of classes of biological objects (genes, enzymes, etc.), and processes that specify potential interactions among these objects (such as enzymatic reactions). GENSIM can employ a theory specified within this framework to predict the outcomes of biological experiments. A program called HYPGENE comes into play when the observed outcome of an experiment does not match the outcome predicted by GENSIM. HYPGENE works backward from the error in GENSIMs prediction to postulate changes to both the theory embodied by GENSIM, and the presumed initial conditions of the experiment. I view HYPGENEs hypothesis generation task as a design problem, and I have adapted AI methods developed for design and planning to this task. These techniques were developed in conjunction with an in-depth study of the discovery of the gene regulation mechanism of attenuation in the E. coli tryptophan operon. Both GENSIM and HYPGENE have been tested on sample problems from the history of attenuation, and produced many of the same solutions as biologists did.

Artificial Intelligence