PubMed Health⌕ Search

PubMed · 15542014

Gene name identification and normalization using a model organism database.

Abstract

Biology has now become an information science, and researchers are increasingly dependent on expert-curated biological databases to organize the findings from the published literature. We report here on a series of experiments related to the application of natural language processing to aid in the curation process for FlyBase. We focused on listing the normalized form of genes and gene products discussed in an article. We broke this into two steps: gene mention tagging in text, followed by normalization of gene names. For gene mention tagging, we adopted a statistical approach. To provide training data, we were able to reverse engineer the gene lists from the associated articles and abstracts, to generate text labeled (imperfectly) with gene mentions. We then evaluated the quality of the noisy training data (precision of 78%, recall 88%) and the quality of the HMM tagger output trained on this noisy data (precision 78%, recall 71%). In order to generate normalized gene lists, we explored two approaches. First, we explored simple pattern matching based on synonym lists to obtain a high recall/low precision system (recall 95%, precision 2%). Using a series of filters, we were able to improve precision to 50% with a recall of 72% (balanced F-measure of 0.59). Our second approach combined the HMM gene mention tagger with various filters to remove ambiguous mentions; this approach achieved an F-measure of 0.72 (precision 88%, recall 61%). These experiments indicate that the lexical resources provided by FlyBase are complete enough to achieve high recall on the gene list task, and that normalization requires accurate disambiguation; different strategies for tagging and normalization trade off recall for precision.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alexander A Morgan, Lynette Hirschman, Marc Colosimo, Alexander S Yeh, Jeff B Colombe. 2004. Gene name identification and normalization using a model organism database.. https://doi.org/10.1016/j.jbi.2004.08.010

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

On the persistence of supplementary resources in biomedical publications.

BACKGROUND: Providing for long-term and consistent public access to scientific data is a growing concern in biomedical research. One aspect of this problem can be demonstrated by evaluating the persistence of supplementary data associated with published biomedical papers. METHODS: We manually evaluated 655 supplementary data links extracted from PubMed abstracts published 1998-2005 (Method 1) as well as a further focused subset of 162 full-text manuscripts published within three representative high-impact biomedical journals between September and December 2004 (Method 2). RESULTS: For Method 1 we found that since 2001, only 71 - 92% of supplementary data were still accessible via the links provided, with 93% of these inaccessible links occurring where supplementary data was not stored with the publishing journal. Of the manuscripts evaluated in Method 2, we found that only 83% of these links were available approximately a year after publication, with 55% of these inaccessible links were at locations outside the journal of publication. CONCLUSION: We conclude that if supplemental data is required to support the publication, journals policies must take-on the responsibility to accept and store such data or require that it be maintained with a credible independent institution or under the terms of a strategic data storage plan specified by the authors. We further recommend that publishers provide automated systems to ensure that supplementary links remain persistent, and that granting bodies such as the NIH develop policies and funding mechanisms to maintain long-term persistent access to these data.

Abstracting and Indexing↗

Scientific papers presented at the European Congress of Radiology: a two-year comparison.

The purpose of this report was to determine the rate at which abstracts orally presented at the European Congress of Radiology (ECR) 2001 were published in 2001-2005 Medline-indexed journals and to compare publication rates and factors with presentations at the ECR in two different periods (2001 and 2000). Absolute and relative publication rates (APR, RPR) and different publication-related factors were analysed. From 991 abstracts originating from 52 countries, 449 articles (APR 45%) were subsequently published in 125 journals, most frequently in European Radiology (n=79, 18%). Country of origin statistically (p<0.0001) influences subsequent publication of the abstract, with Germany having the highest number of presentations (n=300) and derived articles (n=175, RPR 58%) whereas Sweden had the highest RPR (82%). Interventional and physics studies had the highest RPR (59% and 58%, respectively). The ECR meeting has a very high and stable APR (ECR 2001: 45% vs ECR 2000: 47%), and the journal European Radiology had the larger number of related publications (18% RPR following ECR 2001 compared with 14% from ECR 2000). Germany had the highest number of presentations and publications for both meetings. The highest RPR for ECR 2001 was found in interventional and physics studies whereas chest and cardiac studies had the highest RPR for ECR 2000.

Abstracting and Indexing↗