PubMed Health⌕ Search

PubMed · 16464249

A MEDLINE categorization algorithm.

Abstract

BACKGROUND: Categorization is designed to enhance resource description by organizing content description so as to enable the reader to grasp quickly and easily what are the main topics discussed in it. The objective of this work is to propose a categorization algorithm to classify a set of scientific articles indexed with the MeSH thesaurus, and in particular those of the MEDLINE bibliographic database. In a large bibliographic database such as MEDLINE, finding materials of particular interest to a specialty group, or relevant to a particular audience, can be difficult. The categorization refines the retrieval of indexed material. In the CISMeF terminology, metaterms can be considered as super-concepts. They were primarily conceived to improve recall in the CISMeF quality-controlled health gateway. METHODS: The MEDLINE categorization algorithm (MCA) is based on semantic links existing between MeSH terms and metaterms on the one hand and between MeSH subheadings and metaterms on the other hand. These links are used to automatically infer a list of metaterms from any MeSH term/subheading indexing. Medical librarians manually select the semantic links. RESULTS: The MEDLINE categorization algorithm lists the medical specialties relevant to a MEDLINE file by decreasing order of their importance. The MEDLINE categorization algorithm is available on a Web site. It can run on any MEDLINE file in a batch mode. As an example, the top 3 medical specialties for the set of 60 articles published in BioMed Central Medical Informatics & Decision Making, which are currently indexed in MEDLINE are: information science, organization and administration and medical informatics. CONCLUSION: We have presented a MEDLINE categorization algorithm in order to classify the medical specialties addressed in any MEDLINE file in the form of a ranked list of relevant specialties. The categorization method introduced in this paper is based on the manual indexing of resources with MeSH (terms/subheadings) pairs by NLM indexers. This algorithm may be used as a new bibliometric tool.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stefan J Darmoni, Aurelie Névéol, Jean-Marie Renard, Jean-Francois Gehanno, Lina F Soualmia, Badisse Dahamna, Benoit Thirion. 2006-02-07. A MEDLINE categorization algorithm.. https://doi.org/10.1186/1472-6947-6-7

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A comparative evaluation of full-text, concept-based, and context-sensitive search.

OBJECTIVES: Study comparatively (1) concept-based search, using documents pre-indexed by a conceptual hierarchy; (2) context-sensitive search, using structured, labeled documents; and (3) traditional full-text search. Hypotheses were: (1) more contexts lead to better retrieval accuracy; and (2) adding concept-based search to the other searches would improve upon their baseline performances. DESIGN: Use our Vaidurya architecture, for search and retrieval evaluation, of structured documents classified by a conceptual hierarchy, on a clinical guidelines test collection. MEASUREMENTS: Precision computed at different levels of recall to assess the contribution of the retrieval methods. Comparisons of precisions done with recall set at 0.5, using t-tests. RESULTS: Performance increased monotonically with the number of query context elements. Adding context-sensitive elements, mean improvement was 11.1% at recall 0.5. With three contexts, mean query precision was 42% +/- 17% (95% confidence interval [CI], 31% to 53%); with two contexts, 32% +/- 13% (95% CI, 27% to 38%); and one context, 20% +/- 9% (95% CI, 15% to 24%). Adding context-based queries to full-text queries monotonically improved precision beyond the 0.4 level of recall. Mean improvement was 4.5% at recall 0.5. Adding concept-based search to full-text search improved precision to 19.4% at recall 0.5. CONCLUSIONS: The study demonstrated usefulness of concept-based and context-sensitive queries for enhancing the precision of retrieval from a digital library of semi-structured clinical guideline documents. Concept-based searches outperformed free-text queries, especially when baseline precision was low. In general, the more ontological elements used in the query, the greater the resulting precision.

Abstracting and Indexing↗

Improvements to the Red List Index.

The Red List Index uses information from the IUCN Red List to track trends in the projected overall extinction risk of sets of species. It has been widely recognised as an important component of the suite of indicators needed to measure progress towards the international target of significantly reducing the rate of biodiversity loss by 2010. However, further application of the RLI (to non-avian taxa in particular) has revealed some shortcomings in the original formula and approach: It performs inappropriately when a value of zero is reached; RLI values are affected by the frequency of assessments; and newly evaluated species may introduce bias. Here we propose a revision to the formula, and recommend how it should be applied in order to overcome these shortcomings. Two additional advantages of the revisions are that assessment errors are not propagated through time, and the overall level extinction risk can be determined as well as trends in this over time.

Abstracting and Indexing↗