PubMed Health⌕ Search

PubMed · 15369595

Doublet method for very fast autocoding.

Abstract

BACKGROUND: Autocoding (or automatic concept indexing) occurs when a software program extracts terms contained within text and maps them to a standard list of concepts contained in a nomenclature. The purpose of autocoding is to provide a way of organizing large documents by the concepts represented in the text. Because textual data accumulates rapidly in biomedical institutions, the computational methods used to autocode text must be very fast. The purpose of this paper is to describe the doublet method, a new algorithm for very fast autocoding. METHODS: An autocoder was written that transforms plain-text into intercalated word doublets (e.g. "The ciliary body produces aqueous humor" becomes "The ciliary, ciliary body, body produces, produces aqueous, aqueous humor"). Each doublet is checked against an index of doublets extracted from a standard nomenclature. Matching doublets are assigned a numeric code specific for each doublet found in the nomenclature. Text doublets that do not match the index of doublets extracted from the nomenclature are not part of valid nomenclature terms. Runs of matching doublets from text are concatenated and matched against nomenclature terms (also represented as runs of doublets). RESULTS: The doublet autocoder was compared for speed and performance against a previously published phrase autocoder. Both autocoders are Perl scripts, and both autocoders used an identical text (a 170+ Megabyte collection of abstracts collected through a PubMed search) and the same nomenclature (neocl.xml, containing over 102,271 unique names of neoplasms). In side-by-side comparison on the same computer, the doublet method autocoder was 8.4 times faster than the phrase autocoder (211 seconds versus 1,776 seconds). The doublet method codes 0.8 Megabytes of text per second on a desktop computer with a 1.6 GHz processor. In addition, the doublet autocoder successfully matched terms that were missed by the phrase autocoder, while the phrase autocoder found no terms that were missed by the doublet autocoder. CONCLUSIONS: The doublet method of autocoding is a novel algorithm for rapid text autocoding. The method will work with any nomenclature and will parse any ascii plain-text. An implementation of the algorithm in Perl is provided with this article. The algorithm, the Perl implementation, the neoplasm nomenclature, and Perl itself, are all open source materials.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jules J Berman. 2004-09-15. Doublet method for very fast autocoding.. https://doi.org/10.1186/1472-6947-4-16

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Measuring patient and physician participation in exchanges on medications: Dialogue Ratio, Preponderance of Initiative, and Dialogical Roles.

OBJECTIVE: To identify, describe and characterize the patient and physician participation in content production in medication-related exchanges during primary care consultations. METHODS: Descriptive study of audio recordings of 422 medical encounters. MEDICODE, a validated instrument was used to analyze verbal exchanges on medications. Two main indicators of participation were developed: Dialogue Ratio (DR), a 0-1 scale indicating extent of monologue/dialogue; Preponderance of Initiative (PI), a -1 to +1 scale for patient/physician initiative. Participation analyses were conducted by content theme and medication categories (New, Represcribed and Active). RESULTS: We identified 1492 discussions of medications. Categorical analyses identified four communication roles patients and physicians adopted when participating in medication-related exchanges during consultations: (a) Listener, (b) Information Provider, (c) Participant, and (d) Instigator. The mean observed DRs and PIs indicated that monologues and physician initiation dominated medication-related exchanges. CONCLUSION: Four factors are suggested to explain the communicational behaviors observed: (1) patient knowledge about medications, (2) physician expertise, (3) patient experience with the medication, and (4) the act of prescribing. Our data indicate a generally low level of dialogue when discussing medications during primary care encounters since physicians' monologues seem to be the rule rather than the exception, pointing to a lack of mutuality in exchanges on medications. PRACTICE IMPLICATIONS: The proposed concepts offer a unique vocabulary and conceptual framework to help physicians master the necessary content and process skills required to discuss medications with patients.

Abstracting and Indexing↗