To abbreviate or not to abbreviate?
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
OBJECTIVE: To develop methods that automatically map abbreviations to their full forms in biomedical articles. METHODS: The authors developed two methods of mapping defined and undefined abbreviations (defined abbreviations are paired with their full forms in the articles, whereas undefined ones are not). For defined abbreviations, they developed a set of pattern-matching rules to map an abbreviation to its full form and implemented the rules into a software program, AbbRE (for "abbreviation recognition and extraction"). Using the opinions of domain experts as a reference standard, they evaluated the recall and precision of AbbRE for defined abbreviations in ten biomedical articles randomly selected from the ten most frequently cited medical and biological journals. They also measured the percentage of undefined abbreviations in the same set of articles, and they investigated whether they could map undefined abbreviations to any of four public abbreviation databases (GenBank LocusLink, SWISSPROT, LRABR of the UMLS Specialist Lexicon, and BioABACUS). RESULTS: AbbRE had an average 0.70 recall and 0.95 precision for the defined abbreviations. The authors found that an average of 25 percent of abbreviations were defined in biomedical articles and that of a randomly selected subset of undefined abbreviations, 68 percent could be mapped to any of four abbreviation databases. They also found that many abbreviations are ambiguous (i.e., they map to more than one full form in abbreviation databases). CONCLUSION: AbbRE is efficient for mapping defined abbreviations. To couple AbbRE with abbreviation databases for the mapping of undefined abbreviations, not only exhaustive abbreviation databases but also a method to resolve the ambiguity of abbreviations in the databases are needed.
CONTEXT: Abbreviations are used frequently in pathology reports and medical records. Efforts to identify and organize free-text concepts must correctly interpret medical abbreviations. During the past decade, the author has collected more than 12 000 medical abbreviations, concentrating on terms used or interpreted by pathologists. OBJECTIVE: The purpose of the study is to provide readers with a listing of abbreviations. The listing of abbreviations is reviewed for the purpose of determining the variety of ways that long forms are shortened. DESIGN: Abbreviations fell into different classes. These classes seemed amenable to distinct algorithmic approaches to their correct expansions. A discussion of these abbreviation classes was included to assist informaticians who are searching for ways to write software that expands abbreviations found in medical text. Classes were separated by the algorithmic approaches that could be used to map abbreviations to their correct expansions. A Perl implementation was developed to automatically match expansions with Unified Medical Language System concepts. MEASUREMENTS: The abbreviation list contained 12 097 terms; 5772 abbreviations had unique expansions. There were 6325 polysemous abbreviation/expansion pairs. The expansions of 8599 abbreviations mapped to Unified Medical Language System concepts. Three hundred twenty-four abbreviations could be confused with unabbreviated words. Two hundred thirteen abbreviations had different expansions depending on whether the American or the British spellings were used. Nine hundred seventy abbreviations ended in the letter "s."Results.-There were 6 nonexclusive groups of abbreviations classed by expansion algorithm, as follows: (1) ephemeral; (2) hyponymous; (3) monosemous; (4) polysemous; (5) masqueraders of common words; and (6) fatal (abbreviations whose incorrect expansions could easily result in clinical errors). CONCLUSION: Collecting and classifying abbreviations creates a logical approach to the development of class-specific algorithms designed to expand abbreviations. A large listing of medical abbreviations is placed into the public domain. The most current version is available at http://www.pathologyinformatics.org/downloads/abbtwo.htm.
OBJECTIVE: The growth of the biomedical literature presents special challenges for both human readers and automatic algorithms. One such challenge derives from the common and uncontrolled use of abbreviations in the literature. Each additional abbreviation increases the effective size of the vocabulary for a field. Therefore, to create an automatically generated and maintained lexicon of abbreviations, we have developed an algorithm to match abbreviations in text with their expansions. DESIGN: Our method uses a statistical learning algorithm, logistic regression, to score abbreviation expansions based on their resemblance to a training set of human-annotated abbreviations. We applied it to Medstract, a corpus of MEDLINE abstracts in which abbreviations and their expansions have been manually annotated. We then ran the algorithm on all abstracts in MEDLINE, creating a dictionary of biomedical abbreviations. To test the coverage of the database, we used an independently created list of abbreviations from the China Medical Tribune. MEASUREMENTS: We measured the recall and precision of the algorithm in identifying abbreviations from the Medstract corpus. We also measured the recall when searching for abbreviations from the China Medical Tribune against the database. RESULTS: On the Medstract corpus, our algorithm achieves up to 83% recall at 80% precision. Applying the algorithm to all of MEDLINE yielded a database of 781,632 high-scoring abbreviations. Of all the abbreviations in the list from the China Medical Tribune, 88% were in the database. CONCLUSION: We have developed an algorithm to identify abbreviations from text. We are making this available as a public abbreviation server at \url[http://abbreviation.stanford.edu/].
Abbreviations are widely used in medicine. The understanding of abbreviations is important for medical language processing and information retrieval systems. The Unified Medical Language System (UMLS) contains a large number of abbreviations. We hypothesized that extracting and studying the UMLS abbreviations can be helpful for understanding the characteristics of abbreviations in medicine. In this paper, we describe a method for extracting abbreviations from the UMLS. We evaluated the method and studied the ambiguous nature of the abbreviations. In addition, the coverage of the UMLS abbreviations in medical reports was studied. Using our method, we extracted 163,666 unique (abbreviation, full form) pairs from the UMLS with a precision of 97.5%, and a recall of 96%. The UMLS abbreviations were highly ambiguous: 33.1% of abbreviations with six characters or less had multiple meanings; the average number of different full forms for all abbreviations with six characters or less was 2.28. The coverage of the UMLS abbreviations in medical reports was over 66%.
MOTIVATION: Abbreviations are an important type of terminology in the biomedical domain. Although several groups have already created databases of biomedical abbreviations, these are either not public, or are not comprehensive, or focus exclusively on acronym-type abbreviations. We have created another abbreviation database, ADAM, which covers commonly used abbreviations and their definitions (or long-forms) within MEDLINE titles and abstracts, including both acronym and non-acronym abbreviations. RESULTS: A model of recognizing abbreviations and their long-forms from titles and abstracts of MEDLINE (2006 baseline) was employed. After grouping morphological variants, 59 405 abbreviation/long-form pairs were identified. ADAM shows high precision (97.4%) and includes most of the frequently used abbreviations contained in the Unified Medical Language System (UMLS) Lexicon and the Stanford Abbreviation Database. Conversely, one-third of abbreviations in ADAM are novel insofar as they are not included in either database. About 19% of the novel abbreviations are non-acronym-type and these cover at least seven different types of short-form/long-form pairs. AVAILABILITY: A free, public query interface to ADAM is available at http://arrowsmith.psych.uic.edu, and the entire database can be downloaded as a text file.
OBJECTIVES: the abbreviated mental test is widely used in the assessment of cognitive impairment in elderly patients. However, many doctors do not administer the full 10 questions, preferring to estimate the patient's score instead. We have studied the accuracy of doctors in predicting patients' abbreviated mental test scores. METHODS: we assessed 102 patients in the geriatric unit. We asked doctors to predict the patient's abbreviated mental test during the admission interview. A true abbreviated mental test was then recorded. RESULTS: mean age was 80.9 years with a male:female ratio of 27:74. The mean predicted abbreviated mental test score was 6.57 (SD 2.9); the mean actual abbreviated mental test score being 6.36 (SD 3.2). Comparing the two groups, abbreviated mental test scores were predicted most accurately at the extremes and correlation between the two groups of scores was high (P<0.001 Spearman test). Kappa statistics revealed moderate agreement between the two groups, (0.56, 95% CI 0.48-0.63). A predicted score of 5/10 showed the greatest spread of true abbreviated mental test scores (0-10, mean 4.5). However in total, only 31% of the predicted abbreviated mental test scores were accurate, with 42% being incorrect by >1. Using the accepted cut-off of <7/10, this revealed that 13% were underdiagnosed and 19% were overdiagnosed as being cognitively impaired. CONCLUSIONS: clinicians are poor at predicting abbreviated mental tests in the midrange but are more accurate at predicting lower and higher scores. This descriptive study reinforces the importance of using an objective assessment of cognitive impairment rather than clinicians estimating its presence or absence.
OBJECTIVE: The Pediatric Crohn's Disease Activity Index (PCDAI) is a validated measure of disease activity comprised of historical, laboratory and physical examination parameters. It has been suggested that an abbreviated PCDAI may be of similar utility without requiring laboratory evaluations or calculated height velocity. The aim of this study was to compare an abbreviated PCDAI and the original PCDAI and also compare the abbreviated PCDAI and a quality-of-life measurement. METHODS: The authors prospectively analyzed quality of life and disease activity, using the IMPACT-35 Questionnaire, the PCDAI and an abbreviated PCDAI consisting of three historical items (abdominal pain, stools and patient functioning) and three physical examination items (weight, abdomen and perirectal disease). RESULTS: Forty subjects aged 5-24 years (22 males) were included in analysis. Correlations were performed between the original PCDAI, an abbreviated PCDAI and the IMPACT-35. There was a significant, strong correlation between the PCDAI and the abbreviated PCDAI (n = 40, r = 0.849, p <0.001), a significant, moderate correlation between PCDAI and IMPACT-35 (n = 29, r = -0.547, p = 0.002) and a significant, moderate correlation between the abbreviated PCDAI and IMPACT-35 (n = 29, r = -0.579, p <0.001). CONCLUSIONS: An abbreviated PCDAI predicted disease activity as well as the full PCDAI. The IMPACT-35 correlated well with disease activity based on both PCDAI and an abbreviated PCDAI. An abbreviated PCDAI may offer advantages over the original PCDAI and should be prospectively validated in future studies.
MOTIVATION: Biological literature contains many abbreviations with one particular sense in each document. However, most abbreviations do not have a unique sense across the literature. Furthermore, many documents do not contain the long forms of the abbreviations. Resolving an abbreviation in a document consists of retrieving its sense in use. Abbreviation resolution improves accuracy of document retrieval engines and of information extraction systems. RESULTS: We combine an automatic analysis of Medline abstracts and linguistic methods to build a dictionary of abbreviation/sense pairs. The dictionary is used for the resolution of abbreviations occurring with their long forms. Ambiguous global abbreviations are resolved using support vector machines that have been trained on the context of each instance of the abbreviation/sense pairs, previously extracted for the dictionary set-up. The system disambiguates abbreviations with a precision of 98.9% for a recall of 98.2% (98.5% accuracy). This performance is superior in comparison with previously reported research work. AVAILABILITY: The abbreviation resolution module is available at http://www.ebi.ac.uk/Rebholz/software.html.
OBJECTIVE: To help biomedical researchers recognize dynamically introduced abbreviations in biomedical literature, such as gene and protein names, we have constructed a support system called ALICE (Abbreviation LIfter using Corpus-based Extraction). ALICE aims to extract all types of abbreviations with their expansions from a target paper on the fly. METHODS: ALICE extracts an abbreviation and its expansion from the literature by using heuristic pattern-matching rules. This system consists of three phases and potentially identifies valid 320 abbreviation-expansion patterns as combinations of the rules. RESULTS: It achieved 95% recall and 97% precision on randomly selected titles and abstracts from the MEDLINE database. CONCLUSION: ALICE extracted abbreviations and their expansions from the literature efficiently. The subtly compiled heuristics enabled it to extract abbreviations with high recall without significantly reducing precision. ALICE does not only facilitate recognition of an undefined abbreviation in a paper by constructing an abbreviation database or dictionary, but also makes biomedical literature retrieval more accurate. This system is freely available at http://uvdb3.hgc.jp/ALICE/ALICE_index.html.
Abbreviations are widely used in writing, and the understanding of abbreviations is important for natural language processing applications. Abbreviations are not always defined in a document and they are highly ambiguous. A knowledge base that consists of abbreviations with their associated senses and a method to resolve the ambiguities are needed. In this paper, we studied the UMLS coverage, textual variants of senses, and the ambiguity of abbreviations in MEDLINE abstracts. We restricted our study to three-letter abbreviations which were defined using parenthetical expressions. When grouping similar expansions together and representing senses using groups, we found that after ignoring senses where the total number of occurrences within the corresponding group was less than 100, 82.8% of the senses matched the UMLS, covered over 93% of occurrences that were considered, and had an average of 7.74 expansions for each sense. Abbreviations are highly ambiguous: 81.2% of the abbreviations were ambiguous, and had an average of 16.6 senses. However, after ignoring senses with occurrences of less than 5, 64.6% of the abbreviations were ambiguous, and had an average of 4.91 senses.
Although the use of abbreviations not understood by the average reader is discouraged by journal editors, I nevertheless found that 43% of 147 articles published during June 1993 in eight general and surgical journals contained uncommon abbreviations. In 26 (18%) of the 147 articles, all the abbreviations and their explanatory decoding words appeared at the front of the article, either in the abstract or in the first paragraph. This up front position makes easier the reader's back-search. In 37 other articles (25%), at least one uncommon abbreviation was decoded somewhere in the body of the article. In 21 articles (14%) the uncommon abbreviations appeared in the concluding or summating paragraph(s) and the explanatory decoding words were buried in the body of the article, thus making difficult the reader's back-search. Corrective action might include (1) editorial and peer review enforcement of the "no nonstandard abbreviation" policy, which is easily done with computerized word processing; (2) tabulation of all abbreviations with their decoding words either just below the abstract at the front of the article or just above the bibliography at the rear; or (3) expansion of each abbreviation in a footnote at the bottom of the appropriate page.
OBJECTIVE: To determine the effect of the use of abbreviations and acronyms on citation retrieval in MEDLINE searches. METHODS: Twenty common medical abbreviations that retrieved a minimum of 400 citations each in MEDLINE text word searches were studied. Each abbreviation was entered in a MEDLINE subject search to determine whether it mapped to an appropriate medical subject heading (MeSH) term. The MeSH category and the number of citations retrieved were recorded. The abbreviation and its definition were each entered in separate text word searches, and the number of citations retrieved was recorded. Sets were combined to determine the number of identical and unique citations retrieved in the searches. RESULTS: MEDLINE recognized all 20 abbreviations and mapped them to appropriate MeSH headings. MeSH term assignment, however, may be case- and space-sensitive. MeSH term searches retrieved more citations than text word searches for 18 of 20 abbreviations. Comparison of the document sets yielded by each search method revealed a subset of citations common to each. Although all sets retrieved showed overlap, no two were identical. In addition, each citation set contained a proportion of unique documents. CONCLUSION: Retrieval of all unique citations required three searches; subject with abbreviation, text word with abbreviation, and text word with definition. These results have important implications for MEDLINE users.
The National Library of Medicine announces its adoption of the Anglo-American standard for the formulation of journal title abbreviations according to the American National Standard for the Abbreviation of Titles of Periodicals (1969), with individual words abbreviated, in turn, according to the International List of Periodical Title Word Abbreviations (1970). The history of the activity of the specific Z39 Committee of USASI (now ANSI) concerned with journal title abbreviations is reviewed, covering the period from 1962 to the present. A history of the National Clearinghouse for Periodical Title Word Abbreviations and of the International List is also given. Former NLM usage is compared with the forms of the present International List and examples show the major changes in NLM abbreviations. The NLM Rules for Abbreviation of Periodical Titles as derived from the new standard are appended.
The Wechsler Adult Intelligence Scale-Third Edition (WAIS-III) often poses problems for many populations due to the length of administration. Twenty geriatric subjects were administered the full WAIS-III. Three abbreviated forms of the WAIS-III (Satz-Mogel abbreviation; seven-subtest short form; and a clinically derived abbreviation) were evaluated by rescoring original full WAIS-III protocols. Results showed that the abbreviated WAIS-III protocols were highly correlated with complete protocols, and classification rules were the highest for the clinically derived abbreviation. The clinically derived abbreviation was reevaluated in a college LD/ADHD population yielding similarly high correlations. Results support the use of abbreviated forms of the WAIS-III in the evaluation of elderly patients and young adults, and point to the clinically derived abbreviation as providing the smallest discrepancies from FSIQ.
Investigators have validated an abbreviated protocol for testing nonspecific bronchial reactivity with methacholine. We performed a similar validation study with histamine, another bronchoprovocative agent known to induce airflow obstruction. Histamine is pharmacologically distinct from methacholine and, under some circumstances, may provide specific clinical and investigative advantages to methacholine. Twenty-four patients with a clinical history of asthma underwent bronchoprovocative testing using the standard histamine airway protocol recommended by the American Academy of Allergy, Committee on Standardization of Bronchoprovocation. In addition, two abbreviated histamine challenge protocols were tested using the same administration and testing equipment. The abbreviated protocols involved fewer dilutions and dosages of histamine than the standard histamine protocol but covered the same range of cumulative doses. The two abbreviated protocols differed only in the intervals for determination of FEV1 between doses of histamine (30 s vs 3 min). The sequence of these three protocols was randomized for each study subject and each airway challenge was separated by one week. The two abbreviated protocols took significantly less time to administer than the standard protocol--18 min vs 30 min vs 44 min. Both the provocative dose to cause a 20 percent decline in the FEV1 (PD20 FEV1) and the slope of the dose-response curve were not significantly different between the standard protocol and either of the two abbreviated protocols. Moreover, a high degree of agreement was observed between the two abbreviated protocols and the standard histamine protocol for both the PD20 FEV1 and the slope of the dose-response curve. These findings indicate that similar estimates of bronchial reactivity are obtained from either of the abbreviated protocols when compared with the standard histamine protocol.