PubMed Health⌕ Search

Biomedical subjects

John N Dowling

Publications and source records attributed to John N Dowling.

7 recordsLinked to original sources

Inductive creation of an annotation schema for manually indexing clinical conditions from emergency department reports.

Evaluating automated indexing applications requires comparing automatically indexed terms against manual reference standard annotations. However, there are no standard guidelines for determining which words from a textual document to include in manual annotations, and the vague task can result in substantial variation among manual indexers. We applied grounded theory to emergency department reports to create an annotation schema representing syntactic and semantic variables that could be annotated when indexing clinical conditions. We describe the annotation schema, which includes variables representing medical concepts (e.g., symptom, demographics), linguistic form (e.g., noun, adjective), and modifier types (e.g., anatomic location, severity). We measured the schema's quality and found: (1) the schema was comprehensive enough to be applied to 20 unseen reports without changes to the schema; (2) agreement between author annotators applying the schema was high, with an F measure of 93%; and (3) the authors made complementary errors when applying the schema, demonstrating that the schema incorporates both linguistic and medical expertise.

Abstracting and Indexing↗

Generating a reliable reference standard set for syndromic case classification.

OBJECTIVE: To generate and measure the reliability for a reference standard set with representative cases from seven broad syndromic case definitions and several narrower syndromic definitions used for biosurveillance. DESIGN: From 527,228 eligible patients between 1990 and 2003, we generated a set of patients potentially positive for seven syndromes by classifying all eligible patients according to their ICD-9 primary discharge diagnoses. We selected a representative subset of the cases for chart review by physicians, who read emergency department reports and assigned values to 14 variables related to the seven syndromes. MEASUREMENTS: (1) Positive predictive value of the ICD-9 diagnoses; (2) prevalence of the syndromic definitions and related variables; (3) agreement between physician raters demonstrated by kappa, kappa corrected for bias and prevalence, and Finn's r; and (4) reliability of the reference standard classifications demonstrated by generalizability coefficients. RESULTS: Positive predictive value for ICD-9 classification ranged from 0.33 for botulinic to 0.86 for gastrointestinal. We generated between 80 and 566 positive cases for six of the seven syndromic definitions. Rash syndrome exhibited low prevalence (34 cases). Agreement between physician raters was high, with kappa > 0.70 for most variables. Ratings showed no bias. Finn's r was >0.70 for all variables. Generalizability coefficients were >0.70 for all variables but three. CONCLUSION: Of the 27 syndromes generated by the 14 variables, 21 showed high enough prevalence, agreement, and reliability to be used as reference standard definitions against which an automated syndromic classifier could be compared. Syndromic definitions that showed poor agreement or low prevalence include febrile botulinic syndrome, febrile and nonfebrile rash syndrome, respiratory syndrome explained by a nonrespiratory or noninfectious diagnosis, and febrile and nonfebrile gastrointestinal syndrome explained by a nongastrointestinal or noninfectious diagnosis.

Bioterrorism↗

Classification of emergency department chief complaints into 7 syndromes: a retrospective analysis of 527,228 patients.

STUDY OBJECTIVE: Electronic surveillance systems often monitor triage chief complaints in hopes of detecting an outbreak earlier than can be accomplished with traditional reporting methods. We measured the accuracy of a Bayesian chief complaint classifier called CoCo that assigns patients 1 of 7 syndromic categories (respiratory, botulinic, gastrointestinal, neurologic, rash, constitutional, or hemorrhagic) based on free-text triage chief complaints. METHODS: We compared CoCo's classifications with criterion syndromic classification based on International Classification of Diseases, Ninth Revision (ICD-9) discharge diagnoses. We assigned the criterion classification to a patient based on whether the patient's primary diagnosis was a member of a set of ICD-9 codes associated with CoCo's 7 syndromes. We tested CoCo's performance on a set of 527,228 chief complaints from patients registered at the University of Pittsburgh Medical Center emergency department (ED) between 1990 and 2003. We performed a sensitivity analysis by varying the ICD-9 codes in the criterion standard. We also tested CoCo on chief complaints from EDs in a second location (Utah). RESULTS: Approximately 16% (85,569/527,228) of the patients were classified according to the criterion standard into 1 of the 7 syndromes. CoCo's classification performance (number of cases by criterion standard, sensitivity [95% confidence interval (CI)], and specificity [95% CI]) was respiratory (34,916, 63.1 [62.6 to 63.6], 94.3 [94.3 to 94.4]); botulinic (1,961, 30.1 [28.2 to 32.2], 99.3 [99.3 to 99.3]); gastrointestinal (20,431, 69.0 [68.4 to 69.6], 95.6 [95.6 to 95.7]); neurologic (7,393, 67.6 [66.6 to 68.7], 92.7 [92.6 to 92.8]); rash (2,232, 46.8 [44.8 to 48.9], 99.3 [99.3 to 99.3]); constitutional (10,603, 45.8 [44.9 to 46.8], 96.6 [96.6 to 96.7]); and hemorrhagic (8,033, 75.2 [74.3 to 76.2], 98.5 [98.4 to 98.5]). The sensitivity analysis showed that the results were not affected by the choice of ICD-9 codes in the criterion standard. Classification accuracy did not differ on chief complaints from the second location. CONCLUSION: Our results suggest that, for most syndromes, our chief complaint classification system can identify about half of the patients with relevant syndromic presentations, with specificities higher than 90% and positive predictive values ranging from 12% to 44%.

Bayes Theorem↗

Classifying free-text triage chief complaints into syndromic categories with natural language processing.

OBJECTIVE: Develop and evaluate a natural language processing application for classifying chief complaints into syndromic categories for syndromic surveillance. INTRODUCTION: Much of the input data for artificial intelligence applications in the medical field are free-text patient medical records, including dictated medical reports and triage chief complaints. To be useful for automated systems, the free-text must be translated into encoded form. METHODS: We implemented a biosurveillance detection system from Pennsylvania to monitor the 2002 Winter Olympic Games. Because input data was in free-text format, we used a natural language processing text classifier to automatically classify free-text triage chief complaints into syndromic categories used by the biosurveillance system. The classifier was trained on 4700 chief complaints from Pennsylvania. We evaluated the ability of the classifier to classify free-text chief complaints into syndromic categories with a test set of 800 chief complaints from Utah. RESULTS: The classifier produced the following areas under the ROC curve: Constitutional = 0.95; Gastrointestinal = 0.97; Hemorrhagic = 0.99; Neurological = 0.96; Rash = 1.0; Respiratory = 0.99; Other = 0.96. Using information stored in the system's semantic model, we extracted from the Respiratory classifications lower respiratory complaints and lower respiratory complaints with fever with a precision of 0.97 and 0.96, respectively. CONCLUSION: Results suggest that a trainable natural language processing text classifier can accurately extract data from free-text chief complaints for biosurveillance.

Bayes Theorem↗

Fever detection from free-text clinical records for biosurveillance.

Automatic detection of cases of febrile illness may have potential for early detection of outbreaks of infectious disease either by identification of anomalous numbers of febrile illness or in concert with other information in diagnosing specific syndromes, such as febrile respiratory syndrome. At most institutions, febrile information is contained only in free-text clinical records. We compared the sensitivity and specificity of three fever detection algorithms for detecting fever from free-text. Keyword CC and CoCo classified patients based on triage chief complaints; Keyword HP classified patients based on dictated emergency department reports. Keyword HP was the most sensitive (sensitivity 0.98, specificity 0.89), and Keyword CC was the most specific (sensitivity 0.61, specificity 1.0). Because chief complaints are available sooner than emergency department reports, we suggest a combined application that classifies patients based on their chief complaint followed by classification based on their emergency department report, once the report becomes available.

Algorithms↗

Identifying respiratory findings in emergency department reports for biosurveillance using MetaMap.

Clinical conditions described in patients' dictated reports are necessary for automated detection of patients with respiratory illnesses such as inhalational anthrax and pneumonia. We applied MetaMap to emergency department reports to extract a set of 71 clinical conditions relevant to detection of a lower respiratory outbreak. We indexed UMLS terms in emergency department reports with MetaMap, filtered the indexed output with a specialized lexicon of UMLS terms for the domain, and mapped the clinical conditions of interest to concepts in the lexicon. We compared MetaMap's ability to accurately identify the conditions against a physician's manual annotations and evaluated incorrectly indexed features to determine what additional processing is necessary. MetaMap identified the clinical conditions with a recall of 0.72 and a precision of 0.56. Necessary processing beyond MetaMap's indexing includes finding validation, temporal discrimination, anatomic location discrimination, finding-disease discrimination, and contextual inference. Successful identification of clinical conditions in an emergency department report with MetaMap requires processing techniques specific to the clinical question of interest.

Abstracting and Indexing↗

Representative threats for research in public health surveillance.

A large number of biological agents can cause natural or bioterroristic disease outbreaks and each can present in a bewildering number of ways (e.g., a few cases versus many cases, confined to a building versus widely disseminated). This 'problem space' is a challenge for designers of early warning systems for disease outbreaks and the sheer size of this space is a barrier to progress. This paper addresses this problem by deriving nine categories of threats that represent a parsimonious characterization of the problem space. A literature search also identified one or more example outbreaks for each of the nine categories. These outbreaks have occurred in recent times and could be used by researchers in need of actual outbreak data for investigations of the role of different types of surveillance data and algorithms in outbreak detection. The methodological contribution of this research is a Criterion Set of threats for analysis and evaluation of detection systems. This set characterizes the problem space in a tractable manner with less loss of generality than analyses based on one or two selected diseases, which is representative of current analyses.

Bioterrorism↗