PubMed Health⌕ Search

Biomedical subjects

James D Buntrock

Publications and source records attributed to James D Buntrock.

3 recordsLinked to original sources

Automating the assignment of diagnosis codes to patient encounters using example-based and machine learning techniques.

OBJECTIVE: Human classification of diagnoses is a labor intensive process that consumes significant resources. Most medical practices use specially trained medical coders to categorize diagnoses for billing and research purposes. METHODS: We have developed an automated coding system designed to assign codes to clinical diagnoses. The system uses the notion of certainty to recommend subsequent processing. Codes with the highest certainty are generated by matching the diagnostic text to frequent examples in a database of 22 million manually coded entries. These code assignments are not subject to subsequent manual review. Codes at a lower certainty level are assigned by matching to previously infrequently coded examples. The least certain codes are generated by a naïve Bayes classifier. The latter two types of codes are subsequently manually reviewed. MEASUREMENTS: Standard information retrieval accuracy measurements of precision, recall and f-measure were used. Micro- and macro-averaged results were computed. RESULTS At least 48% of all EMR problem list entries at the Mayo Clinic can be automatically classified with macro-averaged 98.0% precision, 98.3% recall and an f-score of 98.2%. An additional 34% of the entries are classified with macro-averaged 90.1% precision, 95.6% recall and 93.1% f-score. The remaining 18% of the entries are classified with macro-averaged 58.5%. CONCLUSION: Over two thirds of all diagnoses are coded automatically with high accuracy. The system has been successfully implemented at the Mayo Clinic, which resulted in a reduction of staff engaged in manual coding from thirty-four coders to seven verifiers.

Abstracting and Indexing↗

Using compound codes for automatic classification of clinical diagnoses.

Classification of diagnoses (a.k.a. coding) is the central part of current concept based medical IR systems. Some classification systems contain over 30,000 distinct codes which makes classifying clinical documents a time consuming labor intensive and error prone process. This paper presents a simple methodology for cleaning up and reusing existing manually coded diagnostic statements mainly extracted from clinical notes to build predictive models using a sparse-feature implementation of a Naïve Bayes classifier. One of the problems addressed is that diagnostic statements often contain several diagnoses and are assigned several codes resulting in a multi-class classification problem. We investigate one possible way of addressing this problem by introducing compound (multiple code) categories. We present experimental results of classifying >16,000 randomly selected diagnostic strings into 19 top level categories. A small improvement (3%) with using compound categories over simple categories indicates that using multiple code categories is a promising solution, although clearly in need of further research and refinement.

Abstracting and Indexing↗

An evaluation of unmediated versus mediated retrieval services.

To understand if unmediated services could serve the data retrieval needs for the Mayo research investigator, a study was conducted to determine researcher interest, ability, and outcome of using a clinical data retrieval system. The results indicate about 25% of the research investigators would use a self-service retrieval tool. However, there is clear evidence a majority of the research investigators are satisfied with and prefer the mediated service because of convenience, retrieval specialist knowledge, and lack of time to perform the search themselves. Approximately 61% of the non-participants indicated they would be willing to pay a fee for continued use of the mediated service. This study confirms the interest in self-service retrieval tools, but the actual interest is lower than anticipated. The recommendation is to continue the use of mediated services and to offer self-service methods as needed, allowing the most options to the research investigator.

Biomedical Research↗