PubMed Health⌕ Search

PubMed · 10557857

A method for measuring interrater agreement on checklists.

Abstract

A method for measuring interrater agreement on checklists is presented. This technique does not assign individual scores to raters, but computes a single agreement score from the concordance of their check mark configurations. An overall coefficient of agreement, called phi, is derived. The agreement coefficient that is expected by chance and the statistical significance of phi are determined by statistical simulation. Despite the dichotomous nature of the checklist agreement (raters either agree or disagree on items), we show that the binomial distribution does not provide a means for testing the statistical significance of phi. A medical education study is used to illustrate the phi methodology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J M Sinacore, K J Connell, A J Olthoff, M H Friedman, M R Gecht. 1999. A method for measuring interrater agreement on checklists.. https://doi.org/10.1177/01632789922034284

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Conditional reliability of admissions interview ratings: extreme ratings are the most informative.

CONTEXT: Admissions interviews are unreliable and have poor predictive validity, yet are the sole measures of non-cognitive skills used by most medical school admissions departments. The low reliability may be due in part to variation in conditional reliability across the rating scale. OBJECTIVES: To describe an empirically derived estimate of conditional reliability and use it to improve the predictive validity of interview ratings. METHODS: A set of medical school interview ratings was compared to a Monte Carlo simulated set to estimate conditional reliability controlling for range restriction, response scale bias and other artefacts. This estimate was used as a weighting function to improve the predictive validity of a second set of interview ratings for predicting non-cognitive measures (USMLE Step II residuals from Step I scores). RESULTS: Compared with the simulated set, both observed sets showed more reliability at low and high rating levels than at moderate levels. Raw interview scores did not predict USMLE Step II scores after controlling for Step I performance (additional r2 = 0.001, not significant). Weighting interview ratings by estimated conditional reliability improved predictive validity (additional r2 = 0.121, P < 0.01). CONCLUSIONS: Conditional reliability is important for understanding the psychometric properties of subjective rating scales. Weighting these measures during the admissions process would improve admissions decisions.

Education, Medical, Undergraduate↗