PubMed Health⌕ Search

PubMed · 2244082

Measuring interrater reliability among multiple raters: an example of methods for nominal data.

Abstract

This paper reviews and critiques various approaches to the measurement of reliability among multiple raters in the case of nominal data. We consider measurement of the overall reliability of a group of raters (using kappa-like statistics) as well as the reliability of individual raters with respect to a group. We introduce modifications of previously published estimators appropriate for measurement of reliability in the case of stratified sampling frames and we interpret these measures in view of standard errors computed using the jackknife. Analyses of a set of 48 anaesthesia case histories in which 42 anaesthesiologists independently rated the appropriateness of care on a nominal scale serve as an example.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

K L Posner, P D Sampson, R A Caplan, R J Ward, F W Cheney. 1990. Measuring interrater reliability among multiple raters: an example of methods for nominal data.. https://doi.org/10.1002/sim.4780090917

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results↗

Verifying DiffEXAFS measurements with differential X-ray diffraction.

Differential EXAFS (DiffEXAFS) is a novel technique for measuring atomic perturbations on a local scale. Here a complementary technique for such studies is presented: differential X-ray diffraction (DiffXRD), which may be used to independently verify DiffEXAFS results whilst using exactly the same experimental apparatus and measurement technique. A test experiment has been conducted to show that DiffXRD can be used to successfully determine the thermal expansion coefficient of SrF(2).

Reproducibility of Results↗