PubMed Health⌕ Search

PubMed · 11293781

Continuous versus categorical data for ROC analysis: some quantitative considerations.

Abstract

RATIONALE AND OBJECTIVES: Several authors have encouraged the use of a quasi-continuous rating scale for data collection in receiver operating characteristic (ROC) curve analysis of diagnostic modalities, rather than rating scales based on five to seven ordinal categories or levels of suspicion. Although many investigators have gone over to this method, a discussion of the issues continues. The present work provides a quantitative analysis from the viewpoint of measurement science. MATERIALS AND METHODS: A simple model of the effect of data discretization or quantization on the measurement of the variance of noisy data was developed. Then Monte Carlo simulations of multiple-reader, multiple-case ROC experiments were performed and analyzed in terms of components-of-variance models to investigate the effect of data quantization in that more complex setting. RESULTS: For single-reader studies, discretization into five categories can reduce the precision of ROC measurements by a large amount. The effect may be attenuated in multireader studies. CONCLUSION: More precise measurements of diagnostic detection performance and thus more efficient use of resources are served by good measurement methods. These are promoted by the use of a quasi-continuous rating scale in ROC studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

R F Wagner, S V Beiden, C E Metz. 2001. Continuous versus categorical data for ROC analysis: some quantitative considerations.. https://doi.org/10.1016/s1076-6332(03)80502-0

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Probability estimation when some observations are grouped.

This paper considers the use of additional questions for decreasing survey non-response rates and an approach for estimating a probability based on the results obtained. In a survey, the respondents are asked to answer an original question and follow-up questions, where the answers for the follow-up questions are grouped answers for the original question. For example, respondents are asked to provide an exact number of incidents, but in cases of 'Do not know' or 'Refuse' responses, they are subsequently asked to pick an answer from a less specific categorical scale. The new estimator obtains smaller variance asymptotically and does not depend on a distribution family. This method is applied to income questions in a survey regarding injury prevention and behaviours. Another application is survey data on intimate partner violence, where some amendments were applied for incorporating post-stratification weights and for using non-random grouping. For additional illustration, an example of parameter estimation on artificially generated data is presented.

Data Collection↗