PubMed Health⌕ Search

PubMed · 15203517

Multiple choice questions revisited.

Abstract

MCQs of the multiple true/false (MTF) variety were widely used in summative assessment 25 years ago. They could test a number of skills in addition to recall of factual knowledge, and were reliable, discriminatory, reproducible and cost-effective. However, there are now considerable doubts about their construct validity, mainly because of the varying responses of examinees to negative countermarking and the 'don't know' option, and the strategies they use when sitting examinations. Extended matching and one-from-five questions are now preferable, and negative countermarking is outmoded. MTF questions are still valuable in formative assessment and revision but are not recommended for summative examinations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

John Anderson. 2004. Multiple choice questions revisited.. https://doi.org/10.1080/0142159042000196141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Conditional reliability of admissions interview ratings: extreme ratings are the most informative.

CONTEXT: Admissions interviews are unreliable and have poor predictive validity, yet are the sole measures of non-cognitive skills used by most medical school admissions departments. The low reliability may be due in part to variation in conditional reliability across the rating scale. OBJECTIVES: To describe an empirically derived estimate of conditional reliability and use it to improve the predictive validity of interview ratings. METHODS: A set of medical school interview ratings was compared to a Monte Carlo simulated set to estimate conditional reliability controlling for range restriction, response scale bias and other artefacts. This estimate was used as a weighting function to improve the predictive validity of a second set of interview ratings for predicting non-cognitive measures (USMLE Step II residuals from Step I scores). RESULTS: Compared with the simulated set, both observed sets showed more reliability at low and high rating levels than at moderate levels. Raw interview scores did not predict USMLE Step II scores after controlling for Step I performance (additional r2 = 0.001, not significant). Weighting interview ratings by estimated conditional reliability improved predictive validity (additional r2 = 0.121, P < 0.01). CONCLUSIONS: Conditional reliability is important for understanding the psychometric properties of subjective rating scales. Weighting these measures during the admissions process would improve admissions decisions.

Education, Medical, Undergraduate↗