PubMed Health⌕ Search

PubMed · 15833721

Measuring teaching effectiveness in a pre-clinical multi-instructor course: a case study in the development and application of a brief instructor rating scale.

Abstract

BACKGROUND: Despite widespread use, misunderstandings persist about student evaluations of teaching. These evaluations have not been well examined in the common medical school setting of the multi-instructor, preclinical lecture course. PURPOSE: The study evaluated the psychometrics of a brief student evaluation of a teaching instrument developed for a multi-instructor 2nd-year course and described its application. METHODS: An 11-item instrument was developed and administered to 276 students to evaluate 27 lecturers per year in 3 years of an introductory clinical psychiatry course. A fully crossed research design allowed for a thorough analysis of variability in ratings. RESULTS: Generalizability analysis showed good reliability and relatively large Student x Lecturer interactions. Profile analysis generated distinct lecturer teaching profiles. CONCLUSIONS: Judicious use of a psychometrically sound student evaluation of a teaching instrument can be used to assist faculty and course development. Administering the evaluation instrument to an entire class produces no better reliability than administration to randomly selected subgroups of students.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Martin H Leamon, Laurie Fields. 2005. Measuring teaching effectiveness in a pre-clinical multi-instructor course: a case study in the development and application of a brief instructor rating scale.. https://doi.org/10.1207/s15328015tlm1702_5

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Conditional reliability of admissions interview ratings: extreme ratings are the most informative.

CONTEXT: Admissions interviews are unreliable and have poor predictive validity, yet are the sole measures of non-cognitive skills used by most medical school admissions departments. The low reliability may be due in part to variation in conditional reliability across the rating scale. OBJECTIVES: To describe an empirically derived estimate of conditional reliability and use it to improve the predictive validity of interview ratings. METHODS: A set of medical school interview ratings was compared to a Monte Carlo simulated set to estimate conditional reliability controlling for range restriction, response scale bias and other artefacts. This estimate was used as a weighting function to improve the predictive validity of a second set of interview ratings for predicting non-cognitive measures (USMLE Step II residuals from Step I scores). RESULTS: Compared with the simulated set, both observed sets showed more reliability at low and high rating levels than at moderate levels. Raw interview scores did not predict USMLE Step II scores after controlling for Step I performance (additional r2 = 0.001, not significant). Weighting interview ratings by estimated conditional reliability improved predictive validity (additional r2 = 0.121, P < 0.01). CONCLUSIONS: Conditional reliability is important for understanding the psychometric properties of subjective rating scales. Weighting these measures during the admissions process would improve admissions decisions.

Education, Medical, Undergraduate↗