PubMed Health⌕ Search

Biomedical subjects

Steven M Downing

Publications and source records attributed to Steven M Downing.

At least 19 recordsLinked to original sources

Validation of a colonoscopy simulation model for skills assessment.

OBJECTIVE: The purpose is to provide initial validation of a novel simulation model's fidelity and ability to assess competence in colonoscopy skills. METHODS: In a prospective, cross-sectional design, each of 39 endoscopists (13 staff, 13 second year fellows, and 13 novices) performed a colonoscopy on a novel bovine simulation model. Staff endoscopists also completed a survey examining different aspects of the model's realism as compared to human colonoscopy. The groups' simulation performances were compared. Additionally, individual performances were correlated to patient-based performance data. RESULTS: Median model realism evaluation scores were favorable for nearly all parameters evaluated with mucosa appearance, endoscopic view, and paradoxical motion parameters receiving the highest scores. During simulation procedures, each group outperformed the less experienced groups in all parameters evaluated. Specifically, median cecal intubation times were: staff 226 s (IQR [interquartile range] 179-273), fellows 340 s (282-568), and novices 1,027 s (970-1,122) (P < 0.05). Median total procedure times on the model were: staff 468 s (416-501), fellows 527 s (459-824), and novices 1,350 s (1,318-1,428) (P < 0.05). Finally, individual cecal intubation times on the simulation model had a very high correlation to their respective patient-based times (r = 0.764). CONCLUSIONS: Overall, this model possesses a favorable degree of realism and is able to easily differentiate users based on their level of colonoscopy experience. More impressive, however, is the strong correlation between individual's simulated intubation times and actual patient-based colonoscopy data. In light of these findings, we speculate that this model has potential to be an effective tool for assessment of colonoscopic competence.

Animals↗

Prior experiences associated with residents' scores on a communication and interpersonal skill OSCE.

OBJECTIVE: This exploratory study investigated whether prior task experience and comfort correlate with scores on an assessment of patient-centered communication. METHODS: A six-station standardized patient exam assessed patient-centered communication of 79 PGY2-3 residents in Internal Medicine and Family Medicine. A survey provided information on prior experiences. t-tests, correlations, and multi-factorial ANOVA explored relationship between scores and experiences. RESULTS: Experience with a task predicted comfort but did not predict communication scores. Comfort was moderately correlated with communication scores for some tasks; residents who were less comfortable were indeed less skilled, but greater comfort did not predict higher scores. Female gender and medical school experiences with standardized patients along with training in patient-centered interviewing were associated with higher scores. Residents without standardized patient experiences in medical school were almost five times more likely to be rejected by patients. CONCLUSIONS: Task experience alone does not guarantee better communication, and may instill a false sense of confidence. Experiences with standardized patients during medical school, especially in combination with interviewing courses, may provide an element of "deliberate practice" and have a long-term impact on communication skills. PRACTICE IMPLICATIONS: The combination of didactic courses and practice with standardized patients may promote a patient-centered approach.

Analysis of Variance↗

Use of flawed multiple-choice items by the New England Journal of Medicine for continuing medical education.

Physicians in the United States are required to complete a minimum number of continuing medical education (CME) credits annually. The goal of CME is to ensure that physicians maintain their knowledge and skills throughout their medical career. The New England Journal of Medicine (NEJM) provides its readers with the opportunity to obtain weekly CME credits. Deviation from established item-writing principles may result in a decrease in validity evidence for tests. This study evaluated the quality of 40 NEJM MCQs using the standard evidence-based principles of effective item writing. Each multiple-choice item reviewed had at least three item flaws, with a mean of 5.1 and a range of 3 to 7. The results of this study demonstrate that the NEJM uses flawed MCQs in its weekly CME program.

Education, Medical, Continuing↗

Developing an institution-based assessment of resident communication and interpersonal skills.

PURPOSE: The authors describe the development and validation of an institution-wide, cross-specialty assessment of residents' communication and interpersonal skills, including related components of patient care and professionalism. METHOD: Residency program faculty, the department of medical education, and the Clinical Performance Center at the University of Illinois at Chicago College of Medicine collaborated to develop six standardized patient-based clinical simulations. The standardized patients rated the residents' performance. The assessment was piloted in 2003 for internal medicine and family medicine and was subsequently adapted for other specialties, including surgery, pediatrics, obstetrics-gynecology, and neurology. We present validity evidence based on the content, internal structure, relationship to other variables, feasibility, acceptability, and impact of the 2003 assessment. RESULTS: Seventy-nine internal medicine and family medicine residents participated in the initial administration of the assessment. A factor analysis of the 18 communication scale items resulted in two factors interpretable as "communication" and "interpersonal skills." Median internal consistency of the scale (coefficient alpha) was 0.91. Generalizability of the assessment ranged from 0.57 to 0.82 across specialties. Case-specific items provided information about group-level deficiencies. Cost of the assessment was about $250 per resident. Once the initial cases had been developed and piloted, they could be adapted for other specialties with minimal additional effort, at a cost saving of about $1,000 per program. CONCLUSION: Centrally developed, institution-wide competency assessment uses resources efficiently to relieve individual programs of the need to "reinvent the wheel" and provides program directors and residents with useful information for individual and programmatic review.

Clinical Competence↗

Procedures for establishing defensible absolute passing scores on performance examinations in health professions education.

BACKGROUND: Establishing credible, defensible, and acceptable passing scores for performance-type examinations in real-world settings is a challenge for health professions educators. Our purpose in this article is to provide step-by-step instructions with worked examples for 5 absolute standard-setting methods that can be used to establish acceptable passing scores for performance examinations such as Objective Structured Clinical Examinations or standardized patient encounters. SUMMARY: All standards reflect the subjective opinions of experts. In this "how-to" article, we demonstrate procedures for systematically capturing these expert opinions using 5 research-based methods (Angoff, Ebel, Hofstee, Borderline Group, and Contrasting Groups). We discuss issues relating to selection of judges, use of performance data, and decision-making processes. CONCLUSIONS: Different standard-setting methods produce different passing scores; there is no "gold standard." The key to defensible standards lies in the choice of credible judges and in the use of a systematic approach to collecting their judgments. Ultimately, all standards are policy decisions.

Clinical Competence↗

The effects of violating standard item writing principles on tests and students: the consequences of using flawed test items on achievement examinations in medical education.

The purpose of this research was to study the effects of violations of standard multiple-choice item writing principles on test characteristics, student scores, and pass-fail outcomes. Four basic science examinations, administered to year-one and year-two medical students, were randomly selected for study. Test items were classified as either standard or flawed by three independent raters, blinded to all item performance data. Flawed test questions violated one or more standard principles of effective item writing. Thirty-six to sixty-five percent of the items on the four tests were flawed. Flawed items were 0-15 percentage points more difficult than standard items measuring the same construct. Over all four examinations, 646 (53%) students passed the standard items while 575 (47%) passed the flawed items. The median passing rate difference between flawed and standard items was 3.5 percentage points, but ranged from -1 to 35 percentage points. Item flaws had little effect on test score reliability or other psychometric quality indices. Results showed that flawed multiple-choice test items, which violate well established and evidence-based principles of effective item writing, disadvantage some medical students. Item flaws introduce the systematic error of construct-irrelevant variance to assessments, thereby reducing the validity evidence for examinations and penalizing some examinees.

Choice Behavior↗

Item analysis to improve reliability for an internal medicine undergraduate OSCE.

Utilization of objective structured clinical examinations (OSCEs) for final assessment of medical students in Internal Medicine requires a representative sample of OSCE stations. The reliability and generalizability of OSCE scores provides validity evidence for OSCE scores and supports its contribution to the final clinical grade of medical students. The objective of this study was to perform item analysis using OSCE stations as the unit of analysis and evaluate the extent to which OSCE score reliability can be improved using item analysis data. OSCE scores from eight cohorts of fourth-year medical students (n = 435) in a 6-year undergraduate program were analyzed. Generalizability (G) coefficients of OSCE scores were computed for each cohort. Item analysis was performed by considering each OSCE station as an item and computing the corrected item-total correlation. OSCE stations which negatively impacted the reliability were deleted and the G-coefficient was recalculated. The G-coefficients of OSCE scores from the eight cohorts ranged from 0.48 to 0.80 (median 0.62). The median number of OSCE stations that negatively impacted the G-coefficient was 3.5 (out of a median of 25 total stations). When the ''problem stations'' were deleted, the median G-coefficient across eight cohorts increased to 0.62--0.72. In conclusion, item analysis of OSCE stations is useful and should be performed to improve the reliability of total OSCE scores. Problem stations can then be identified and improved.

Clinical Competence↗

Sources of validity evidence for an internal medicine student evaluation system: an evaluative study of assessment methods.

BACKGROUND: Medical students' final clinical grades in internal medicine are based on the results of multiple assessments that reflect not only the students' knowledge, but also their skills and attitudes. OBJECTIVE: To examine the sources of validity evidence for internal medicine final assessment results comprising scores from 3 evaluations and 2 examinations. METHODS: The final assessment scores of 8 cohorts of Year 4 medical students in a 6-year undergraduate programme were analysed. The final assessment scores consisted of scores in ward evaluations (WEs), preceptor evaluations (PREs), outpatient clinic evaluations (OPCs), general knowledge and problem-solving multiple-choice questions (MCQs), and objective structured clinical examinations (OSCEs). Sources of validity evidence examined were content, response process, internal structure, relationship to other variables, and consequences. RESULTS: The median generalisability coefficient of the OSCEs was 0.62. The internal consistency reliability of the MCQs was 0.84. Scores for OSCEs correlated well with WE, PRE and MCQ scores with observed (disattenuated) correlation of 0.36 (0.77), 0.33 (0.71) and 0.48 (0.69), respectively. Scores for WEs and PREs correlated better with OSCE than MCQ scores. Sources of validity evidence including content, response process, internal structure and relationship to other variables were shown for most components. CONCLUSION: There is sufficient validity evidence to support the utilisation of various types of assessment scores for final clinical grades at the end of an internal medicine rotation. Validity evidence should be examined for any final student evaluation system in order to establish the meaningfulness of the student assessment scores.

Clinical Competence↗

Resident performance on the Council on Resident Education in Obstetrics and Gynecology (CREOG) In-Training Examination: years 1996 through 2002.

OBJECTIVE: This study was undertaken to evaluate the Council on Resident Education in Obstetrics and Gynecology (CREOG) In-Training Examination scores for significant trends. STUDY DESIGN: The percent-correct scores for each of the 6 published examination objectives from 7 consecutive years were analyzed. The data set was analyzed by multivariate analysis of variance using gender, examination year, and postgraduate year as categorical variables, and each year was analyzed separately by gender and postgraduate year. Scores of residents who took the examination for 4 consecutive years were analyzed by using repeated measures analysis of covariance. RESULTS: Variation by examination year appeared random, although scores monotonically increased with postgraduate year for all objectives and all years. The mean relative scores of women were higher than men on the primary/preventive care objective, but the reverse was true for the general considerations objective. CONCLUSION: The CREOG In-Training Examination appears to be a dependable measure of residents' improvement in cognitive knowledge.

Adult↗

Validity threats: overcoming interference with proposed interpretations of assessment data.

CONTEXT: Factors that interfere with the ability to interpret assessment scores or ratings in the proposed manner threaten validity. To be interpreted in a meaningful manner, all assessments in medical education require sound, scientific evidence of validity. PURPOSE: The purpose of this essay is to discuss 2 major threats to validity: construct under-representation (CU) and construct-irrelevant variance (CIV). Examples of each type of threat for written, performance and clinical performance examinations are provided. DISCUSSION: The CU threat to validity refers to undersampling the content domain. Using too few items, cases or clinical performance observations to adequately generalise to the domain represents CU. Variables that systematically (rather than randomly) interfere with the ability to meaningfully interpret scores or ratings represent CIV. Issues such as flawed test items written at inappropriate reading levels or statistically biased questions represent CIV in written tests. For performance examinations, such as standardised patient examinations, flawed cases or cases that are too difficult for student ability contribute CIV to the assessment. For clinical performance data, systematic rater error, such as halo or central tendency error, represents CIV. The term face validity is rejected as representative of any type of legitimate validity evidence, although the fact that the appearance of the assessment may be an important characteristic other than validity is acknowledged. CONCLUSIONS: There are multiple threats to validity in all types of assessment in medical education. Methods to eliminate or control validity threats are suggested.

Bias↗

Toward meaningful evaluation of clinical competence: the role of direct observation in clerkship ratings.

PROBLEM STATEMENT AND PURPOSE: The lack of direct observation by faculty may affect meaningful judgments of clinical competence. The purpose of this study was to explore the influence of direct observation on reliability and validity evidence for family medicine clerkship ratings of clinical performance. METHOD: Preceptors rating family medicine clerks (n = 172) on a 16-item evaluation instrument noted the data-source for each rating: note review, case discussion, and/or direct observation. Mean data-source scores were computed and categorized as low, medium or high, with the high-score group including the most direct observation. Analyses examined the influence of data-source on interrater agreement and associations between clerkship clinical scores (CCS) and scores from the National Board of Medical Examiners (NBME(R)) subject examination as well as a fourth-year standardized patient-based clinical competence examination (M4CCE). RESULTS: Interrater reliability increased as a function of data-source; for the low, medium, and high groups, intraclass correlation coefficients were.29,.50, and.74, respectively. For the high-score group, there were significant positive correlations between CCS and NBME score (r =.311, p =.054); and between CCS and M4CCE (r =.423, p =.009). CONCLUSION: Reliability and validity evidence for clinical competence is enhanced when more direct observation is included as a basis for clerkship ratings.

Certification↗

Reliability: on the reproducibility of assessment data.

CONTEXT: All assessment data, like other scientific experimental data, must be reproducible in order to be meaningfully interpreted. PURPOSE: The purpose of this paper is to discuss applications of reliability to the most common assessment methods in medical education. Typical methods of estimating reliability are discussed intuitively and non-mathematically. SUMMARY: Reliability refers to the consistency of assessment outcomes. The exact type of consistency of greatest interest depends on the type of assessment, its purpose and the consequential use of the data. Written tests of cognitive achievement look to internal test consistency, using estimation methods derived from the test-retest design. Rater-based assessment data, such as ratings of clinical performance on the wards, require interrater consistency or agreement. Objective structured clinical examinations, simulated patient examinations and other performance-type assessments generally require generalisability theory analysis to account for various sources of measurement error in complex designs and to estimate the consistency of the generalisations to a universe or domain of skills. CONCLUSIONS: Reliability is a major source of validity evidence for assessments. Low reliability indicates that large variations in scores can be expected upon retesting. Inconsistent assessment scores are difficult or impossible to interpret meaningfully and thus reduce validity evidence. Reliability coefficients allow the quantification and estimation of the random errors of measurement in assessments, such that overall assessment can be improved.

Bias↗

What is the impact of commercial test preparation courses on medical examination performance?

BACKGROUND: Commercial test preparation courses are part of the fabric of U.S. medical education. They are also big business with 2,000 sales for 1 firm listed at nearly $250 million. This article systematically reviews and evaluates research published in peer-reviewed journals and in the "grey literature" that addresses the impact of commercial test preparation courses on standardized, undergraduate medical examinations. SUMMARY: Thirteen computerized English language databases were searched using 29 search terms and search concepts from their onset to October 1, 2002. Also manually searched was medical education conference proceedings and publications after the end date; and medical education journal editors were contacted about articles accepted for publication, but not yet in print, that were deemed pertinent to this review. Studies that met three criteria were selected: (a) a commercial test preparation course or service was an educational intervention, (b) the outcome variable was one of several standardized medical examinations, and (c) results are published in a peer-reviewed journal or another outlet that insures scholarly scrutiny. The criteria were applied and data extracted by consensus of 2 reviewers. The search identified 11 empirical studies, of which 10 (8 journal articles, 2 unpublished reports) are included in this review. Qualitative data synthesis and tabular presentation of research methods and outcomes are used. CONCLUSION: The articles and unpublished reports reveal that current research lacks control and rigor; the incremental validity of the commercial courses on medical examination performance, if any, is extremely small; and evidence in support of the courses is weak or nonexistent; almost no details are given about the form and conduct of the commercial test preparation courses; studies are confined to courses in preparation for the Medical College Admission Test, the former National Board of Medical Examiners Part 1, and the United States Medical Licensing Examination Step 1, not tests of clinical science; and that cost-benefit analyses of the test preparation courses have not been done. It is concluded that the utility and value of commercial test preparation courses in medicine have not been demonstrated, and that evaluation apprehension in the medical profession and aggressive marketing practices are most likely responsible for commercial course prosperity.

Advertising↗

Item response theory: applications of modern test theory in medical education.

CONTEXT: Item response theory (IRT) measurement models are discussed in the context of their potential usefulness in various medical education settings such as assessment of achievement and evaluation of clinical performance. PURPOSE: The purpose of this article is to compare and contrast IRT measurement with the more familiar classical measurement theory (CMT) and to explore the benefits of IRT applications in typical medical education settings. SUMMARY: CMT, the more common measurement model used in medical education, is straightforward and intuitive. Its limitation is that it is sample-dependent, in that all statistics are confounded with the particular sample of examinees who completed the assessment. Examinee scores from IRT are independent of the particular sample of test questions or assessment stimuli. Also, item characteristics, such as item difficulty, are independent of the particular sample of examinees. The IRT characteristic of invariance permits easy equating of examination scores, which places scores on a constant measurement scale and permits the legitimate comparison of student ability change over time. Three common IRT models and their statistical assumptions are discussed. IRT applications in computer-adaptive testing and as a method useful for adjusting rater error in clinical performance assessments are overviewed. CONCLUSIONS: IRT measurement is a powerful tool used to solve a major problem of CMT, that is, the confounding of examinee ability with item characteristics. IRT measurement addresses important issues in medical education, such as eliminating rater error from performance assessments.

Clinical Competence↗

Establishing passing standards for classroom achievement tests in medical education: a comparative study of four methods.

PURPOSE: The purpose of this research was to evaluate the Direct Borderline standard-setting method, designed for classroom instructor use, and to compare the characteristics of this newer method to three well-established methods. Most standard-setting methods were designed for large-scale assessments, and most research has taken place in the context of high-stakes examinations. METHOD: Four absolute standard-setting methods (Nedelsky, Direct Borderline, Hofstee, and Ebel) were studied for year 1 and 2 basic science examinations. RESULTS: The Direct Borderline method produced passing scores similar to the Nedelsky method and was reproducible. The Hofstee and Ebel methods produced the lowest passing scores. Standard errors at the passing score were the same or lower for the Direct Borderline method compared with the Nedelsky method. CONCLUSIONS: The Direct Borderline method has reasonable psychometric characteristics and may be practical for faculty to use in establishing absolute passing standards for classroom achievement tests.

Achievement↗