PubMed HealthSearch

Biomedical subjects

G Regehr

Publications and source records attributed to G Regehr.

At least 19 recordsLinked to original sources

Computer-assisted learning versus a lecture and feedback seminar for teaching a basic surgical technical skill.

BACKGROUND: Rapid improvements in computer technology allow us to consider the use of computer-assisted learning (CAL) for teaching technical skills in surgical training. The objective of this study was to compare in a prospective, randomized fashion, CAL with a lecture and feedback seminar (LFS) for the purpose of teaching a basic surgical skill. METHODS: Freshman medical students were randomly assigned to spend 1 hour in either a CAL or LFS session. Both sessions were designed to teach them to tie a two-handed square knot. Students in both groups were given knot tying boards and those in the CAL group were asked to interact with the CAL program. Students in the LFS group were given a slide presentation and were given individualized feedback as they practiced this skill. At the end of the session the students were videotaped tying two complete knots. The tapes were independently analyzed, in a blinded fashion, by three surgeons. The total time for the task was recorded, the knots were evaluated for squareness, and each subject was scored for the quality of performance. RESULTS: Data from 82 subjects were available for the final analysis. Comparison of the two groups demonstrated no significant difference between the proportion of subjects who were able to tie a square knot. There was no difference between the average time required to perform the task. The CAL group had significantly lower quality of performance (t = 5.37, P <0.0001). CONCLUSIONS: CAL and LFS were equally effective in conveying the cognitive information associated with this skill. However, the significantly lower performance score demonstrates that the students in the CAL group did not attain a proficiency in this skill equal to the students in the LFS group. Comments by the students suggest that the lack of feedback in this model of CAL was the significant difference between these two educational methods.

Computer-Assisted Instruction

Validation of an objective structured clinical examination in psychiatry.

PURPOSE: To examine the validity of a psychiatry clerkship's objective structured clinical examination (OSCE). METHOD: In 1996, 33 clinical clerks and 17 psychiatry residents at the University of Toronto participated in an eight-station OSCE evaluated by psychiatrist-examiners using binary checklists and global ratings. Prior to the OSCE, communication course instructors were asked to rank the clerks on interviewing ability, and faculty supervisors were asked to identify the OSCE stations on which the clerks were likely to do well or poorly. RESULTS: Mean OSCE scores were significantly higher for the residents than for the clerks on global ratings but not on checklists. The communication instructors accurately predicted the clerks' rankings on the global scores but not their scores on the checklists. The faculty supervisors predicted with moderate accuracy the clerks' success on the OSCE stations as measured by the checklists but not by the global ratings. The residents rated the OSCE scenarios as highly realistic. CONCLUSIONS: The evidence of construct and concurrent validity together with high ratings of realism suggest that a psychiatry OSCE can be a valid assessment of clerks' clinical competence.

Clinical Clerkship

Comparing the psychometric properties of checklists and global rating scales for assessing performance on an OSCE-format examination.

PURPOSE: To compare the psychometric properties of checklists, global rating scales preceded by a checklist, and global rating scales alone in assessing surgery residents' performances on an OSCE-like technical skills examination. METHOD: In 1996, 53 general surgery residents with one to six years of postgraduate training participated in a performance-based examination of technical skills consisting of eight 15-minute stations (bench-model simulations of operative procedures in general surgery). Two qualified surgeons marked at each station, one using a task-specific checklist (C) and a subsequent global rating scale (Gc), the other using a global rating scale only (G). RESULTS: Interstation reliabilities measured by Cronbach's alpha were .79 for C, .89 for Gc, and .85 for G. A series of multiple regressions predicting level of training from test scores revealed an R2 of .584 for C alone, which increased to .711 when Gc was entered after (p < .001), and increased to .704 when G was entered after C (p < .001). However, R2 for Gc alone was .711, and for G alone was .704, neither of which changed when C was entered into the prediction (p > .10). The R2 for Gc and G predicting level of training (.725) was not significantly greater than that of either Gc or G alone. A very similar pattern of results was seen when C, Gc, and G were used to predict independent evaluations of the operative outcomes. CONCLUSIONS: Global rating scales scored by experts showed higher inter-station reliability, better construct validity, and better concurrent validity than did checklists. Further, the presence of the checklists did not improve the reliability or validity of the global rating scale over that of the global rating scale alone. These results suggest that global rating scales administered by experts are a more appropriate summative measure when assessing candidates on performance-based examinations.

Educational Measurement

Using videotaped benchmarks to improve the self-assessment ability of family practice residents.

PURPOSE: To address methodologic and statistical problems of previous studies of self-assessment by exposing participants to relevant standards, anchoring rating scales, and providing practice in the use of the assessment tool. METHOD: Fifty first- and second-year family practice residents performed a ten-minute patient interview with a difficult communication problem. Following each interview, the resident and two experts independently evaluated the resident's communication skills. The resident was then shown a videotape of four performances (ranging in quality from poor to good) of the same scenario. The resident evaluated the communication skills displayed in each performance and then reevaluated his or her own performance. RESULTS: The correlation between experts' evaluations and residents' self-evaluations was moderate immediately after the interview (r = 0.38) but increased significantly after the residents viewed the videotape (r = 0.52). This effect was more pronounced for first-year residents (0.22 to 0.45) than for second-year residents (0.53 to 0.65), although the difference was not significant. Post-hoc analysis revealed that neither initial nor post-benchmark self-assessment ability was related to the ability to accurately evaluate the benchmarks in a manner consistent with the experts. CONCLUSIONS: The ability to self-assess does not seem strongly tied to the ability to assess the performances of others on the same task. Nonetheless, providing a set of benchmarks against which trainees can compare their own performances improves their ability to self-evaluate even if the qualities of the benchmarks are not explicitly identified.

Adult

The integration of child psychiatry into a psychiatry clerkship OSCE.

OBJECTIVE: To integrate child psychiatry into a psychiatry clerkship OBJECTIVE Structured Clinical Examination (OSCE). METHOD: Child psychiatry OSCE stations were designed to evaluate clerk' skills in the identification of 4 common conditions. Child psychiatrists wrote case scenarios and checklists and supported standardized patient (SP) training for the stations. A bank of 4 child psychiatry OSCE stations is now available for use in the psychiatry OSCE. Child psychiatry faculty have been trained as examiners for ongoing administration of his OSCE. RESULTS: This bank of child psychiatry OSCE stations has examined 402 clerks. Mean student scores for content were 68% to 86% and for process were 69% to 76%. Station reliability and examiner feedback were acceptable. CONCLUSIONS: Child psychiatry has been successfully integrated into a psychiatry clerkship OSCE. Although the commitment in terms of monetary and faculty costs has been considerable, the accompanying educational benefits of such integration warranted this expense.

Child

A model for predicting depression in victims of rape.

This article proposes a model for understanding the factors contributing to long-standing depression in women who have been raped. A path analysis of data obtained from 71 women who had been raped revealed that women with generalized beliefs that they could not control events in their lives were more likely to attribute responsibility for their rape to permanent intrapsychic factors and were more likely to be depressed. Women who perceived that they had higher levels of internal control tended to have higher levels of education, were more likely to be employed, and were less likely to be depressed more than one year after having been raped. Childhood sexual abuse was not associated with internal control or attributions of causality or depression in this analysis. Implications for the determination of prognosis and treatment recommendations in civil litigation assessments are discussed.

Adolescent

Testing technical skill via an innovative "bench station" examination.

BACKGROUND: A new approach to testing operative technical skills, the Objective Structured Assessment of Technical Skill (OSATS), formally assesses discrete segments of surgical tasks using bench model simulations. This study examines the interstation reliability and construct validity of a large-scale administration of the OSATS. METHODS: A 2-hour, eight-station OSATS was administered to 48 general surgery residents. Residents were assessed at each station by one of 48 surgeons who evaluated the resident using two methods of scoring: task-specific checklists and global rating scales. RESULTS: Interstation reliability was 0.78 for the checklist score, and 0.85 for the global score. Analysis of variance revealed a significant effect of training for both the checklist score, F(3,44) = 20.08, P <0.001, and the global score, F(3,44) = 24.63, P <0.001. CONCLUSIONS: The OSATS demonstrates high reliability and construct validity, suggesting that we can effectively measure residents' technical ability outside the operating room using bench model simulations.

Clinical Competence

A new assessment tool: the patient assessment and management examination.

BACKGROUND: The major goal of certification is to assure the public that the candidate is competent in all facets required of the position. The patient assessment and management examination (PAME) was developed to enable a more comprehensive assessment of competence in the practice of surgery. METHODS: A six-station, 3-hour, standardized-patient-based evaluation was developed. Each station was scored using a set of five-point global rating scales. PAME results were compared to the last two in training evaluation reports (ITER), the clinical knowledge component of the ITER (ITER-CK), an in-house oral examination (OE), and the Canadian Association of General Surgeons' multiple-choice examination (CAGS). RESULTS: Eighteen senior general surgery residents were evaluated. Overall reliability was 0.70 (Cronbach's alpha). Fifth-year residents scored significantly better than fourth-year residents (t = 3.062; p = 0.0074), with 1 year of training accounting for 37% of the variance in scores. Correlations between the PAME and each of the other measures were ITER, 0.24; ITER-CK, 0.38; OE, -0.13; and CAGS, 0.061, with the PAME demonstrating better reliability and stronger evidence of validity than any other. CONCLUSIONS: The PAME had better psychometric properties than other measures and assessed areas often not evaluated. This type of evaluation may be useful for feedback, remediation, or certification decisions.

Adult

Methodological problems in the retrospective computation of responsiveness to change: the lesson of Cronbach.

OBJECTIVE: To examine the relation between responsiveness coefficients derived directly from a calculation of average change resulting from a treatment intervention (Responsiveness-Treatment or RT) and those derived from retrospective analysis of changed and unchanged groups (Responsiveness Retrospective or RR) based on a global measure of change. METHOD: Two approaches were used. First, we used simulation methods to examine the analytical relationship between the RT and RR coefficients. We then located eight studies where it was possible to compute both RT and RR coefficients. As anticipated from theoretical arguments, the RR coefficients were larger than the RT coefficients (1.50 versus 0.41, p < .0001). Within study there was no predictable relationship between the two indices. Across studies, the magnitude of the RR coefficient was strongly related to the correlation with the retrospective global scale, and unrelated to the magnitude of the RT coefficient. The simulated curves fit well with the observed data, and substantiated the observation that the relation between RT and RR coefficients is complex and only weakly related to the size of the treatment effect. CONCLUSION: Retrospective methods of computing responsiveness yield little information about the ability of an instrument to detect treatment effects, and should not be used as a basis for choice of an instrument for applications to clinical trials.

Computer Simulation

Objective structured assessment of technical skill (OSATS) for surgical residents.

BACKGROUND: The technical skill of surgical trainees is not well assessed. This study aimed (1) to compare the reliability of three scoring systems, (2) to compare live and bench formats and (3) to assess construct validity of a test of operative skill. METHODS: Parallel examinations of operative skill, one using live animals and one using simulations, were developed. Performance was graded using operation-specific checklists, detailed global rating forms and pass/fail judgements. Twenty surgical residents each took both formats. RESULTS: Disattenuated correlations between live and bench scores were high (0.69-0.72). Mean interrater reliability across stations ranged from 0.64 to 0.72. Internal consistency was moderate to high (alpha: 0.61-0.74) for the live format using the checklist and for live and bench formats using global ratings. Global ratings discriminated between resident levels for both formats (bench: F(2,17) = 4.45, P < 0.05; live: F(2,17) = 3.55, P < 0.05), checklists did not. CONCLUSION: This preliminary study suggests that the Objective Structured Assessment of Technical Skill can reliably and validly assess surgical skills. Global ratings are a better method of assessment than task-specific checklists. Bench model simulation gives equivalent results to use of live animals for this test format.

Clinical Competence

An objective structured clinical examination for evaluating psychiatric clinical clerks.

PURPOSE: To assess the feasibility, reliability, and validity of an objective structured clinical examination (OSCE) for psychiatric clinical clerks. METHOD: In 1995 two parallel forms of a ten-station OSCE (eight clinical stations, two writing stations) were developed at the University of Toronto Faculty of Medicine Each 12-minute performance-based clinical station was assessed by a faculty psychiatrist using both a checklist for each student's performance content and a global-rating scale of the performance process. The students' clinical-station scores were calculated as the average of their content and process scores (expressed as percentages). Examiners also recorded an overall judgment of each students' performance (pass, borderline, or fail) and wrote [in collaboration with the standardized patient (SP) at that station] comments on each student's performance. There were two criteria for a passing grade: a total mark of 60% or higher across all ten stations and a "pass" or "borderline" mark in at least five of the eight clinical stations. Each OSCE form was administered three times. RESULTS: The first form was used to examine 94 clerks, the second form to examine 98 clerks. The students' mean scores for the two forms were 70.47% (SD, 6.33%) and 67.66% (SD, 7.05%), respectively. In addition to the standard evaluation information collected on the students, several critical incidents occurred (e.g., a student's loss of control of emotions) that may identify potential problems in professional conduct. The direct cost for one administration of the examination was approximately Can$3,300: the largest portion of this was for the SPs' time spent in training and performing their roles. CONCLUSION: Preliminary evidence suggests that a psychiatry OSCE is feasible for assessing complex psychiatric skills. However, careful attention must be paid to SP training, examination monitoring, detection of critical incidents, and provision of feedback to students, faculty, and SPs. The university's previous system of oral examinations required approximately 600 faculty hours per year. The OSCE requires approximately 450 faculty hours, and the 150 hours saved almost cover the Can$20,000 that the examination costs each year. In all, the OSCE is an evaluation system that has demonstrable reliability and is more enjoyable for both the faculty and the students.

Clinical Clerkship

Who should rate candidates in an objective structured clinical examination?

PURPOSE: To determine who is the better rater of history taking in an objective structured clinical examination (OSCE): a physician or a standardized patient (SP). METHOD: During the 1991 pilot administration of an OSCE for the Medical Council of Canada's qualifying examination, five history-taking stations were videotaped. Candidates at these stations were scored by three raters: a physician (MD), an SP observer (SPO), and an SP rating from recall (SPR). To determine the validity of each rater's scores, these scores were compared with a "gold standard", which was the average of videotape ratings by three physicians, each scoring independently. Analysis included both correlations with the standard and a repeated-measures analysis of variance (ANOVA) comparing raters' mean scores on each station with mean scores of the gold standard. RESULTS: Ninety-one videotapes were scored by the "gold-standard" physicians. Correlations with the standard showed no clear preference for MD, SPO, or SPR raters. ANOVAs revealed significant differences from the standard on three stations for the SPR, two stations for the SPO, and one stations for the MD. CONCLUSIONS: An MD rater is less likely to differ from a standard established by a consensus of MD ratings than are SP raters rating from recall. If an MD cannot be used, an SP observer is preferable to an SP rating from recall.

Analysis of Variance

Issues in cognitive psychology: implications for professional education.

Education and cognitive psychology have tended to pursue parallel rather than overlapping paths. Yet there is, or should be, considerable common ground, since both have major interests in learning and memory. This paper presents a number of topics in cognitive psychology, summarizes the findings in the field, and explores the implications for teaching and learning. THE ORGANIZATION OF LONG-TERM MEMORY: The acquisition of expertise in an area can be characterized by the development of idiosyncratic memory structures called semantic networks, which are meaningful sets of connections among abstract concepts and/or specific experiences. Information (such as the assumptions and hypotheses that are necessary to diagnose and manage cases) is retrieved through the activation of these networks. Thus, when teaching, new information must be embedded meaningfully in relevant, previously existing knowledge to ensure that it will be retrievable when necessary. INFLUENCES ON STORAGE AND RETRIEVAL FROM MEMORY: A wide variety of variables affect the capacity to store and retrieve information from memory, including meaning, the context and manner in which information is learned, and relevant practice in retrieval. Educational strategies must, therefore, be directed at three goals--to enhance meaning, to reduce dependence on context, and to provide repeated relevant practice in retrieving information. PROBLEM SOLVING AND TRANSFER: Much of the development of expertise involves the transition from using general problem-solving routines to using specialized knowledge that reduces the need for classic "problem solving." Two manifestations of this specialized knowledge are the use of analogy and the specialization of general routines in specific domains. To develop these specialized forms of knowledge, the learner must have extensive practice in using relevant problem-solving routines and in identifying the situations in which a particular routine is likely to be useful. CONCEPT FORMATION: Experts possess both abstract proto-typical information about categories and an extensive set of separate, specific examples of categories, which have been obtained through individual experience. Both these sources of information are used in categorization and diagnostic classifications. Thus, it is important for educators to be aware that experience with sample cases is not just an opportunity to apply and practice the rules "at the end of the chapter." Instead, experience with cases provides an alternative method of reasoning that is independent of, but equally useful to, analytical rules. DECISION MAKING: Experts clearly do not use classic formal decision theory, but rather make use of heuristics, or shortcuts, when making decisions. Nonetheless, experts generally make appropriate decisions. This suggests that the shortcuts are useful more often than not. Rather than teaching learners to avoid heuristics, then, it might be more reasonable to help them recognize those relatively infrequent situations where their heuristics are likely to fail.

Cognition