Paper Summary
Share...

Direct link:

Evaluating Linking Errors for a Large-Scale High-Stakes Clinical Skills Licensing Exam

Sat, April 9, 4:05 to 5:35pm, Convention Center, Floor: Level One, Room 143 B

Abstract

Perspective: Clinical skills exams have been widely used as part of medical licensure in North America. However, very few programs have conducted formal linking/equating procedures to establish score comparability over time. Furthermore, little is known about the accuracy of the linking/equating in this context. An objective structured clinical examination (OSCE) is currently used to assess internationally trained medical graduates’ readiness for entry into Canadian residency programs. It consists of ten OSCE stations depicting various clinical scenarios which are portrayed by standardized patients (SPs). The performances of candidates on these stations are rated by physician examiners on nine clinical competencies using Likert scales. Multiple test forms are administered within a given year, for security reasons. To assure score comparability across forms, a linking procedure was implemented to account for form difficulty using Levine’s observed score linear equating under a common-item nonequivalent groups design. Candidate scores were linked through anchor stations to the base form to enable score comparison and application of the same pass mark across forms.
Objective: The accuracy of the linking procedure for OSCEs could be affected by various random or systematic errors. Standard errors (SEs) of linking can be estimated using a bootstrap method. This method has been used in many studies to examine the SEs of linking/equating, but applications with performance assessments are considerably rarer. The purpose of this study was to estimate the SEs of linking for the OSCE described above using the bootstrap method with small sample real examinee data.
Methods and Data Sources: Data from four test forms were used. Each form was linked through a different set of anchor stations to the base form. For bootstrapping, a 2x2 design was used for each form under the following conditions:
• 2 data conditions:
o Unpurified (including all examinees)
o Purified (excluding repeaters and other non-targeted examinees)

• 2 bootstrap sample sizes:
o N=a random sample of 100 candidates from each form was used for bootstrapping.
o N=size of the operational data
 Unpurified: N equals 315, 514, 601, 262, 309 for Base Form, Forms A, B, C, D, respectively
 Purified: N equals 196, 280, 337, 134, 149 for Base Form, Forms A, B, C, D, respectively
For each condition, a pair of bootstrap samples was drawn with replacement from the forms to be linked and linking was subsequently performed. This process was replicated 10,000 times. The SEs of linking were calculated as the standard deviation of the linked scores over replications.
Results: Preliminary results indicate that SEs are smaller near the mean than at the extremes and lower for unpurified data. The effects of the data condition and the bootstrap sample size on mean SEs of linking were compared using ANOVA.
Significance: When multiple test forms are used for making high-stakes decisions, the accuracy of linking is critical to ensure that scores have the same meaning, regardless of when the test form was administered or which form a candidate completed. The implications of our findings for high-stakes performance assessments will be further discussed in the full paper.

Author