Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
Although there are ways to provide automatic scoring when item parameters are known prior to testing, this is impossible for many exams. For example, some vendors may not support real time IRT scoring functionality. Additionally, item parameters may not be available prior to testing due to limitations in how many items can be pretested. In these situations, using IRT parameters to provide immediate feedback is infeasible, thus any immediate feedback provided to examinees will need to be provided using raw scores.
Using simple number correct scores to score exams can be problematic as the raw passing score can fluctuate from exam form to exam form (Crocker & Algina, 1986). For example, the raw score needed to pass Form 1 could be a 45. However, Form 2 may be more difficult, so the raw score needed to pass the exam would be lower, say a 42. Thus, across multiple forms of a test a range of raw scores map onto the passing score. Typically, IRT is used to equate the two forms eliminating this issue (Crocker & Algina, 1986; Kolen & Brennan 2014; Thissen & Wainer, 2001), but absent known IRT parameters immediate IRT scoring is not possible.
One solution to this problem is to use the current item pool to build several exams and find a range of raw scores that could potentially map onto the passing score. This range of raw scores would represent a range of indeterminacy where examinees have to wait for official scoring. Examinees scoring higher than this range will receive an immediate pass, and examinees scoring below this range will receive an immediate fail. Then, several ranges could be built to determine the classification accuracy and the percent of examinees that would not receive immediate scoring. This paper demonstrates how to implement and evaluate this method and provides recommendations for use and possible policy implications. Additionally, some research will be presented on the impact of implementing this method in a credentialing exam as well as lessons learned from using the method.
Whitney Smiley, American Board of Internal Medicine
Jerome C Clauser, American Board of Internal Medicine