Paper Summary

Effects of Rater Monitoring on Rater Bias in Grading Essays on a High-Stakes Licensing Examination

Mon, April 16, 4:05 to 6:05pm, Pan Pacific, Floor: Restaurant Level, Pacific Rim 2

Abstract

This study assessed the consistency of rater grading of essays when rater training includes monitoring as well as initial training to consensus. Essays by 112 bar examiner candidates were evaluated by raters who received pre-training with monitoring after every 12th paper while 191 were graded by raters receiving initial training only. Maximum Effect Sizes (MES) and R-squared values served as criterion measures. Raters with monitoring produced smaller MES and R-square values for 5 of 7 questions than were obtained for values based upon group differences only compared to 1 (MES) or 2 (R-squared) of 7 for the training without monitoring. Monitoring grading of performance tasks can be a useful tool in a high-stakes assessment program for controlling general rater bias

Authors