Paper Summary

Cross-Classified Random-Effects Models for Assessing Rater Severity and Differential Rater Functioning

Sat, April 14, 8:15 to 9:45am, Marriott Pinnacle, Floor: Fourth Level, Ambleside

Abstract

Previous research has demonstrated use of the cross-classified random effects model (CCrem; Goldstein, 2010) to model the cross-classification of item scores by examinees and items (Goldstein, 2010; Van den Noortgate, De Boeck, &Meulders, 2003). And other researchers have implicitly (e.g., Mariano & Junker, 2007; Patz, Junker, Johnson, & Mariano, 2002; Wang & Wilson, 2005) and explicitly (Beretvas, 2009; Verhelst & Verstralen, 2001) explained how a CCrem could be used to model items’ difficulty, person ability and rater severity. In her paper, Beretvas (2009) detailed some of the benefits of modeling raters’ effects as random rather than fixed. Beretvas also distinguished raters’ data scenarios in which the CCrem versus the cross-classified multiple membership random effects model (CCMMrem; Rasbash & Browne, 2001) might be preferred.

The current study extends Beretvas’ (2009) study by demonstrating when and how to use the cross-classified random effects model (CCrem) and the cross-classified multiple membership random effects model (CCMMrem) for assessing differential rater functioning (DRF). Distinctions are made between scenarios when the CCrem versus the CCMMrem are necessary. The equivalence of the more conventional fixed effects model for raters with the basic CCrem model is demonstrated. Benefits of the random-effects over the fixed-effects parameterization are detailed.

Real and simulated data will be used to demonstrate use of the different models in the following set of scenarios entailing multiple rated item scores per participant: (a) one score for each item rated by one rater with the same rater for all items per participant; (b) one score for each item rated by one rater with different raters for different items per participant; (c) one score for each item representing the overall score from multiple raters per item with the same set of raters for each item per participant; (d) one score for each item representing the overall score from multiple raters with different possible sets of raters for each item per participant; (e) multiple scores for each item representing the score of each of multiple raters with the same set of raters for each item per participant; and (f) multiple scores for each item representing the score of each of multiple raters with different sets of raters for each item per participant.

For each scenario, parameterization of the appropriate model necessary for obtaining each of the following parameters: person ability, item difficulty and potential DRF will be provided. The paper will focus on interpretation of the relevant parameters, data setup for estimation of the models. The paper will also provide well-commented BUGS code necessary for estimation of each model.

Author