Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
Rater-mediated assessments draw on raters to evaluate latent constructs of interest. To arrive at such judgments, raters are usually tasked with interpreting and discriminating among different responses, behaviors, interactions or other types of indicators believed to reflect the latent constructs of interest. An important assumption underlying meaningful comparisons in these types of assessments is that measurement is invariant across raters. Measurement invariance across raters suggests that raters use indicators in similar ways so that the relationships between a latent trait and its indicators do not depend on which rater conducted the evaluation. When measurement is invariant across raters, judgments or scores given by different raters can be placed on a common scale and meaningful comparisons concerning the latent trait can be made.
In this study, I developed and investigated a set of psychometric methods that accommodate, as best as possible, measurement noninvariance among raters. To accommodate rater noninvariance, I drew on the concept of approximate measurement invariance and developed a cross-classified multilevel random item effects item response model. The approach relaxes assumptions of measurement invariance across raters by explicitly adjusting for the different ways in which raters use instruments to establish an approximately invariant scale. Using the proposed framework, I investigated the plausibility of rater measurement invariance in classroom assessments of teaching quality using data from the Measures of Effective Teaching study (MET) and the Mathematical Quality of Instruction (MQI) and Framework for Teaching (FFT) instruments. In turn, I examined three primary questions: (1) To what extent is measurement invariance among raters plausible in rater-mediated assessments of teaching? (2) To what does accommodating measurement noninvariance through random item effects improve the fit of the model to the data? (3) To what extent does adopting an approximate measurement invariance approach improve the predictive validity of the teaching observation scores as compared to more conventional approaches?