Paper Summary
Share...

Direct link:

Rater Drift in Classroom Observation Scores

Mon, April 7, 2:15 to 3:45pm, Convention Center, Floor: 100 Level, 116

Abstract

Objective
The objective of this study is to analyze and understand drift in ratings of classrooms in the Understanding Teaching Quality (UTQ) study, and provide implications for designs collecting ratings over time.

Theoretical Framework
The variance decompositions performed to date on classroom observation ratings have indicated relatively large contributions of both raters and individual rating events to errors in scores, and studies in this and other contexts have indicated that raters drift in overall severity. We need further understanding of the characteristics and relative contributions of these errors to support the improvement of future rating designs in classroom observation systems for research and practice.

Data
The UTQ study took place in middle schools in three large school systems from the same metropolitan area in the southeastern United States. It includes 231 mathematics and 227 English language arts teachers of sixth, seventh, or eight graders. Two classrooms (or sections) of students were observed for each of two lessons for each teacher for a total of 916 sections and 1,829 lessons scored. Thirty-four percent of the sections were grade 6, 29 were grade 7, 36 were grade 8 and the rest were mixed grade. Our analysis of rater drift used scores on the Classroom Assessment Scoring System-Secondary (CLASS-S) and Framework for Teaching (FFT) protocols from both live and video observations taken over a two-year observation and scoring period.

Methods
We fit a series of generalizability ā€œGā€ study models augmented to capture instructional day trends (changes in teaching) and scoring day trends (changes in raters). We used model estimates to plot rater scoring trends and applied a model-based approach to decompose the variance in the scores. With these methods we investigated the following research questions with respect to the CLASS-S and FFT:
1) What are the characteristics of overall time trends in ratings in the UTQ study?
2) How much do raters differ in their severity, and does this change over time?
3) How much does rater severity drift contribute to error in ratings?

Results
We found significant rater drift in observation scores. Raters start out being very lenient, giving relatively high scores but quickly become more severe lowering their scores on average. Drift was smallest for evaluations of classroom management but scores in all domains showed some amount of drift during the entire two years of scoring. Raters did not converge to a common level of scoring for any domain or overall score. In fact, for CLASS-S the variability among raters was increasing during the many months of video scoring. While instructional day trends were negligible, rater drift contributed up to 18% of the variance in scores.

Significance of Study
Observations of teaching are essential for research on teaching and many interventions that are mediated by classroom practice. UTQ has shown that rater drift could be a substantial source of error in such studies. Studies that confound rater experience with assessments of individual teachers will have extra error and may make biased conclusions about some teachers.

Authors