Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Browse Posters
Search Tips
Register for SRCD23
Personal Schedule
Welcome Letter
Program Guide
Change Preferences / Time Zone
Sign In
Introduction: As part of a larger study examining how teacher behaviors affect student bullying and defending behaviors, the Classroom Assessment Scoring System – Upper Elementary (CLASS, Pianta et al., 2008) was used to augment teacher and student survey data. CLASS is designed for observers to rate the teacher-student interactions and has become a popular tool in various fields. However, previous findings have questioned whether CLASS can effectively capture the quality of interactions due to rating differences among raters, which could cause small associations between CLASS and children’s outcomes (Perlman et al., 2016). This study posed the following questions: How do the observation results compare to the data in the CLASS manual? What is the correlation between the ratings in spring and fall? How much do raters differ in each domain and dimension?
Method: Ten trained and certified raters used the CLASS Upper Elementary protocol to rate 69 fourth and fifth grade classrooms in four diverse school districts in the southwest US in fall and spring 2021-2022. Each classroom was observed by two raters four times in one school year. Observations consisted of two to six cycles (20 minutes observation, 10 minutes coding) (M = 3.86, SD=.78, n = 381). Raters were randomly assigned to match rater availability with school schedule, so rater pairs varied across classes.
Results: Table 1 compares the data reported in the CLASS manual with results obtained in the current study. CFA identified the same three factors as those in the manual, but our fit statistics were not as good a fit to the data. Reliability of all dimensions was assessed with the ICC (.55 to 1.00) and Linear Weighted Kappa for each domain (.22 to .70). The percentage of rater pairs whose average scores for each dimension and domain were within 1 point of each other, ranged from 62% to 98%. Mean ratings were calculated for each dimension and domain at both time points. Correlations between fall and spring ranged from .21 to .63.
To assess rater effects, average scores were computed for each dimension and domain for each rater using all data, as well as fall and spring separately. From these data, within teacher Z-scores were calculated. Figure 1 presents those results for the full dataset. ANOVA analyses revealed significant differences among raters (p < .05) on all but two scores. Post hoc analyses using Tukey’s HSD revealed that Observer #3 rated significantly lower than all other raters on eight scales, and Observer #10 was significantly higher on five. Overall, scores were significantly different from others on each dimension/domain.
Discussion:
Our findings revealed areas of concern in inter-rater reliability, inter-rater agreement, and factor loadings. Although raters were trained and certified, results of the current study detected significant differences in their ratings on most dimensions and domains, suggesting there are potential rater effects; steps to understand and reduce bias in rater scores are needed. Although the CLASS has been widely used with strong validity evidence, the current findings indicate that potential rater effect/bias should be evaluated.