Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Register for SRCD21
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Classroom observations play an increasingly important role in decision-making in developmental settings, in particular in educational contexts, ranging from impact evaluation of intervention and prevention efforts to decisions about hiring, firing, and funding in school, state, and federal accountability systems. Despite the growing reliance on classroom observations, more knowledge is needed to make strategic design and methodological decisions (e.g., determine number of raters) that will generate precise and reliable scores. These strategic decisions require information that enables users to parse the relative benefit of each component while also considering feasibility (i.e., time, cost).
Researchers typically rely on Cronbach’s alpha as a “catch all” reliability statistic to determine the precision of scores. However, Cronbach’s alpha only considers one source of variance, items, even though many factors determine the precision of reliability estimates including the number of items, raters, and occasions. In consequence, Cronbach’s alpha systematically overestimates reliability, and does not provide key information needed to improve the precision of classroom observation systems (Hill et al., 2012).
This paper uses Generalizability Theory (Brennan, 2001) to estimate the reliability and precision of a classroom observation tool, the Adapted Teaching Styles Rating Scale (A-TSRS). The A-TSRS captures classroom structure and teaching practices that support children’s social and emotional development. For this paper, A-TSRS data were collected as part of a larger intervention study during the fall of 2017. The study took place in six charter schools in a mid-size city serving primarily low-income Latinx students and included 65 pre-Kindergarten through fourth grade classrooms. In each classroom, each rater completed the A-TSRS two times, referred to below as cycles. This means that each teacher received two scores on the A-TSRS during one observation period. Two raters were present for a random 20% of observations.
We began our analysis of precision and reliability with a mixed model to decompose different sources of variance that comprise the observed score. That information illuminates aspects of the design that contribute the most to imprecision, or variance in the estimates. Table 1 presents estimates for eleven sources of variance in observational ratings for the A-TSRS. One key finding is that the most variance is attributed to items, meaning that some items are more or less agreeable than others (i.e., tend to be scored higher or lower), regardless of person, cycle, or rater. Using the variance estimates, combinations of raters, items, and cycles were then modeled in order to assess precision for up to three raters cycles, referred to as the Decision Study (see Figure 1 below for estimates with 10 items). Holding items constant, the Decision Study indicates a significant gain in precision achieved from adding just one additional cycle of observation.
The final presentation will provide additional details about the multi-phase analysis process. In addition to presenting the full results, we will underscore the need to be clear about the design elements that drive precision, particularly when using tools for high stakes purposes and will revisit how Generalizability Theory can inform strategic decision-making in developmental settings and policy contexts.