Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
Objectives or purposes
Evidence from a handful of large-scale studies suggests that although observers can be trained to score reliably using observation protocols, there are concerns related to initial training and calibration activities designed to keep observers scoring accurately over time (e.g., Bell, et al, 2012; BMGF, 2012). Studies offer little insight into how educational practitioners understand and score observation protocols. This lack of clarity on the factors that facilitate and constrain educators’ learning and use of observation systems makes it difficult to implement training and quality control processes at scale. This study of the use of observation protocols describes which aspects of teaching are easiest and hardest for observers to learn and uses think aloud data to advance hypotheses about why those patterns exist.
Perspective
The current generation of protocols requires observers to make holistic judgments about classroom interactions and their meaning. Protocols typically require observers to watch for many different features of interaction, creating a challenge in parsing the incoming stream of information for key elements for focus. Further, because classroom interactions vary in so many ways, observers must learn to apply scoring criteria to a wide range of lesson topics, subject areas, and grade levels. Finally, observers are not blank slates. Most observers are former teachers and have ideas about what counts as high quality teaching and learning (Lortie, 1975). This study takes the perspective that scoring observations of teacher-student interactions is a complex socio-cognitive process that we are only beginning to understand.
Methods and Data Sources
Drawing on data from the Measures of Effective Teaching Project and the Understanding Teaching Quality study, we investigate the reliability of observation scores on four observation protocols. Scores were created by 243 observers across almost 2,000 teachers. We combine these analyses (e.g., percent agreement, correlations, and kappa) with grounded theory analyses (Strauss & Corbin, 1998) of think aloud and interview data from 13 observers and master observers. The think aloud sessions required all observers to watch and score a lesson per the protocol, talking aloud about their thinking during their scoring, and answer a standardized set of clarifying questions after the scoring was complete.
Results
Results suggest observers struggled to agree with each other and master observers on a number of dimensions across the observation protocols. In general, reliability was below 80% exact agreement. Observers had the worst reliability on instructional and emotional dimensions of teaching and the best reliability on classroom organization and management dimensions. Think aloud data suggest these difficulties come from poor observer confidence, the use of strategies for scoring that differ from master observers, confusion about how to apply specific scoring rules, and foundational misunderstandings of core protocol constructs (e.g., student understanding, etc.).
Significance of the work
This study begins to nominate areas of teaching that observers struggle with and potential reasons they may struggle with those areas. These findings contribute both to our understanding of the validity of observation scores and how to improve the training of observers.
Courtney A. Bell, ETS
Yi Qi, Educational Testing Service
Andrew J. Croft, ETS
Dawn Leusner, ETS
Daniel F. McCaffrey, ETS
Drew H. Gitomer, Rutgers University - New Brunswick/Piscataway
Robert Pianta, University of Virginia