Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
Recently, interest in using standard classroom observation protocols to evaluate teaching has grown substantially. Observation scores are ordinal and hierarchical with one or more ratings observed for a teacher on each of multiple lessons. The constructs of teaching may differ at teacher- or lesson-levels. Correlation in measurement errors for multiple dimensions may also occur because raters rate each of the multiple dimensions of an observation protocol simultaneously, which may confound the analyses of a measure’s underlying structure. The use of a hierarchical approach to examine classroom observation protocol data enables more accurate analyses of the factor structure of observation scores.
The purpose of this study is to illustrate the use of two latent, hierarchical approaches (one that considers the ordinal nature of the data and one that treats ordinal data as continuous) to analyze the factor structure of the Classroom Assessment Scoring System – Secondary (CLASS-S). We compared the results of these (i.e., the multivariate hierarchical ordinal-response model and the multivariate hierarchical continuous response) models to a standard approach consisting of averaging ratings from different lessons to create teacher-level dimension scores. We used data from a) the Understanding Teacher Quality study conducted with three large middle school systems in southeastern United States with 458 math or language arts teachers, which were primarily female and white, and b) the Toward an Understanding of Classroom Context study conducted with 82 teachers from middle and high schools of an urban fringe mid-Atlantic school district with teachers and students that were primarily black. In the two samples, about 50% of students were eligible for free or reduced price meals.
Study results revealed that ignoring data nesting had a larger impact on results from data analyses than whether data was treated as ordinal or continuous; which yielded similar results at all data levels. The use of a hierarchical model revealed a different (one) factor structure than that the three-factor structure derived from the standard approach. Study results suggest that uncovering the structure in teaching requires separating rating errors, which are large relative to other sources of variance. Rating errors appear to spillover to analyses conducted at the lesson- or teacher-levels if ratings are averaged. Rating errors (e.g., due to the background of the raters, prior experience, and difficulty in understanding and applying the protocol) may have occurred despite raters receiving substantial training and attending calibration exercises before and during the study (Bejar, 2012). Study findings imply that a total rather than a domain score should be used to represent the quality of classroom interaction because separate scores offer little distinct information about teaching, and variation across domains may be driven by rater error rather than systematic differences in the practices of teachers.
This study makes a significant contribution to the field because it is the first study to use a hierarchical model to analyze the factor structure of classroom observation protocol data. Previous studies have been conducted using standard approaches, which have led to incorrect results due to ignoring data nesting.
Daniel McCaffrey, RAND Corporation
Kun Yuan, RAND Corporation
Terrance D. Savitsky
John H. Lockwood, ETS
Maria O. Edelen