Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
What to do in Chicago
Personal Schedule
Sign In
X (Twitter)
A consensus has developed that any high-quality evaluation system requires the use of multiple measures. While a main goal of such a system is equity, it can be a substantial challenge to implement such a system fairly. We examine the components of the multiple measures system from a large urban school district, which includes principal and peer observation ratings (based on the Charlotte Danielson Framework) and value-added ratings. This work is especially relevant as the implementation of statewide evaluation systems accelerates. Evaluation systems that do not account for differences in rater severity, timing, and other factors may yield different results from systems that do account for these factors. In addition, systems that do not report the precision or reliability of the observational measures, may be omitting useful information that could be factored into a summative evaluation.
The evaluation data used in this analysis includes ratings from multiple observers at different times. The district allows peer and principal observers to create a holistic evaluation independently in the spring, resulting in an overall score in each subdomain; these ratings are combined to yield an overall rating according to a weighted average specified by the district. An alternative method of combining this observational data is to use a multi-faceted Rasch model. In this process, each component has its own effect on the overall rating; we can therefore report a measure of difficulty or leniency for each. We explore components such as teacher, observer, time of year, whether the observation was formal or informal, whether it was performed by a peer or principal, the teacher’s subject area or grade level, and the subdomain, and explore whether better results are achieved through unified or separate models.
Using the multi-faceted Rasch model, the authors found that the reliability of the resulting teachers’ ratings is high; there is a substantial impact of including rater differences; and it is easier to achieve a high score in some subdomains than in others. The impact of other components, such as subject area/grade level, is more complex. Further analysis was restricted to teachers in grades 3, 4 and 5, and expands beyond the traditionally-tested areas of math and reading to include assessments in other subjects. When comparing the Rasch measure with the district’s holistic written evaluations, the authors find that the correlation with value-added results is typically higher for the Rasch measure than for written evaluations, while correlation with student pretest is also higher for the Rasch measure. The authors investigate potential explanations for these differences (holistic evaluations may be subject to more error, or they may include additional information, such as an adjustment for a teacher working in a particularly difficult classroom) by using structural equation modeling to determine the relative weights of each in-person observation on the overall rating, and to determine whether any additional information seems to have been included in the holistic rating.
Stephen Ponisciak, University of Notre Dame
Nandita Gawade, Wisconsin Center for Education Research
Yang Wang, Education Analytics
Robert H. Meyer, University of Wisconsin - Madison