Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
Teacher survey self-report measures are often used to gather data about instruction because they are easy to administer and relatively low-cost. However, teachers’ survey self-reports can have low alignment with observation ratings on the quality of instruction (Spillane & Zeuli, 1999; Stecher et al. 2006). One problem with teacher survey measures is a lack of calibration; a high level of an instructional practice as perceived by one teacher may be a lower level as perceived by another. In this paper, we describe the use of “anchoring vignettes” (King et al., 2004; Kapteyn et al. 2007) within surveys to calibrate teachers’ responses so that they are more comparable, and we explore how teachers’ raw and calibrated survey responses align with a composite teacher performance measure and observation ratings of instruction.
Our data was gathered from 61 mathematics teachers in grades 4-9 participating in the Gates Foundation’s Measures of Effective Extension (METX) project. Teachers completed our online survey within one to two days after video-recording a mathematics lesson. In the survey, teachers rated themselves on six “dimensions” for their videotaped lesson: (1) mathematical vocabulary; (2) questioning; (3) student effort; (4) use of instructional time; (5) cognitively challenging tasks; and (6) remediation. For each dimension, teachers also rated and ranked four short anchoring vignettes representing hypothetical classrooms where differing levels of the dimension were present. Teachers’ vignette ratings were used to calibrate their self-ratings.
We administered the survey in two waves from January to June 2013, with questions for three randomly-chosen dimensions included in each wave. In the first wave, self-ratings came before the vignette ratings for each dimension; in the second wave, self-ratings came after vignette ratings. The videotaped lessons were scored by trained raters (two raters per video) who used a rubric that captures the same dimensions as the survey.
We also have a composite teacher performance measure for each teacher, drawn from the original Measures of Effective Teaching project. The composite measure ranked teachers based on an array of measures that included value-added, student perceptions, and classroom observations.
Preliminary findings suggest that anchoring vignettes represent a promising innovation for measuring teachers’ instruction through survey self-reports. Specifically, we found:
• Teachers’ survey responses that are calibrated through the use of anchoring vignettes have increased variation compared to teachers’ raw survey responses, particularly for the cognitive challenge dimension;
• Teachers’ calibrated survey responses regarding mathematical vocabulary and cognitively challenging tasks are more correlated with the composite measure of teacher performance compared to raw survey responses;
• If teachers gave their self-rating after rating the vignettes, rather than before, the entire collection of calibrated self-ratings are signficantly correlated with the composite performance measure (p<.05).
The paper also investigates whether teachers rank the vignettes the way we intended; which dimensions teachers and observers rate reliably; and how survey and observation ratings compare. These findings suggest that anchoring vignettes improve the accuracy of teachers’ self-reports, which has implications for how researchers and practitioners can efficiently gather and learn from instructional data.
Julia Heath Kaufman, RAND Corporation
John Engberg, RAND Corporation
Laura S. Hamilton, RAND Corporation
Heather C. Hill, Harvard Graduate School of Education
Kristin Umland, University of New Mexico
Kun Yuan, RAND Corporation
Daniel F. McCaffrey, ETS