Paper Summary
Share...

Direct link:

Observing Special Education Teachers in High-Stakes Teacher Evaluation Systems

Fri, April 5, 2:25 to 3:55pm, Metro Toronto Convention Centre, Floor: 200 Level, Room 203A

Abstract

Objective
To date, no research studies have examined the validity of general observation systems for use with special educators, a group that represents approximately 14.5% of the teaching workforce. Further, researchers have raised concerns about whether such systems appropriately account for the kinds of instructional practices most commonly used by special educators (Author, 2015b). To help identify effective special education teachers and to help those who are struggling to improve, it will be critical to districts that they have accurate, reliable observation data.

Theoretical Framework
In the present study, we investigate the extent to which the Framework for Teaching (FFT), a widely-adopted observation scheme, reliably and validly captures special educators' practice. To do so, we adopt Kane's validity argument approach, in which we appraise the plausibility of the argument that judgments of special educator teaching quality can be made on the basis of FFT scores. This argument is to be validated through an empirical evaluation of four related sets of inferences: 1) scoring, 2) generalization, 3) extrapolation, and 4) interpretation.

Methods and Data Sources
In conducting our validation, we collected classroom observation data from 80 special education teachers in Rhode Island and Idaho. We included elementary and middle school teachers in grades 3-8. Each classroom teacher was videotaped four times in 2016-2017; teachers also completed surveys and were interviewed about their backgrounds, teaching context, and evaluation experiences. All 320 teacher lessons were scored on both FFT and on a set of observation instruments more closely aligned with special education teaching practices, the Classroom Observation Student-Teacher Interactions (COSTI) and the Quality of Classroom Instruction (QCI) (Doabler et al., 2014; 2015). Scores on each protocol are also compared to teachers' value-added scores.

Results
Overall, special educators' lessons scored as expected in the classroom environment domain (Domain 2). Across the 321 lessons, the mean scores ranged from 2.23 on 2b: Culture to 2.94 on 2e: Space, on FFT's 4-point scale. In contrast, scores in the Instructional domain (Domain 3) were consistently low across all components, with average scores ranging from 1.78 to 2.30, far lower than previously-published studies using FFT. When comparing FFT scores to scores drawn from the special education instrument, the two measures rank ordered teachers in similar ways. However, the results indicate a ceiling on FFT scores between basic and proficient; regardless of how strong teachers' lessons looked on the QCI, they were unlikely to score highly on the FFT

Scholarly Significance
There are complex conceptual and logistic challenges in adopting a common evaluation tool across all teachers in a state. While this question can be viewed in narrow, psychometric terms, we argue that the results of this may have profound implications for how we think about the twin goals for using observation scores - teacher improvement and teacher sorting. For the purposes of teacher sorting, it is helpful to see that FFT and QCI rank order teachers' similarly. However, for improvement purposes, decreased scores could lead to wasted time and resources as well as teacher discouragement and frustration.

Authors