Paper Summary
Share...

Direct link:

Developing Multiple Measures of Teacher Effectiveness in the Pittsburgh Public Schools

Fri, April 4, 12:25 to 1:55pm, Convention Center, Floor: Terrace Level, Terrace IV

Abstract

Background and Purpose
Pittsburgh is an innovator in the nationwide movement to evaluate, enhance, and reward effective teaching. Working collaboratively with the Pittsburgh Federation of Teachers, the Pittsburgh Public Schools (PPS) have been developing more intensive professional development opportunities, a modified career ladder, financial awards for highly effective teaching, and several new measures of teacher effectiveness. This study examines multiple measures of teaching effectiveness in Pittsburgh, with the aim of assisting PPS’ efforts to refine those measures to create a picture of teacher performance that is richer, more valid, and more comprehensive than any single measure could produce on its own.

The PPS teacher evaluation system includes three types of measures. Value-added measures (VAMs) use student test score gains to identify each teacher’s contributions to student achievement, bringing in data on up to three years of teaching. Professional practice measures rely on principal assessments using Pittsburgh’s Research-based Inclusive System of Evaluation (RISE), which is based on Charlotte Danielson’s widely used Framework for Teaching (Danielson, 2011). The student survey measures, called Tripod (developed by Ronald Ferguson of Harvard University and administered by Cambridge Education), incorporate students’ perceptions of teachers and the classroom environment.

Framework
The study has two major aims: (1) to understand the extent to which RISE and Tripod produce differentiation in teacher scores; and (2) to examine the extent to which RISE and Tripod ratings are correlated with each other and with teachers’ estimated value added.

Methods and Data
The analyses in this study use unique teacher IDs to link PPS’s three sources of teacher effectiveness data (value-added, RISE, and Tripod), focusing on the 2011-12 school year. The analyses also incorporate value-added data covering up to two previous years of performance. RISE data include teachers’ scores (as rated by their principals) on 12 to 24 components that are divided into the following four domains. The questions on the Tripod student survey fall into one of the following seven constructs (see Kane and Staiger 2012 for more information).

Descriptive statistics will be calculated to determine the amount of variation in the RISE and Tripod data (similar analyses have already been conducted for Pittsburgh’s VAM estimates; Johnson et al., 2012). We will begin by examining the distribution of RISE and Tripod ratings measures to see how much differentiation among teachers is present; and we will use the component RISE and Tripod measures to assess their internal consistency. We will also examine the extent to which RISE and Tripod scores are related to student and teacher characteristics.

Contributions/Significance
The analysis will also help identify ways to simplify and improve the utility of the observational metric. It can also inform how teachers lacking value-added scores might be rated with only observational and student survey measures. Finally, the analysis may guide the district’s efforts to create a summary measure of teacher effectiveness that includes several component measures. The findings should be useful not only to educators in Pittsburgh but to others who are working on the development of new systems of teacher evaluation.

Author