Paper Summary
Share...

Direct link:

Beyond High-Stakes Tests: Teacher Effects on Other Educational Outcomes

Sat, April 14, 2:15 to 3:45pm, Vancouver Convention Centre, Floor: Second Level, East Room 11

Abstract

In this paper, we explore two largely untested assumptions on proposals to evaluate teachers based on value-added measures of teacher effectiveness. First, policy-makers supporting these systems believe that the high-stakes tests from which value-added measures are derived provide sufficient and accurate information—information that is not contaminated by well-documented behavioral responses to accountability systems—to determine which teachers are “good.” Second, these policy-makers assume that value-added measures are predictive of other outcomes, such as long-term educational attainment, earnings, and participation in civic life, that we care about, and thus the use of test scores to make evaluation and compensation decisions is not problematic. In other words, the belief is that a “good teacher”—as measured by value-added to test scores—is good at all components of the job.
We present evidence that the choice of test can make a difference to value-added comes from recent research comparing value-added measures on multiple tests of the same content area. Since 1998, the Houston Independent School District (HISD) has administered two standardized tests every year—the state TAKS and the nationally-normed Stanford Achievement Test. Using data from HISD, we calculated separate value-added measures for 4th and 5th grade teachers for the two tests. These measures were based on the same students, tested in the same subject, at approximately the same time of year. The key difference between the two tests is their stakes: one is used for accountability purposes, while the other is not. While there are other reasons why value-added measures may vary across tests, our interest is primarily in the role high stakes play in their divergence. We found that the correlation between teachers' value-added on the two tests is modest at best: 0.52 in reading and 0.60 in math. This modest correlation yields many inconsistent rankings of effective teachers.
A second set of issues with measuring teacher effectiveness stems from the fact that public education has multiple goals, including promotion of social and behavioral skills. Analysts and policy-makers measure test scores because they assume they are correlated with these outcomes and current policy assumes that a “good” teacher is effective across domains, but in fact we know little about whether teachers good at supporting one set of competencies are similarly capable at supporting others. We build on this literature by reporting the results of a recent study by Jennings and DiPrete (2010), which analyzed data from the Early Childhood Longitudinal Study – Kindergarten Cohort and asked whether kindergarten teachers who are good at promoting academic growth are also good at promoting social growth. Jennings and DiPrete found that the social-academic correlation ranged from .13-.17. These results suggest that social and behavioral teaching is a distinct competency from the teaching of math or reading.

Authors