Paper Summary
Share...

Direct link:

Combining Multiple Measures of Teacher Practice and Performance: Technical and Conceptual Considerations

Sat, April 18, 8:15 to 9:45am, Marriott, Floor: Third Level, Cook

Abstract

There is increasing consensus that teacher evaluation systems should use multiple measures in order to adequately assess teacher performance. The debate among policymakers and the public seems to focus largely on the best approaches to estimate teacher contributions to student achievement (e.g. through Value Added Indicators). However, in practice nearly all efforts to revamp teacher evaluation systems in the U.S. base the majority of a teacher’s rating on indicators other than student achievement. These may include information from classroom observations, principal reports, parent or student surveys, and official records.
The critical question at the center of the current policy debate is thus how to combine these measures for assessing teachers (and teaching.) However, limited guidance is available for addressing the variety of conceptual, technical, and policy issues that emerge in trying to answer this question. Much systematic discussion is therefore needed around questions like what indicators will be collected or created, from what sources, and for what purposes? To what extent do these indicators provide unique or overlapping information? How can these indicators be used to improve performance? In this paper, we examine these issues in the context of performance appraisal for teacher education programs and larger scale teacher evaluation systems.
A variety of models may be used when the purpose is combining or aggregating different indicators into single composite measures. Conjunctive and disjunctive models (Mehrens, 1989) combine indicators discretely with basis on non-linear decision rules. Compensatory models on the other hand are used to create optimal linear combinations of measures; these may seek to maximize relationships a specific (and arbitrary) criterion indicator, either measured (Aamodt & Kimbrough, 1985) or unmeasured (Darlington, 1970). Finally, canonical or factor analysis models may be used to examine empirical correlations among indicators and create composites to maximize shared variance, reliability, or stability.
Importantly, however, for many other purposes (including prominently formative evaluation of teachers and teacher candidates) it is more natural to conceive of multiple measures not being combined but used in combination to provide richer and more useful information to inform efforts to improve teacher performance (Schmidt & Kaplan, 1971.) Thus, how to combine multiple measures of teaching is only partly a technical issue. Teacher evaluation systems must make decisions about the value of each measure and consider conceptual and practical issues related to definitions of effectiveness, local context, program goals and priorities.
In considering these decisions it is useful to conceive of multiple measures as components of a validity argument for specific inferences to be derived from the various indicators for specific purposes (Kane, 2006). The validity framework highlights the need to rely on assumptions about the nature of the theoretical construct being measured (i.e. teacher quality or effectiveness) and the ways in which it may be operationalized in practice. It also makes clear that different uses will lead to different validity arguments, and require different sources of evidence and support.

Authors