Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
If all interpretations and uses of test scores must be validated, then the interpretations and uses of “growth” scores must also be validated. In this paper the argument-based approach to validity (Kane 1992, 2006) is used to create three initial interpretive arguments for the interpretation and use of three aggregate growth scores, in the context of a simplified version of teacher pay for performance. The aggregate scores, which represent different conceptualizations of growth, are: median student growth percentiles (Betebenner, 2009), regression coefficients or “teacher” effects from a “value-added” covariate adjustment model (McCaffery et al, 2004) and unadjusted mean gain scores. An alternative approach rooted in the evaluation literature (e.g. Wholey, 1979) is also used to examine the same interpretation and use of the growth scores.
The argument-based approach focuses on the interpretive argument, which lays out the inferences and assumptions need for a given interpretation and use (Kane, 2006). Evidence is then gathered and used to falsify or support each assumption. Here the interpretive arguments and types of supporting evidence are created through the synthesis of the growth modeling literature and expert consultation. The value in this approach is threefold, in that it:
1. Lays out the assumptions and evidence for three growth scores in a detailed and systematic way;
2. Compares and contrasts three interpretive arguments and their sets of accompanying assumptions and evidence; and
3. Conceptually reformulates the interpretive argument for growth scores.
The findings indicate that the assumptions, as defined by Kane (1992; 2006), needed to support the interpretation and use of each growth score are almost identical; however reasonableness of the assumptions, as well as their implications and supporting evidence, varies across models. This paper also shows how some assumptions can be defined with greater specificity in order to differentiate between models.
However, the argument based approach has several limitations. Most importantly, the approach is better suited to the examination of interpretations rather than use, since the line of reasoning starts at the “observed performances [and leads] to the conclusions and decisions” (Kane, 2006, p. 23). In contrast, the alternative evaluation approach uses a theory of action to explain the mechanisms that act on the inputs to produce a set of outcomes (see Author, 2010; Wholey, 1979). This difference is borne out in the findings, in that the evaluation approach is superior in addressing the results of score use, but fails to explain how a “feasible set of performance measures” (Wholey, 1979, pg. 27) is obtained. This leads to the conclusion that if both interpretation and use are the purview of validity, then the argument-based approach should be situated within the evaluative approach, or equivalently that the interpretive argument is subsumed by the theory of action. The resulting model of validity may be helpful in conversations about the responsibilities of test developers and test users, as it provides a starting point for validating interpretations and uses of growth scores that is rooted in program evaluation and psychometric validation traditions.
Nathan Dadey, University of Colorado - Boulder
Brian Gong, National Center for the Improvement of Educational Assessment, Inc.