Paper Summary

Exploring Unintended or Inappropriate Interpretations, Uses, and Consequences of a Test

Mon, April 16, 8:15 to 9:45am, Marriott Pinnacle, Floor: Fourth Level, Ambleside

Abstract

This presentation provides an alternative viewpoint regarding validation. Whereas some areas of consensus exist related to validation and appropriate roles in the validation process, we argue that gathering evidence related to consequences of test score use is not a part of validation and that conflating the necessarily distinct process of validation of score meaning and justification of score use confound both conclusions about score interpretation and the desirability of using the test.
The theoretical foundations of validity are well established in the work of Messick (1989), Kane (2006) and others. There is much common ground between those foundations and the alternative position described here in which it is argued that validation of score meaning and justification of score use must be distinguished. Strong support also exists for that position in the theoretical work of Borsboom and his colleagues (2004, 2005, 2007); Cizek and colleagues (2008, 2010, 2011), Mehrens (1997) and others.
The presentation will present the case for why the validation of a test does not involve efforts to refute rival score interpretations (outside of theoretically plausible, alternative hypotheses about score meaning), nor to justify uses of test scores (outside of those intended uses), nor to investigate and evaluate consequences of score use. To be sure, all of those efforts may be worthwhile; however, it will be demonstrated that they are squarely outside the realm of the validation task, and more appropriately accounted for via distinct methods, investigations, and evidentiary warrants.
Validity theory remains troubled regarding how to account for consequences of test use (see e.g., Brennan, 2006; Frisbie, 2005; Hood, 2009; Kane, 2009). The attempts to incorporate justification of score use into theoretical conceptions of how to ensure scientifically-supported score meaning continue to be frustrated by the logical incompatibility of the two. A new conceptualization is necessary. The framework presented in this presentation will provide an encompassing system that subsumes concerns about score meaning and test use, while differentiating these related inquiries into two parallel endeavors. The first endeavor is one that gathers and evaluates support for test score inferences; that is, validation. The second endeavor is one that gathers and evaluates support for test use; that is, justification.
There is great practical significance of this reconceptualization. In addition to providing a clearer roadmap for investigating score meaning, new research and development must address standards and procedures for justification of score use. Unlike the well-established standards and methods that exist for validity, little work has been done to articulate guidelines or procedures for justification of test use. This presentation will include beginning suggestions for such guidelines and procedures. One hoped-for consequence of differentiating between validity of score inferences and justification of test use is that rigor regarding both efforts will also advance and parallel each other. It is hoped that the reconceptualization will facilitate more complete validation practice, enhance the quality and utility of test results, and enable developers and users to improve the outcomes for all those that are the ultimate beneficiaries of high-quality test information.

Author