Paper Summary

What Counts as Evidence? An Empirical Review of Validity Studies in Educational and Psychological Measurement

Sat, April 14, 10:35am to 12:05pm, Pan Pacific, Floor: Lobby Level, Oceanview 1&2

Abstract

Objective: This study reviews past and present validation practices in the journal Educational and Psychological Measurement (EPM) to address three primary research questions:

1. What sorts of validity evidence are presented to support test and measurement interpretations in educational and psychological measurement?
2. How is the evidence integrated and presented to support these interpretations?
3. How has this changed over time?

Theoretical Framework: This research is situated within the framework of modern validity theory (e.g. AERA, APA, & NCME, 1999; Kane, 2006; Messick, 1989; Sireci, 2009; Zumbo, 2007, 2009). Validity is no longer seen as a static property of tests, but rather as an integrated judgment about the degree of justifiability of inferences and decisions we make based upon test scores (Kane, 2006; Messick, 1989). As validity has become more expansive it has also become more complex, giving rise to debates in the field about what sorts of evidence are needed in different contexts and how best to synthesize that evidence. An analysis of validation practices in educational and psychological measurement extending previous work (e.g. Cizek, Rosenberg, & Koons, 2008; Hogan & Agnello, 2004; Jonson & Plake, 1998) would be helpful to inform such discussions.

Methods and Data Source: We used the Validity Studies section of the journal EPM to operationalize validation research practices. All studies published from 1960-1969 (Volumes 20-29; n = 265) and from 2000-2009 (Volumes 60-69; n = 293) were compiled. We randomly selected 20 studies from each decade to serve as the sample for analysis. Using the five categories of validity evidence described in the Standards, we coded whether the evidence presented in each study was based on a) content, b) response processes, c) internal structure analyses, d) relations to other variables, or e) consequences of testing. We also noted whether the authors cited recent work in validity theory, whether they used an explicit theoretical framework to synthesize evidence, and whether reliability estimates were presented.

Results: Two sources of evidence (based on response processes and consequences of testing) are rarely presented in past or current studies, a finding reported elsewhere (Cizek et al., 2008). Over time, the largest change in practice was an increase in studies presenting internal structure analyses and reliability estimates. Although evidence presented in validation studies has become more complex over time, generally including a wider range of evidence, researchers do not often present their work within the various frameworks of modern validity theory, which includes the Standards.

Significance: This research can be used to a) inform validity theory by describing the context in which these theories might be applied, and b) inform our measurement practice by highlighting under-analyzed aspects of test use and interpretation, and hence potentially unconsidered threats to our test-based inferences. Some aspects of current validity theory, such as the role of consequences, are particularly controversial (Cizek, Bowen, & Church, 2010; Hubley & Zumbo, in press); we hope this analysis of validation practice will help to inform these debates and make them as relevant as possible to researchers.

Authors