Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Personal Schedule
Sign In
X (Twitter)
We report on what we have learned in our efforts to build a validity argument for assessments embedded in a digital learning system (DLS) for middle grades mathematics (Author, 2015). Our validation approach borrows from Kane’s interpretive argument (2004), argument mapping (Wigmore, 1913; Toulmin, 1958), and Popper’s concepts of falsifiability and auxiliary theories (1962). We also integrated guidelines from the standards for educational and psychological testing (2014) and the CCSSO’s “Criteria for High-Quality Assessment” (2014).
Our validation work was performed on the assessments within Math-Mapper 6-8 DLS. These diagnostic assessments were built around Learning Trajectories (LT). LTs document landmarks and obstacles that students may encounter as they progress from a naïve to sophisticated understanding of a target concept (Authors, 2014).
We designed our system of assessments with four objectives in mind: Firstly, score reports will provide actionable and accurate student- and class-level feedback so that teachers can plan and inform instruction in a theory-driven manner. Secondly, if teachers interpret reports according to our guidelines, they will draw valid conclusions concerning the progress of students. Thirdly, students will know what they understand more precisely, and see a clear path to improve. Fourthly, if teachers use the conclusions to adapt instruction, learning gaps will close, misconceptions will diminish, and overall learning will increase.
Given the central role LTs play in our system and score reporting, the internal structure of the tests are a critical element of our validity argument. Therefore, we made two predictions with respect to the internal structure: LT items would be essentially unidimensional, and LT levels would positively correlate with item difficulty. To test these predictions, we conducted exploratory and confirmatory factor analyses, examined scatter plots of LT level and Rasch item difficulty (Wilson, 2005).
We have collected two years of field test data from two school districts. Sample sizes ranged from 200 to 2000 responses per test. The predicted correlation between LT level and item difficulty did surface in the majority, but not all of the LTs. We will present our interpretation of these findings in light of Popper’s (1962) concept of auxiliary theories. This presentation will have scholarly and scientific significance because it exemplifies one approach to integrating a validity argument into test development with the goal of improving the quality and coherence of validation arguments.
Jere Confrey, North Carolina State University
G. Gianopulos, North Carolina State University
Meetal Shah, North Carolina State University