Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Personal Schedule
Sign In
X (Twitter)
Diagnostic classroom assessments that utilize learning trajectories (LTs) offer opportunities to improve mathematics learning (Author_C, 2019a). Validation studies of these LT-based assessments are critical to ensure that claims about learning align with the evidence collected (Graf & van Rijn, 2016). Most validation studies of LTs have centered around single LTs (e.g., Author_D, 2011; Steedle & Shavelson, 2009), however, there has been limited research examining validity evidence across multiple content areas. Math Mapper 6-8 (MM) is an LT-aligned diagnostic assessment system comprised of a learning map with nine big ideas, 23 relational learning clusters (RLCs), and 62 LTs encompassing the middle school curriculum. Our presentation reports on results from validation work conducted in all 62 LTs to gain a better understanding of both the LTs themselves and the validation process for LT-aligned classroom assessments.
The validation approach used in this study builds on a validation framework derived from one by Pellegrino, Dibello, and Goldman (2016) that focuses on cognition, instruction and implementation. Our version uses those components to interrogate six purposes for using an LT-based assessment system. We created a validation argument (Table 2) that builds claims to the use of data in the classroom (Author_C, 2019b). A critical aspect of our validation approach involves a deep collaboration between experts in the learning sciences and psychometrics (Author_C, 2019c).
Data Sources and Methods
This presentation focuses on first round validation studies that covered the first three claims of the validation argument and utilized data from approximately 37,000, 8-10 item MM assessments collected over three years from students in grades 6-8 at three middle schools. The validation studies for each RLC consisted of item response theory (IRT) modeling of the data, sequential regressions to identify potentially non-conforming items, and learning science reviews of items, LTs, and RLCs based on the results. The analysis was aimed at identifying and addressing variations of item difficulty and refining the expectations of how LTs are structured in terms of ordered levels with acceptable within-level variations.
Results and Implications
As a result of this validation work, a wide range of changes were made at the item, LT, and RLC level across the map (see Figure 2). This presentation investigates these variations by examining the differences among the LTs from learning science and psychometric perspectives. For instance, preliminary comparisons among the LTs with low correlations between item difficulty and LT level suggests that these LTs correspond with content that either has a limited research base or does not receive curricular emphasis. Other comparisons among LTs highlight differences in linearity, steepness of slope, and spread of difficulties within a level which provide insights into how fine-grained LTs, designed to impact instruction, should be considered. Data from the following school year will also be used to evaluate how item adjustments from the first round performed. The results of this study have significant implications for making generalizations about the collection and analysis of validity evidence from classroom assessments at scale.
Jere Confrey, North Carolina State University
Emily Toutkoushian, The American Board of Anesthesiology
Meetal Shah, LearnPlatform