Search
On-Site Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Sign In
X (Twitter)
High stakes admissions testing is typically carried out using items with pre-specified difficulties, with these difficulties calibrated using pilot testing. Though this approach often works well in practice, factors such as item order and time pressure can modify the testing context enough that pre-calibrated parameters are no longer valid, threatening score validity. As an example, look to the 2014 administration of ENEM, the national Brazilian college entrance exam. ENEM has a single yearly administration where all students take the same items at the same time under time pressure. To discourage cheating, ENEM randomly assigns students one of four colored booklets. Though each booklet contains the same items, the pages are presented in a different order. As these booklets are assigned at random, the ability distributions for students for each color booklet are identical in expectation. Despite this, in 2014 we observe that students taking the blue book perform the worst in math, implying that item order has an effect on student performance and the color of your booklet could partially determine whether or not you attend college.
Previous approaches that model position effects as variation in item parameters imply that these effects are homogenous across respondents. Similarly, assuming that individual ability decays over the course of a testing session implies that position effects are independent of item content and stem from individual test endurance, contrary to observations by Kingston & Durans (1984). To address this gap, we propose an item response model that treats position effects as both person-side and item-side by modeling heterogeneity in individual response processes over the course of the test. Here an individual’s encounter with an item is treated as a smoothly varying mixture of how the student would interact with the item if encountered early in the test and how the student would interact with the item were it encountered late in the test, weighted by actual item position. This directly models a difference in response processes for students who are “fresh” and “fatigued,” estimating ability net of individual endurance. We apply this model to ENEM 2014, showing that the score gap between booklets disappears.