Search
On-Site Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Sign In
Using large-scale standardized tests to monitor student achievement overall and for student subgroups assumes tests provide valid measures of student learning that are comparable across student subgroups. Prior research has shown that score comparisons across student gender groups may be complicated by the type of test items used. Relative to female students, male students tend to earn relatively higher scores on tests using greater proportions of multiple-choice (MC) items compared to tests with greater proportions of constructed-response (CR) items (Reardon et al., 2018; Schwabe et al., 2015; Taylor & Lee, 2012), although the exact pattern and magnitude of the differences can vary across content and context (Willingham & Cole, 1997). This paper uses differential item functioning (DIF; Holland & Wainer, 1993) analyses to quantify potential item format gender bias among three recent, nationally representative samples of US 15-year-old students participating in the PISA Math and Reading tests in 2009, 2012, and 2015. The psychometric analyses investigate associations between item features and DIF to provide new empirical evidence about the magnitude of these differences among US high school students. The implications for state accountability testing, where the proportion of MC and CR items per test varies widely across states and grades, are discussed.