Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
Objectives or purposes: The internal consistency and validity of academic achievement tests are significantly different for EO and ELL students (Abedi, 2002, 2006). The objectives of this paper are to 1) determine if we find DIF in the scores of diverse middle school students in their performance on vocabulary assessment, and (2) test competing models of ability in calculating DIF. These results will inform both student learning assessment and choice of ability estimates in DIF analyses.
Theoretical framework: We assume that assessments should be similarly reliable and accurate for all students, although there is strong evidence that is not the case. We are informed by theories that emphasize the importance of vocabulary knowledge for skilled reading (Perfetti & Stafura, 2014) and see vocabulary learning as tightly connected to other language skills in L1 and L2.
Methods: Differential item functioning (DIF) were addressed through logistic regression and item response models.
Logistic Regression Approach. To detect whether vocabulary items function differently for EOs and ELLs, we conducted logistic regression analyses for each vocabulary item (Swaminathan and Rogers, 1990) using SAS.
IRT Approach. DIF was also investigated with the 1PL descriptive item response model (dIRT) using IRTPRO 2.1 software (Cai, Thissen, & du Toit, 2011). DIF testing consisted of two steps, and did not require designated anchors. In the first step, we will fit a model wherein mean = 0 and variance = 1 for EO students, the mean and variance for ELL students was estimated, and item difficulty was constrained to be equal for EO and ELL students. In the second step, a model was fitted with the ELL mean and variance fixed to the values obtained in the first step.
Data sources: The primary measures are the Word Generation Assessment of Academic Vocabulary and the reading and vocabulary tests of the Gates. We are conducting analysis of the pre-test assessments from the Word Generation trial. The sample is 13,780 students who attended thirteen middle schools in a large diverse urban district in California.
Results: Figure 1 presents relations between different estimates of uniform DIF for each item on the 1011 test form. The difference between the two estimates stems from the use of either the Gates (left panel) or a latent score (right panel) as the estimate of Vocabulary Ability. Although the estimates for uniform DIF for a given item were highly correlated across approaches, DIF classification varies across models. The DIF classification based on the model with the latent ability score was in line with research suggesting that most items showing DIF favor EOs since the tests were normed on or developed for EO students.
Scientific or scholarly significance: DIF analysis often uses total score as a default for ability. Our results suggest that a richer conceptualization of student verbal knowledge is needed to understand the performance of students on vocabulary items across language proficiency levels. In this paper we draw practical implication from these findings, and explore the DIF results at the item level across word characteristic such as frequency, dispersion, semantic precision and polysemy.
Paulina Kulesz, University of Houston Times
Autumn McIlraith, University of Houston
Joshua Fahey Lawrence, University of Oslo
David J. Francis, University of Houston