Paper Summary
Share...

Direct link:

Using Mixture Modeling to Make Meaning of Differential Item Functioning

Sun, April 19, 8:15 to 10:15am, Virtual Room

Abstract

This paper discusses theoretical frameworks for using the combination of differential item functioning (DIF) and mixture model findings for validity arguments. Additionally, it discusses the philosophical consequences of invoking the notion of measurement invariance when using mixture models, namely causal theories of measurement (Hood, 2009). While some researchers are explicitly working on the statistical aspects of measurement invariance in mixture modelling including covariate inclusion and DIF within mixture models (c.f. Cohen & Bolt, 2005; Cole, Bauer, & Hussong, 2019; Masyn, 2017), certain theoretical and philosophical problems remain. For use-cases, identifying reasons for DIF is important. Mixture models may help with understanding aspects of measure and test fairness (and describing possible causes of DIF in non-mixture settings), especially as DIF researchers increasingly acknowledge that person groupings are not as simple as the categories often assigned (Mislevy, 2018). Emergence of latent classes indicate that the data generation process exists differently for some respondents to a test or survey. A more complicated vision exists when measurement non-invariance is invoked within a finite mixture model where model parameters differ across classes for manifest person groupings (Lubke & Luningham, 2017).
This paper discusses the theoretical challenges of interpreting DIF and extends this thinking to mixture models. An idealized example is provided using a reading measure that has some items requiring Spanish language knowledge. Here, measurement non-invariance within the mixture model may have implications for item selection and may provide important insights about the determinants of item response probabilities.
Currently, finding unobserved heterogeneity in a sample-space is an indication of potential measurement non-invariance and a threat to test validity. In the motivating example from the reading test, certain items have English-Spanish cognates. Thus, it would be unfair to compare students who speak English with those who speak both English and Spanish because of the presence of construct irrelevant variance (Messick, 1995). However, Borsboom, Mellenbergh, & Van Heerden (2002), note that there may be two kinds of possible measurement: absolute measurement and relative measurement. Therefore, the items may still be of interest for assessing both groups of students. Under relative measurement, students would be compared to fellow language-speakers. The extent to which this measure is valid is influenced by the interpretation of the mixture model and the ability to classify readers assuming that the model identifies a real categorical latent variable. Finding measurement non-invariance among these groups could indicate other typologies and violation of possible relative measurement.
As noted by Masyn (2017), given the framework that mixture models are one form of general latent variable model, the logic of measure validation applies as it does to continuous latent variables. This fact perhaps bolsters the claims of Borsboom et al. (2002) that DIF is not an immediate reason to discard items, when, in fact, these items may be informative for relative measurement. Mixture modelling may provide a path forward for Borsboom et al.’s (2002) relative measurement. What this entails for the ontology of the latent variable remains a question to ponder (Slaney & Racine, 2017).

Author