Paper Summary

Examining the Viability of Recent Models for Forced Choice Data

Sun, April 15, 8:15 to 10:15am, Marriott Pinnacle, Floor: Third Level, Pinnacle I

Abstract

Forced-choice tests (a.k.a., ipsative tests) can be used to provide information about an individual’s relative standings across multiple dimensions (i.e., they can be used for intra-individual comparison). Moreover, ipsative tests can also be potentially useful in addressing testing issues associated with social desirability and faking. To extend the usefulness of forced-choice tests, and to extract scores from ipsative instruments that can be used for inter-individual comparisons, Stark, Chernyshenko and Drasgow (2005) proposed an item response theory (IRT) model called the multi-unidimensional pairwise-preference (MUPP) model.

Recently, several researchers have extended the MUPP model in various directions to handle more diverse situations that can be encountered in practice. In the first extension, Leenen, Ponsoda, de la Torre, and Romero (2010) reformulated the MUPP model within a Bayesian framework, and proposed a Markov chain Monte Carlo (MCMC) procedure for estimating the model parameters. Their findings indicate that the proposed procedure can be used with the MUPP model even with unknown component parameters and latent correlation structures.
In the second extension, de la Torre, Ponsoda, Leenen, and Hontangas (2011) proposed the PICK (pick the “best” component), MOLE (pick the “best” and “worst” components) and LUCE (rank order the components) models as extensions of the MUPP model to accommodate more than two components, and different forced-choice formats. Their results reveal that estimates based on these models perform appreciably better than the traditional scores (e.g., the number of times components from a dimension is selected). The correlations between the true and estimated latent traits for the proposed models ranged from 0.61 to 0.88, whereas the range was only from 0.36 to 0.52 for the traditional scores.

The current paper is a two-part study that further extends these current works. The goal of the first part of the study is to investigate in-depth how the design of the test based on the MUPP model relates to the quality of parameter recovery obtained using MCMC. In particular, we designed a simulation study, where the following factors are varied: test length, the number of latent traits measured, the number of different components per dimension, the difference between (the location parameters of) the components in a pair, and the proportion of unidimensional pairs (i.e., with both components from the same dimension) in the test. The goal of the second part of the study is to systematically examine and compare the behavior of traditional and IRT latent trait estimates obtained from various types of forced-choice instruments. We specifically manipulated the following factors: item format (PICK, MOLE or LUCE), test length, item discrimination, range and variability of the item location parameters, and size of discrepancy between true and assumed correlation structures. For both parts of the study, the correspondence between the true and estimated parameter values is evaluated using several goodness-of-recovery statistics and compared across the conditions of the simulation design. Taken as a whole, the current work provides a solid basis for evaluating the viability and practical utility of these recently proposed models and procedures for forced choice data.

Authors