Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Visiting Washington, D.C.
Personal Schedule
Sign In
X (Twitter)
Perspective and Purpose
There are many statistical approaches for testing measurement invariance. Common to all these methods is the need for researchers to conduct tests on each item, one by one, to discover non-invariance. As such, testing becomes more onerous as the number of scale items and response options increase. Further, if more than two or three groups are involved, this process needs to be repeated for each pair of groups individually. This results in a myriad of complicated, statistical tests where errors of inference are likely.
The alignment method (Asparouhov & Muthén, 2014) was recently proposed as an alternative means of estimating confirmatory factor models across many groups. The method automates the estimation of model parameters across many groups and testing for noninvariance. Simulation studies show that the method works well with items are continuous (Asparouhov & Muthén, 2014), but no simulation research has yet been conducted with polytomous items.
The primary purpose of the current study was to understand under what circumstances the alignment method should be used when scale items are polytomous. As with any new method, researchers need procedures to detect violations of the assumptions and guidance about how those violations will impact their results.
Data Source and Methods
Item response data from a large scale assessment were analyzed to obtain realistic specifications for data generation; we then manipulated those parameters to meet various conditions (Table 1). Five-hundred replications of each condition were run with the alignment method. Then we calculated the bias, MSE, and coverage of measurement parameters and the hit rate of the testing procedure.
Results
Overall the measurement model part of the method was excellent at recovering the true parameters and produced estimates with little bias. The method uniformly performed well when the magnitude of noninvariance was small (loadings = -.10, thresholds = -.20), even if 43% of the items were noninvariant. Across all of the different estimates there were generally only substantial issues in conditions with 29-43% noninvariance of a medium or large magnitude, regardless the location of the noninvariance.
However, the noninvariance testing procedure was too conservative, flagging noninvariant items as invariant a majority of the time, across all conditions. The procedure worked better for the thresholds than for the loadings and when the magnitude was large (see Figure 1). It appears that there is an interaction between the ability of the procedure to estimate the parameter when noninvariance is extreme and the distribution of item responses across the categories, such that if a parameter is poorly estimated due to noninvariance and has skewed thresholds, the procedure cannot detect noninvariance.
Significance
For researchers who want to use the alignment method a simple set of recommendations and criteria needs to be established. This study is a first step toward making those recommendations. Because of space limitations not all recommendations are included here, but they include investigating the mean structure of the items, rules of thumb for mean differences estimated between items, and use of an R2 statistic provided in the alignment output.