Paper Summary
Share...

Direct link:

Class Enumeration in Mixture Modeling: A Summary of Recommendations and a Look Forward

Fri, April 14, 2:50 to 4:20pm CDT (2:50 to 4:20pm CDT), Radisson Blu Aqua Hotel, Chicago, Floor: 2nd Floor, Caspian

Abstract

A critical step in estimating mixture models is deciding the number of classes to retain, or the class enumeration process. The process involves fitting a number of latent classes, comparing fit information, then deciding on the best fitting latent class model based on recommended fit indices found in the mixture modeling literature, along with substantive theory to help when solutions are ambiguous. However, as with most recommendations, there is a range of different recommendations, and adopting any particular recommendation may result in selecting different numbers of latent classes. Recommendations on the performance of the fit indices are most commonly based on simulation studies, many of which use different data generation models and vary under different conditions. This paper aims to summarize the results of the vast array of simulation studies conducted and revisit the current recommendations when enumerating mixture models.
The adoption of a common set of fit indices has become widely used to determine the number of classes in latent class model: Information criteria (IC), where the lower values indicate superior fit; LMR and BLRT, likelihood-based tests that provide p-values to determine a statistically significant improvement in model fit, and Bayesian indices including the Bayes Factor and correct modeling probability, which provide the probability of each model being “true” or “correct.” While these are the more common fit indices used in enumeration, not one index is favorable over the other in every mixture modeling scenario.
Due to the variety in the nature of simulation studies, there are different recommendations about when to consider certain fit indices for specific types of mixture models, some varying by the type of mixture model. For example, an earlier study found that sample size adjusted BIC (ssBIC) was the most accurate fit statistic when detecting population differences within a structural equation modeling framework and is commonly used today (Henson et al., 2007). This study also found that the classification likelihood Bayesian information criteria (ICL-BIC) and the classification likelihood information criterion (CLC) perform consistently well across all simulation conditions when comparing more than one class, but these indices are not commonly used. However, a similar study examined the ability of multiple fit indices for latent profiles and found that the CLC and ICL-BIC correctly identified the one-class model in all conditions when variance and covariances were freely estimated (Peugh & Fan, 2013).
In addition to fit indices, other techniques that be used to help with enumeration include cross-validation of the latent classes. Using cross-validation techniques with commonly used fit indices improved class enumeration accuracy when class separation increased, sample size increased, indicators increased, and the number of classes decreased (Whittaker & Miller, 2021).
This paper explores and summarizes the current practices for class enumeration and highlights trends in the recommendations, focusing on studies that may not have received much attention. We summarize recommendations and provide insight into the promise of future simulation studies, including the need to better understand enumeration with mixture modeling and auxiliary variables, enumeration for mixture models with random effects, among others.

Authors