Individual Submission Summary
Share...

Direct link:

Towards Best Practices for Treating Missing Data in Mixture Models: Results from Latent Profile Analysis

Thu, March 21, 2:15 to 3:45pm, Hilton Baltimore, Floor: Level 2, Key 4

Integrative Statement

It is well established that missing data rates of as little as 5-10% can bias estimates and threaten statistical inference in applied developmental research (Little et al., 2013). Consequently, methodologists recommend that researchers adopt “state of the art” (Schafer & Graham, 2002) missing data procedures, such as full information maximum likelihood (FIML) or multiple imputation. Despite the widespread adoption of these missing data strategies, the treatment of missing data remains problematic when conducting person-centered analyses. The goal of this presentation is to (a) highlight the real-world complications that arise when conducting latent profile analysis in the presence of missing data, and (b) to provide practical suggestions for applied researchers for best practices.

To clarify concepts, I will rely on a pedagogical example of a simple latent profile analysis performed on data gathered from N = 2,066 Tanzanian children aged 18-37 months. Specifically, profiles were formed using age-adjusted motor, cognition, language, and socioemotional scores from the BSID-III. The data also contains information on stimulation (using the MICS) and the height-for-age anthropometrics (HAZ). Nearly 10% of observations were missing at least one BSID subscale score. Further, HAZ predicts missingness, suggesting that the missingness is not completely at random and may depend on variables external to the profile indicators. A two-class model was selected, and the profiles were named (1) socioemotional development as-predicted and (2) socioemotional development under-predicted. The socioemotional under-predicted profile tended to have lower socioemotional scores relative to the subscores on the other domain, whereas the as-predicted group tended to have scores that were reflective of their adjusted scores on the other domains.

Using the real-world example, I will highlight the limitations of FIML and multiple imputation for handling missing data. Although FIML may be an appropriate strategy in some situations, most often it is inadequate in the real world since missingness often depends on variables other than the profile indicators (such as HAZ or MICS). On the other hand, multiple imputation is far more flexible in handling real-world data because it can easily incorporate this auxiliary information (Enders, 2010). Even so, multiple imputation is known to artificially decrease class separation (Enders & Gottschall, 2011; Sterba, 2016).

Importantly, I will show that the documented poor performance of multiple imputation in LPA is not a limitation of multiple imputation as a general strategy for dealing with missing indicator data. Instead, the poor performance reflects the shortcoming of the choices for the models used to generate imputations. In fact, I will demonstrate that tree-imputation models are a robust alternative to the FIML procedures and the generic imputation practices that currently dominate research. I will discuss the conditions in which tree-imputation outperforms FIML and other multiple imputation strategies. Finally, I will illustrate how researchers can implement tree imputation real-world BSID-III data and detail best practices for pooling model fit statistics across multiply imputed data sets to conduct enumeration.

Authors