Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Personal Schedule
Sign In
X (Twitter)
Planned missing data designs in large surveys can efficiently reduce respondents’ burden and lower the cost associated with data collection, without cutting down on the questionnaire items. If the missing data are not appropriately planned, descriptive and potential causal parameter estimates will be biased. For a fixed sample size, the extent of bias depends on three major characteristics of design and data: the missing percentage, the overlap percentage (i.e., the portion of cases where two items are observed jointly), and the distribution of variables. This paper implemented two simulation studies to investigate how the bias in marginal means, correlations and regression coefficients depends on the chosen planned missing data designs and the related characteristics.
The first simulation study investigates the bias in means, correlations and regression coefficients by systematically varying the overlap percentage, missing percentage, joint distribution of the data and sample size. The results show that the estimates of means are unbiased for large survey data (i.e., sample sizes exceeding 1000 cases) even when overlap is zero and the missing percentage is high. However, the recovery of correlations and regression coefficients requires positive overlap. The bias in correlations is negligibly small when there is 20% or more overlap for continuous data. A low missing percentage is of minor importance for bias reduction as long as the sample size is large (at least 1000). Regarding the regression coefficients, the bias is negligibly small when overlap exceeds 20% for multivariate normal data. For skewed and categorical data, the coefficients are estimated with larger bias and less reliability than for multivariate normal data, though with 33% overlap or higher all biases are still within 0.05 standard deviations of the outcome variable.
The second simulation study compares a two-form design, a three-form design, and two optimal block designs with 50% and 33% missingness. The results show that all designs recover the means of variables without bias. The biases in correlations are negligibly small for all designs except for the two-form design which has no overlap across forms. For the regression coefficients, the two-form design again performs the worst due to no overlap. With the same amount of missingness, the optimal block design largely reduces the bias due to its 20% of overlap. Furthermore, the three-form design and the optimal block design produce negligibly small bias and more reliable estimates due to more overlap and less missingness. The results show that given the same amount of missing percentage, an optimal block design that ensures sufficient overlap and maximizes the efficiency at the same time can be a better choice than a multiple-form design.