Paper Summary
Share...

Direct link:

Challenges, Best Practices, and Innovations in Complex Sample Data Analysis

Mon, April 11, 10:00 to 11:30am, Marriott Marquis, Floor: Level Two, Marquis Salon 4

Abstract

Purpose
The purpose of this study is to: a) examine current published literature that has analyzed complex survey data and determine if/how the sampling design was accommodated (with a special focus on longitudinal studies and matched comparison studies using propensity score analysis); and b) to discuss challenges and best practices in the application and reporting of complex samples analyses.
Theoretical Framework
Complex sample data are data collected by non-simple random sample designs, and large-scale data collection efforts are generally always complex, employing sampling designs such as multi-stage, stratified, and cluster features (Kalton, 1983). Complex sample designs are analogous to a “symphonic arrangement… harmonically [combining] numerous elements in order to produce the desired outcome” (Turner, 2003, pp. 2-3), and two analytic issues must therefore be addressed when analyzing complex sample data: 1) homogenous clusters created by multi-stage sampling; and 2) unequal selection probabilities (e.g., created by oversampling, adjustment for nonresponse). Longitudinal designs carry the extra consideration of non-response bias; an issue that may be alleviated via the use of FIML with auxiliary variables. Similarly, analysis of complex samples with propensity score analysis (PSA) has added complexity given the multiple process of PSA, and the point at which survey weights should be incorporated in PSA is still up for debate (Hahs-Vaughn, 2015).
Methods
A content analysis was conducted of journal articles using ECLS-K data published 2010-2015 and identified in EBSCOhost. The ECLS-K is a complex sample, is used for substantive research in various disciplines (e.g., economics, school psychology, music education, speech language) (e.g., Cho, 2012; Mashburn, 2010; Miksza & Gault, 2014; Titley, D'Amato, & Koehler-Hak, 2014), and is often used as demonstration data. Search terms included “ECLS” and “early childhood longitudinal study.”
Data Source
There were 172 articles identified. Items developed by Author (year) were coded from articles subject to full review: 1) accommodation of complex sampling design; and 2) description of complex survey adjustment. Level of accommodation was also coded [e.g., fully addressed; partially addressed with evidence; mentioned but no evidence]. Results will be interpreted in aggregate and disaggregated for PSA studies. For longitudinal studies, studies were coded based on the use of non-response adjusted panel weights versus FIML estimation using only base year sampling weights with auxiliary information.
Results
Preliminary results suggest that less than 20% of published studies appropriately made some type of adjustment for the sampling design. However, these studies generally only mentioned the type of accommodation without providing supporting evidence (e.g., survey weight name).
Significance
Results illustrate the widespread use of complex samples but the limited understanding of how to appropriately adjust for the design. Results are similar to previous research (Author, year) despite the plethora of research that illustrates problems that exist when complex survey design is not addressed (e.g., Kaplan & Ferguson, 1999; Lee, Forthofer, & Lorimor, 1989; Stapleton, 2002). Research consistently shows that failing to address the complex sampling issues underestimates standard errors and thus results in inflated test statistics.

Author