Paper Summary
Share...

Direct link:

Approaches to Incorporating Late Pretests in Experiments

Sun, April 6, 8:15 to 9:45am, Convention Center, Floor: 100 Level, 113B

Abstract

Ideally, baseline data collection should occur before the start of program/intervention implementation, but unforeseen factors may delay it, leading to “contaminated” baseline measures (henceforth: “late” pretests). Incorporating late pretests in impact regressions has two consequences (Schochet, 2008). First, resulting impact estimates may be biased, reflecting early intervention effects. Magnitude and direction of this bias depends on the size and direction of the early effects (e.g., large and positive early effects cause impact estimates with large downward bias). Second, late pretests tend to be correlated with the treatment indicator(s) included in impact regressions due to the early treatment effect, yielding larger standard errors than uncontaminated pretests. Despite these adverse effects, using late pretests in impact analyses may still be preferable because they may explain a significant portion of the outcome variance and help with precision, offsetting the bias they introduce. This bias-precision tradeoff depends on the size of early treatment effects, growth trajectory of treatment effects, and how well late pretests explain posttest measures.

Following Schochet, we assess the extent of early treatment effects and compare estimated impacts for the BB and BB+SSR conditions using three estimators. The first is the posttest-only estimator that does not use any pretest measures. The second is the ANCOVA estimator that controls for the pretest measures as they are. The third is an alternative ANCOVA estimator that removes the contaminated portion of pretests, yielding unbiased impact estimates while still explaining some portion of the outcome. This approach entails: (i) building a model that uses the contaminated pretest as the outcome and all available exogenous covariates and the time between school start and baseline testing as covariates, (ii) estimating this model using only the control students’ pretest measures, and (iii) creating predicted pretests for all students using the estimated model but setting the time between school start and pretest date to zero. Note that this process yields predicted pretest scores that are free of early treatment effects because the prediction model is estimated using covariates that were collected before the treatment and only control students’ scores.

Analyses conducted thus far show that controlling for contaminated pretest measures in impact regressions can lead to substantially biased impact estimates, although their effect on the precision of impact estimates is less profound, causing one to question the overall merit of using them. The preliminary application of our decontamination approach with a limited set of predictors yields unbiased estimates, but it does not seem to help much with precision. These analyses will be extended and improved in time for the conference with the use of additional administrative data — scores from an assessment that teachers complete twice a year to measure children’s progress, which are presumably better presumably better predictors of study-administered pretests— that are currently being obtained.

Authors