Individual Submission Summary
Share...

Direct link:

Increasing the Generalizability of Survey Experiments

Sat, August 31, 8:00 to 9:30am, Omni, Senate Room

Abstract

Political scientists have devoted considerable attention to the external validity of experiments, primarily in terms of sample characteristics and mode effects. However, much less attention has been paid to an important aspect of external validity that is completely under the researcher’s control – the substantive topic. Upon developing a hypothesis, researchers typically select a substantive topic, such as immigration attitudes, then design stimuli for that issue. Whatever effect is found is then taken as evidence for or against the theory. But the chosen topic is just one from a larger population of possible topics, and it is unclear how well the observed effect will generalize to different substantive topics. As a result, many articles conclude with caveats about experimental results being “circumscribed by our focus on a single issue” (Chong and Druckman 2012, 14), or “limited to only two controversial issues” (Bartels and Mutz 2009, 258).

Researchers sometimes address this threat to external validity by reporting multiple experiments or multiple arms of an experiment, each on a different topic. However, this approach is costly both in time and resources, and only incrementally increases our confidence in the generalizability of the results across issues. Beyond an individual study, researchers can conduct a systematic literature review of different tests of a theory or conduct a meta-analysis. However, studies often differ in many ways, making it difficult to know why study results differ. As a recent example, Bullock (2011, 509) reviewed previous findings on the relative impact of party cues and policy information and concluded that “variation in these findings defeats most attempts to generalize.”

In this paper, we propose and implement a novel experimental design that allows researchers to overcome these problems. We first identified a population of more than 150 relevant issues from contemporary polling. We then conducted a party cue experiment that involved randomly assigning respondents to one of those issues, then randomly assigning the respondent to the treatment or control condition. This design allows the estimation of a single average treatment effect that is representative of the larger population of issues. Using multilevel modeling, we also show how to efficiently estimate issue-specific treatment effects, as well as how to directly test hypotheses about how issue characteristics (e.g., easy vs. hard) moderate treatment effects. Finally, we provide practical advice on selecting respondent and issue sample sizes.

Authors