Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
Tests of the replicability of original findings play a critical role in scientific progress. The purpose of this paper is to highlight ways in which replication tests, if not conducted properly, can appear to move a field toward greater clarity and truth when, in reality, they are undermining those scientific ideals. Specifically, this paper demonstrates how discretion exercised at two stages of the replication process—experimental design and data analysis—can cause replicators to arrive at incorrect conclusions about the replicability of an original finding.
This paper presents an account of one published replication debate in which replicator degrees of freedom at both the design and analysis stages played an important role: Bryan, Walton, Rogers, and Dweck (2011) published the findings of two field experiments in which participants who were randomly assigned to complete a brief, online pre-election survey on their thoughts and attitudes about “being a voter” in an upcoming election subsequently voted at a significantly higher rate (validated turnout) than participants who instead were assigned to answer questions about “voting.” An initial replication test by Gerber, Huber, Biggers and Hendry (2016), using a sample more than 14 times as big as the initial experiments combined, found no effect. The original authors countered by pointing out a number of important deviations from the original study’s context and design. The replicators then conducted a second replication test in a sample more than 10 times the size employed in the original research, correcting the most serious contextual difference from the original experiments (Gerber, Huber & Fang, 2018). The replicators reported that, even after correcting for the contextual differences, they found no evidence at all for the original effect. Here, we re-analyze the data from that second replication test and demonstrate that it, in fact, found quite strong evidence for the original effect but that evidence was obscured by the replicators’ choice of model specifications.
We take an approach to re-analyzing the data, called Specification Curve Analysis (Simonsohn, Simmons & Nelson, 2015), that is specifically designed to minimize the influence of researchers’ analytical decisions on results. This involves running all reasonable model specifications to expose the effect of various analytical choices on the results. A bootstrapping-based significance test then quantifies the strength of the evidence for the hypothesis.
A test of 540 different model specifications yields clear evidence that Gerber and colleagues (2018) did indeed replicate the original finding. Median point estimates of the size of the increase in validated voter turnout caused by this minimal intervention ranged from 3.6 to 4.2 percentage points—likely an underestimate when considering that the experiment contained a number of serious design flaws that biased against finding an effect. This result is robust in an expanded set of 2,388 different model specifications, many of which unreasonably bias the test against finding a statistically significant result. Notably, in every set of specifications we examined, the model specifications using the exact set of questionable analytical choices made by the replicating authors were quite dramatic outliers, yielding point estimates that were much lower than any other model specification we looked at.
In sum, this re-analysis makes an important contribution to the broader ongoing debate in replication science, providing direct evidence that (a) replication authors can make analysis and reporting choices that make a successful replication test appear to have failed and that are not apparent without very careful scrutiny and (b) when an apparent failure to replicate an original finding contains material deviations from the experimental design of the original study, even a minimal effort to address those deviations can result in a successful replication. In demonstrating these points, we provide perhaps the strongest evidence to date affirming an argument that has been made prominently in recent debates about replication science: sample size alone is not a useful metric for judging the quality of a replication test. Without a more thoughtful set of criteria for evaluating such tests and much more careful scrutiny of replicators’ methods, well-intentioned initiatives to improve the reproducibility of scientific findings by encouraging the publication of replication tests could have a profoundly ironic and counterproductive result: trading one sort of misleading research finding (false-positive original findings) for another (false-negative replication tests). This is a bad trade because the latter sort of misleading finding undoes the field’s hard-won progress toward improved scientific understanding.
Christopher J. Bryan, University of Chicago
David Yeager, Stanford University
Joseph O'Brien, University of Texas, Austin