Individual Submission Summary
Share...

Direct link:

Routine Researcher Variability and the Reliability of Replications

Thu, August 29, 8:00 to 9:30am, Marriott, Washington 2

Abstract

Scientific progress depends on the accuracy of cumulative insights researchers gain through empirical observation. However, these insights depend not only on the quality of our data but also on the reliability with which these same researchers transform available data into conclusions about the political phenomena they intend to study. If researchers differ in the conclusions they draw from identical data, this constitutes a distinct source of error in our empirical knowledge base left unexplored in the political methodology literature. In this paper, we analyze the variability in results among a population of researchers conducting the same research using identical data on the same population. We call this researcher variability and identify it as a potential threat to the reliability and, by extension, credibility of research. We identify two types: non-routine researcher variability that derives from deliberate choices researchers make and routine researcher variability that derives from undeliberate actions researchers take under constraints. While non-routine variability can be controlled through observation and curation of researchers’ choices routine researcher variability cannot, at least not entirely, making it a worrisome and underappreciated source of error. In this paper, we present the first-ever estimates of this phenomenon, in the specific context of replications in a macro-comparative setting, where replication usually takes the form of verifiability and robustness testing. Replication and the usage of macro-comparative secondary survey data entail constraints on researcher choice that enable us to observe researcher error that is not based on deliberate choices made by researchers.

We draw on data from an experiment within a large-scale crowdsourced replication effort involving 85 independent research teams. Teams were randomly assigned to one of two conditions that systematically vary the degree of constraint experienced by the involved researchers. The first condition involved instructions to replicate a published study under full information about the original study; the second condition involved instructions to replicate an anonymous version of the same published study under limited information about the original study. All analyses reported in this paper have been pre-registered on the Open Science Framework before the beginning of data collection. In the pre-analysis plan we hypothesize (a) that routine researcher variability will occur in the studied replication setting and (b) that the magnitude of routine researcher variability will depend on the amount of information provided to replicators. We provide a direct test of the hypothesis that less information about an original study (i.e. deficient research documentation practices) will reduce the constraints for researchers throughout the process of replication and thus create more opportunities for routine variability to impact their research procedures and substantive conclusions. We generate systematic summaries of replication results generated by the independent researcher teams and use these to estimate the degree of routine researcher variability amongst them. The results support both hypotheses: In both experimental conditions, replicators vary considerably in the numerical results they produce, but this variability of replication outcomes is significantly higher in the reduced-information condition (i.e., with fewer constraints on replicators). We follow up on these findings with simulations to demonstrate how the observed patterns of routine researcher variability may affect the reliability of conclusions political scientists draw from their data in replications in macro-comparative research settings. We also offer solutions to the problems of routine researcher variability in such settings and provide guidance regarding the number of replicators needed to alleviate reliability concerns in macro-comparative political inquiry. Finally, we discuss the extent of the problem in the wider context of empirical political research and practical implications of our findings for reporting practices that are in line with the goal of maximizing the replicability of published political research.

Authors