Individual Submission Summary
Share...

Direct link:

Replicable Privacy: Enabling Replication of Research on Sensitive Datasets

Thu, August 29, 8:00 to 9:30am, Marriott, Washington 2

Abstract

The data deluge produced by the mass adoption of the Internet has amplified the capacity of quantitative social science to conduct empirical analysis about the most relevant societal issues. This increased capacity can be illustrated by the use of data from social media platforms to understand political behavior, polarisation or censorship. However, the involvement of personal data introduces the need to comply with a complex set of restrictions, as illustrated by the data protection regulation of the European Union (GDPR). Regulatory compliance can restrain research by avoiding access to the data, limiting the capacity of researchers to share and combine datasets, and even impeding the release of datasets to replicate the results. This paper describes an approach to facilitate research on sensitive datasets building on a set of tools developed at the Institute for Quantitative Social Science, at Harvard University. First, we apply differential privacy to produce a data release which prevents re-identification of the individuals, while preserving the statistical properties of the datasets. Second, we use the Dataverse technology with a Datatags implementaion to facilitate uploading and sharing of the sensitive data. Third, we show the utility of this approach by releasing a privacy preserving version of the Internet Correlates, a dataset of Internet connectivity of unprecendented precision.

Authors