Individual Submission Summary
Share...

Direct link:

Estimating the Presence and Volume of Duplicate Voting

Fri, August 30, 2:00 to 3:30pm, Hilton, Columbia 11

Abstract

Immediately after the 2016 presidential election, President Trump claimed “millions and millions” of people voted multiple times which ultimately cost him the popular vote. A few months later, the Kobach commission was formed. Despite former commissioner Matthew Dunlap (D-ME) claiming little evidence was found supporting President Trump’s claims, Kris Kobach asserted that voter fraud was widespread and Dunlap was “willfully blind to the voter fraud in front of his nose.” In April 2018, President Trump doubled down on his original claim arguing that “millions and millions of people” vote more than once. A day before the most recent midterm election, President Trump again alluded to a voter fraud crisis saying “There are a lot of people that try and get in illegally and actually vote illegally…” Polling shows many of these claims may be hitting home, especially among Republicans where a majority still (52%) view voter fraud as a major issue.

The present study estimates the volume of duplicate voting and considers whether there are any indications of strategic behavior. For example, organizations, like the Brennan Center and Project Vote, have reviewed instances where citizens have been charged with voter fraud, but noticeably miss cases were voter fraud was never prosecuted. Scholars at Arizona State University took a similar approach, but instead of relying on popular accounts of voter fraud prosecutions, they contacted election officials and asked for known incidents of voter fraud. Here too, there is a strong selection effect which makes it difficult to estimate the actual duplicate voting rate. Cottrell et al., Christensen, and journalists at the Washington Post avoid the issue of selection by using the turnout of various groups to estimate the likelihood of voter fraud, but in doing so introduce an ecological fallacy since strong assumptions have to be made in order to tie aggregate turnout rates to individual-level behavior.

The best studies have approached the problem by matching various voting records to determine either the number of deceased voters who cast a vote or the duplicate voting rate. Here, detecting duplicate voting is characterized as a birthday problem. Although birthday probabilities can be calculated, the same cannot be said for first names, last names, and dates-of-birth which makes it difficult to use the same approach to estimate the probability of duplicate voting. For example, McDonald assumes birth years are uniformly distributed between 18 and 81, but we know some states are older than other and that birth years are often clustered within states. Similarly, first and last names are also not uniformly distributed, meaning we may be able to estimate the probability that two individuals should have the same birthdate, but we cannot easily estimate the likelihood of two individuals sharing the same name.

We overcome these limitations by estimating the probability distribution associated with first name, last name, and date-of-birth triads using 146,623,919 registered voters distributed across 9,875,306 county dyads. We then compare the county-level duplicate voting rate obtained from a national voter file and the voter files from the majority of the states to estimated probability distributions obtained from 9,875,306,000 bootstrapped samples. Using our approach, we can determine with various levels of statistical precision when and where duplicate voting occurs. Estimating the duplicate voting rate not only speaks to an important political debate, but it also allows us to determine whether there are any characteristics associated with duplicate voting – something noticeably lacking from the current popular and academic literature.

Authors