Individual Submission Summary
Share...

Direct link:

Download

Harnessing Accurate Bias in Large Language Models to Explore Public Opinion (Pre-Recorded)

Thu, September 30, 8:00 to 9:30am PDT (8:00 to 9:30am PDT), TBA

Abstract

In the machine learning community, a language model (LM), such as GPT-2 (1), BERT (2), T5 (3), or GPT-3 (4), is a conditional probability distribution over words. When trained at scale, LMs exhibit a remarkable ability to capture the syntax and semantics of natural language and show state of the art performance on a wide variety of natural language tasks, such as text summarization (5), conversation (6), text generation (7,8), and zero-shot learning (9).

The machine learning community rightfully worries about mitigating algorithmic bias in LMs (10,11). However, striking similarities between the statistical learning algorithms used by LMs to predict the next word/token, and theories of statistical learning in human cognition (12) to describe how the brain predicts the next moment (13,14), provide reason to hope that LMs can be profitably used as a model, even if still far from a technical match, for human cognition. Furthermore, LMs show surprising sensitivity to detail, and it is therefore possible that bias in LMs is not always “simple'” (such as assuming all doctors are male), but a complex and nuanced bias that is accurately intertwined with time, geography, personal identity, and cultural milieu (15).

In this paper, our interdisciplinary team (3 computer scientists, 3 political psychologists) presents initial evidence for what we call “accurate bias” in the largest LM available, GPT-3. We show that when properly conditioned, GPT-3 capably identifies and reproduces the complex patterns of relationships between ideas and attitudes identified in humans. We first show that we can successfully “condition” GPT-3 on a mix of traditional political science demographics (location, political ideology, ethnicity, sex, education, etc.) to create large “samples” of different (silicon) “individuals.” We then show that distributions of responses from these “individuals” to a variety of American National Election Study (ANES) questions do not differ systematically from human ANES responses to these same questions. Moreover, the degree of correlation between GPT-3 responses to variables one might expect to be correlated (church attendance and vote from Trump in 2016, for example), largely matches that in the human data.

We then show GPT-3 capable of producing text outside the bounds set by the structure of a survey that is “accurate” enough that humans cannot tell the difference between it and real human text. Busby et. al.’s “Pigeonholing Partisans” experiment asked individuals to list four words to describe Republicans and Democrats (one of which is their political ingroup, the other their outgroup). We generated a GPT-3 sample that matched the demographics of this human sample and asked it to the same thing. We then randomly presented both the human and GPT-3 word lists to over 3000 human coders and asked them to identify the political ideology of the individual who wrote the words. To pass as human, the free text generated by GPT-3 would need to show “accurate bias” for the specific type of individual for which it was conditioned listing words that in tenor and content mirror those listed by humans with a particular partisanship. In such tasks, humans naturally group certain types of words together to indicate their overall feeling towards the group. Thus, for word lists produced by humans, the partisanship of the writer should be generally discernible by our human coders. We asked: is GPT-3 capable of producing text that replicates these specific biases, such that our human coders cannot tell the difference? Our results suggest yes.

We conclude the paper with a discussion of the import of these findings for the study of political psychology. In particular, we note that unlike human brains, we can “open the black box” of these silicon brains, using methods to identify the patterns of associations, and their relative weights, behind the responses we present in the paper. These insights can then be used to both generate new theory and expand on old, making significant contributions to the study of human psychology across a variety of fields.

1. A. Radford et al., OpenAI Blog 1 (2019).
2. J. Devlin et. al., arXiv:1810.04805 (2018).
3. C. Raffel et. al., Journal of Machine Learning Research 21 (2020).
4. T.B. Brown et. al., arXiv:2005.14165 (2020).
5. C. Raffel et. al., Journal of Machine Learning Research 21 (2020).
6. D. Adiwardana et. al., arXiv:2001.09977 (2020).
7. Z. Dai et. al., arXiv:1901.02860 (2019).
8. N.S. Keska et. al., arXiv:1909.05858 (2019).
9. B. Krause et. al., arXiv:2009.06367 (2020).
10. W.L. Hamilton et. al., arXiv:1605.09096 (2016).
11. N. Garg et. al., Proceedings of the National Academy of Sciences 115, E3635 (2018).
12. A. Schapiro et. al, Brain mapping 3, 501 (2015).
13. M. Lodge and C. S. Taber, The Rationalizing Voter (Cambridge University Press, 2013).
14. L. F. Barrett, How Emotions Are Made (Houghton Mifflin Harcourt, 2017).
15. A. Caliskan et. al., Science 356, no. 6334 (2017): 183-86.

Authors