Individual Submission Summary
Share...

Direct link:

Semi-Supervised Topic Models for Political Rhetoric and Survey Open-ends

Sat, August 31, 10:00 to 11:30am, Omni, Senate Room

Abstract

While topic models have become increasingly popular in political science, researchers often use them to classify documents without systematic validation or interpretation, leading to potentially misleading conclusions. Using elected officials’ social media posts from a three-year period and a nationally-representative open-ended survey corpus, we present two methodological contributions meant to address these problems. First, we apply semi-supervised topic models that allow researchers to iteratively develop and refine a set of topics, making them more coherent and interpretable by leveraging known keywords that cleanly map onto concepts of interest. We compare this model with existing alternatives and demonstrate that totally unsupervised models risk biasing results by over- or understating the prevalence of key concepts across documents. Second, we develop and deploy a validation strategy based on human coding to determine whether or not the topics accurately capture meaningful concepts. We show that topics based upon unsupervised methods often cannot be reliably used as coding instruments. Furthermore, for the subset of topics that we are able to code with high levels of inter-rater reliability, we find that simple word lists – specifically, the lists of anchor terms developed for each topic – agree more closely with human classifications and perform as well or better than the topics from the fitted model in nearly every case. These findings suggest that unsupervised topic models risk generating misleading quantities of interest, but that semi-supervised models with a small amount of human validation can help minimize those risks. We also show how these models make substantive empirical analysis possible at scale: we show that some topics in congressional social media posts are associated with much higher levels of online engagement.

Authors