Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
Unsupervised text analysis requires validation for each new data set and each new model. Ignoring this step can lead to substantively wrong conclusions based on flawed or even meaningless results. However, researchers rarely take the needed steps to validate their topic model results in a transparent and replicable way, instead relying on subjective interpretations of model outputs (e.g., top words or word clouds) and qualitative judgements about the topic meanings. This article builds on previous work in computer science (Chang et al. 2009) using crowdsourcing to efficiently and transparently validate topic models using inexpensive human resources on Amazon Mechanical Turk (AMT) via the SentimentIt platform (Carlson and Montgomery 2017). Unlike previous work, our framework validates the topics as well as the substantive topic labels assigned by researchers. Furthermore, we design and test multiple task structures which are recommended for different substantive needs and appear to be both easier for workers to complete and provide better discrimination across models than those provided in Chang et al. (2009). We also take additional steps to evaluate the reliability of worker judgments in these tasks. We illustrate this method with a novel analysis of Facebook posts by members of Congress and provide easy-to-follow examples that will enable researchers to apply this method to their own text-analysis tasks. Our aim is that by providing new tools, researchers will take model validation more seriously and eschew the current practice of providing readers with no ability to verify the meaning of their model results.
Jacob M. Montgomery, Washington University in St. Louis
Luwei Ying, Washington University in St. Louis
Brandon Michael Stewart, Princeton University