Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
Supervised machine learning methods are employed at an increasing rate in political science. Such models require costly manual labeling of documents. In this paper, we introduce active learning, a framework in which data to be labeled by human coders is not chosen at random but rather targeted in such a way that the required amount of data to train a machine learning model can be minimized. We study its benefits using the example of text data. We perform a series of experiments to identify the conditions where active learning can reduce the cost of labeling text data. We perform these experiments on three corpora that vary in size, document length, and domain. We find that in all cases where the document class of interest is not perfectly balanced, researchers can label a fraction of the labeled documents one would need using random sampling (or `passive' learning) to achieve equally performing classifiers. We further investigate how varying levels of inter-coder reliability affect the active learning procedures and find that with low-reliability, active learning can lead to decreased performance.
Blake Miller, London School of Economics and Political Science
Fridolin Jakob Linder, Pennsylvania State University
Walter R. Mebane, University of Michigan, Ann Arbor