Individual Submission Summary
Share...

Direct link:

Active Learning Approaches for Labeling Text

Sun, September 1, 10:00 to 11:30am, Marriott, Washington 5

Abstract

Supervised machine learning methods are employed at an increasing rate in political science. Such models require costly manual labeling of documents. In this paper, we introduce active learning, a framework in which data to be labeled by human coders is not chosen at random but rather targeted in such a way that the required amount of data to train a machine learning model can be minimized. We study its benefits using the example of text data. We perform a series of experiments to identify the conditions where active learning can reduce the cost of labeling text data. We perform these experiments on three corpora that vary in size, document length, and domain. We find that in all cases where the document class of interest is not perfectly balanced, researchers can label a fraction of the labeled documents one would need using random sampling (or `passive' learning) to achieve equally performing classifiers. We further investigate how varying levels of inter-coder reliability affect the active learning procedures and find that with low-reliability, active learning can lead to decreased performance.

Authors