Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Browse Sessions by Fields of Interest
Browse Papers by Fields of Interest
Search Tips
Conference
Location
About APSA
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Most recent work in political science attempting to classify text along various dimensions has relied upon supervised models where human-coded text is used to train a classifier to segment the rest of the corpus into relevant categories. While this approach often improves classification over dictionary approaches or simple bag-of-words models, it requires significant resources for coding training data (Barberá et al. 2021). This cost can grow exponentially if multiple classification tasks are needed (ex., categorizing salience, stance, and conflict in a text) or if the corpus involves multiple languages. Such costs and time limitations have made the production of multi-dimensional, multi-lingual text-derived datasets rare.
This paper examines the potential to use a new and popular Large Language Model (LLM), GPT-3, to classify newspaper articles in news articles on multiple dimensions and in multiple languages. GPT-3 is a generative language model built on a large transformer network and is trained on hundreds of billions of tokens. Compared to other transformer-based LLMs, such as ones based on BERT, GPT-3 claims to have expanded multilinguistic flexibility and better zero-shot classification performance. We evaluate the performance of GPT-3 on a human-coded dataset of news articles about government Covid response drawn from identical time periods for the United Kingdom, the United States, and Mexico.
In comparison to previous approaches, the performance of GPT-3 is equal to or better than other larger language models and significantly better than traditional machine learning-based classification models. Our results show GPT-3 regularly outperforms other classification models in terms of accuracy, precision, and recall, even with very small training sets. Moreover, using GPT-3 for zero-shot classification performs well for tasks such as sentiment analysis. These results appear to be robust to the language (English or Spanish) or moderate changes in model hyperparameters. GPT-3 can also be used to efficiently filter out irrelevant articles from the corpus, a task that has proved challenging using other methods.
The results of this study suggest that GPT-3 can be used to classify text data in both English and Spanish, and likely other languages, with a high degree of recall and precision. In addition, GPT-3 is able to achieve this performance with either no or minimal amount of training data. This has significant implications for political science research, particularly for projects that require the analysis of multilingual datasets. Large language models may also increase the number of researchers who are able to perform classification tasks with text data, as the resources needed in terms of both time and money are far lower than for traditional classification methods. Overall, we show that tools like GPT-3, most publicly noted for text production, are potentially highly valuable tools for researchers wanting to analyze text corpora across domains and languages.