Individual Submission Summary
Share...

Direct link:

Visual Clustering for Exploration, Stratified Random Sampling and Cost Reduction

Fri, September 1, 10:00 to 11:30am PDT (10:00 to 11:30am PDT), Virtual, Virtual 7

Abstract

Abstract

Compared to text and audio, images can be an especially effective form of political communication. It has become relatively easy to automatically label images for many features of interest (such as protests, famous people or facial expressions). As a result, scholars are increasingly using large-N image analysis to investigate contemporary political attitudes and behavior. We address three emerging needs of image scholarship. First, researchers may want to visually explore an image corpus to discern patterns before they begin assigning labels. Second, they may want to annotate images for the presence of complex theoretical mechanisms that cannot be easily assigned using existing automated methods. Third, they may be primarily interested in studying human annotation decisions. We demonstrate how unsupervised image clustering can help researchers address each of these needs when dealing with large unbalanced image corpora. In addition, we assess the performance of many unsupervised image clustering techniques and choices available to researchers. Finally, we illustrate the usefulness of the method using two original image datasets.

Description

The proliferation of digital media and mobile communications, in combination with advances in deep learning, has made the study of visual information an increasingly important domain of social science research. A growing number of scholars are turning to convolutional neural networks (CNNs), a type of supervised deep learning algorithm, to study large image corpora. We identify three important limitations of using supervised learning CNNs for image labeling. We then explore how different unsupervised image-clustering approaches can help to address these limitations.

The first limitation is that, in order to train a CNN, researchers need to know ahead of time the particular objects or features they are looking for in the images, which may not always be the case. Once trained, supervised algorithms automatically assign the same universe of labels to other images, but it will only do so for the features it has been trained to predict. Unsupervised clustering can help researchers to discover unknown relevant features, so that they can later train a supervised classifier that accurately measures their quantity of interest.

A second limitation is that CNNs may not be very good at labeling images for complex theoretical constructs. For example, existing research demonstrates that CNNs can successfully capture the emotions reflected in faces present in an image. However, they do less well in predicting the emotions viewers feel when viewing an image. In such cases, manual annotation may be the preferred option.

A final limitation is that at some point in an image labeling project, researchers may be primarily interested in the human label generating process. For example, they might want to know whether particular groups of people (i.e. Republicans v. Democrats) have systematically differing

emotional reactions to the same images or it may be important to know whether some people are better at spotting instances of visual misinformation. For these last two limitations, particularly when working with unbalanced data (which is often the case in the social sciences), unsupervised clustering can help researchers draw stratified image samples for manual annotation.

There are many decisions one needs to make when building an unsupervised image-clustering algorithm, including: (a) how to numerically represent the images (e.g. using a pre-trained model to generate image embeddings; whether to fine-tune the pre-trained model first; if so, how to fine-tune the model, etc.), (b) the particular clustering algorithm to use (e.g. k-means, hierarchical, etc.), and (c) how to assess the performance of the algorithm. We illustrate the performance of different clustering options on two original image corpus: a dataset of Twitter messages/images related to a wide range of social movements, and a dataset of images from news articles.

Authors