Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
Twitter is a rich source of material for the study of political phenomena. We apply word embedding methods to a selection of users who Tweeted some observation about the 2016 election in order to study how incident observations vary by partisan association. Our indicators for partisan association are the cosine similarities between users’ self-descriptions and a set of keywords that refer to the major parties' presidential campaigns. To validate the cosine similarities as reflecting "partisanship,” in some sense, we assess how well the cosine similarities predict other partisan Twitter behavior, such as which political accounts each user retweets or favorites. We use a loss function based on the cosine similarities' success predicting partisan retweet behavior to choose the best set of preprocessing steps to use on the text from which the word embedding vectors are computed. We use a finite sampling argument to motivate averaging the results from multiple runs of the word embedding algorithm to minimize estimation error in the cosine similarities. We also use the word and document vectors estimated from Twitter descriptions to predict document vectors using comment text from partisan subreddits (e.g. “the_donald” and "hillaryclinton") on the popular media aggregator Reddit. Cosine similarities derived from these predicted document vectors successfully predict the partisan bent of the subreddit the comments are taken from. Observing that the word and document vector space produced by Twitter predicts partisan behavior on another social media site provides further validation of this methodology.
The method used above turns to the Twitter user's description, which is a user-provided biography. The technical details of method used to generate the cosine similarities and partisan association scores is as follows: we first obtain document vectors for each user description and word vectors for every unique word across the descriptions using doc2vec. We then compare these user descriptions to a selection of partisan keywords, such as "clinton," "trump," "democrat," "republican," etc. by calculating the cosine similarities between the document vectors and the word vectors for these keywords. These cosine similarities can be used as the dissimilarity matrix for metric unfolding, with the first dimension representing the partisan association scale. In contrast to the state-of-the-art approaches of estimating these associations, our method analyzes text instead of analyzing which political actors and partisan entities a user follows. These latter approaches often make unrealistic assumptions about why a user chooses to follow someone, require dealing with Twitter API's rate limit, and are difficult to update with changing political environments, such as users purposely not following "mainstream" politicians.
The informal idea of this proposed method is to recognize patterns in the words surrounding explicit partisan associations of some users and to use that to estimate each user's partisan association. The logic is that even if someone does not explicitly describe or express a partisan association, they may describe themselves in the fashion of other users who do explicitly express partisan associations. Doc2vec captures the meaning of words and topics of documents by predicting its linguistic contexts, making it a well-suited algorithm for this proposed method. As a toy example, let us assume that several users describe themselves as "latte sipping Clinton supporters" in their user descriptions on Twitter. Several users may also simply describe themselves as "latte lovers." But the cosine similarity between the user descriptions of the latter group to the "clinton" word embedding will be high, because other user descriptions that explicitly mention the word "Clinton" often have "latte" within their linguistic contexts. On the other hand, let us assume that not many users explicitly mention supporting Trump and enjoying lattes in their user descriptions. Thus, the cosine similarity between the user descriptions that mention the word "latte" and the "trump" word embedding will be lower. Thus, we can estimate the partisan associations of expressed textual elements of users even if the user does not explicitly express a partisan association. Our methodology has the advantage over other similar methodologies in that it only requires some, but not all, Twitter users to be explicitly partisan.
We also compare our results to other methods that measure partisan association, such as Pablo Barbera's estimation of ideology based on Twitter networks. For example, we find that of the 215,247 users that make some observation about the 2016 election, 194,336 users had user descriptions. On the other hand, only about half of the users followed the minimum number of politicians needed to have their partisan association estimated using Barbera's method.
Patrick Y. Wu, University of Michigan, Ann Arbor
Walter R. Mebane, University of Michigan, Ann Arbor
Joe Klaver, University of Michigan
Logan Woods, University of Michigan
Alejandro Pineda, University of Michigan
Blake Miller, London School of Economics and Political Science