Individual Submission Summary
Share...

Direct link:

Embedding Regression

Sun, September 13, 2:00 to 3:30pm MDT (2:00 to 3:30pm MDT), TBA

Abstract

Political scientists commonly seek to make statements about how a word's usage and meaning varies over contexts---whether that be time, partisan identity, or some other document-level covariate. A promising avenue is "word embeddings" that are specific to a domain, and that simultaneously allow for statements of uncertainty and statistical inference. We introduce the "a la Carte on Text embedding regression model" (ConText regression model) for this exact purpose. In particular, we extend and validate a simple model-based method of "retrofitting" pre-trained embeddings to local contexts that requires minimal input data and out-performs well-known competitors for studying changes in meaning across groups and times. Our approach allows us to speak descriptively of "effects" of covariates on the way that words are understood, and to comment on whether a particular use is statistically significantly different to another. We provide experimental and observational evidence of performance of the model, along with open-source software.

Authors