Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Search Tips
Virtual Exhibit Hall
Personal Schedule
Sign In
X (Twitter)
We consider the properties of word embeddings, a 'deep learning' technique for representing semantic meaning---the popularity of which has exploded in recent times. While political
scientists are rightly attracted to the potential of such methods, there is little to guide them on fundamental questions concerning the fitting of,
and inference about, these models. This paper provides an overview of embeddings conceptually, and then provides advice for practitioners based on experiments we conduct on political corpora
of various sizes and in various languages. In particular, we explore key parameters choices regarding the size of the vocabulary, 'context window' and vector length, showing in each case how it affects the difficultly and speed of the estimation process, along with the quality of the inferences one is able to draw in each case. Overall, with important caveats, we show that
findings are robust to such choices for large corpora. We also report generally positive results on the technical and substantive 'human' performance of readily available pre-trained---as opposed to locally fit---embeddings for political science queries. Our validation framework is not architecture specific, and as such can be used to evaluate new embedding architectures as they emerge in this rapidly growing area.