Personal Schedule
Change Preferences / Time Zone
Sign In
Search
Browse By Day
Browse By Time
Browse By Person
Browse By Mini-Conference
Browse By Division
Browse By Session or Event Type
Browse Sessions by Fields of Interest
Browse Papers by Fields of Interest
Search Tips
Location
About APSA
Bluesky
X (Twitter)
INTRODUCTION: Unsupervised topic modeling for cross-country political texts poses substantial challenges. Even when texts are translated into a common language, country-specific lexicon often overly influences text representation, producing clusters driven by geography rather than substantively meaningful policy or political content. Similar problems arise in single-country studies when regional or local markers overwhelm thematic structure. Supervised approaches are unsuited for exploratory analysis and topic discovery, as they require ex ante specification of relevant categories. Generative large language models face related limitations, as effective prompt engineering presupposes prior knowledge of what researchers seek to extract. These challenges motivate the development of unsupervised methods that improve cross-national comparability while preserving interpretability.
OBJECTIVE: This paper provides methodological guidance for applying embedding-based topic models in comparative political research, building on BERTopic. It introduces a simple debiasing strategy to improve cross-national comparability and extends standard topic modeling pipelines with a spectral method to recover an interpretable measure of within-topic polarity.
DATA: The study analyzes a corpus of party manifestos translated into English by the Comparative Manifesto Project (CMP), covering ten countries with diverse institutional settings and levels of economic development. The sample includes Argentina, Brazil, Canada, Germany, Mexico, Portugal, Spain, Sweden, the United Kingdom, and the United States. I process texts into paragraph-sized chunks and use CMP macro-topic codes for external validation.
METHODS: I encode text chunks using pretrained sentence-transformer embeddings and identify micro-topics by combining nonlinear dimensionality reduction with density-based clustering, following the approach used in BERTopic and Top2Vec to organize high-dimensional text representations. My main methodological contribution is a debiasing step that removes country-specific structure from the embedding space by projecting out principal components derived from country-level embedding centroids. I then aggregate micro-topics using hierarchical clustering over topic representations. To enhance interpretability, I also compute the Fiedler vector within each macro-topic, using spectral decomposition of a similarity graph to estimate a continuous polarity dimension.
RESULTS: Country-subspace debiasing improves both external alignment and internal topic quality relative to an unadjusted baseline, with the strongest gains emerging when I remove two principal country components. While overall agreement with CMP labels remains modest, debiasing consistently strengthens alignment and produces more coherent macro-topics. Within macro-topics, Fiedler-vector scoring yields substantively interpretable polarity axes. For example, the polarity measure can distinguish text segments that emphasize environmental protection from those that focus on infrastructure expansion.
CONTRIBUTION: The paper offers a lightweight and reproducible extension to embedding-based topic modeling pipelines that improves cross-national comparability without requiring supervised labels or prompt engineering. By combining country-subspace debiasing with spectral polarity estimation, the approach allows researchers to recover meaningful within-topic structure. The paper also provides guidelines on how to use a similar pipeline with static document embeddings for researchers with access to more limited computational resources or larger corpora.