Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Browse Artificial Intelligence Presentations
Program Calendar
Sign In
Search Tips
Background
Interviews are a primary tool for capturing variation in policy implementation across contexts, including multi‑site programs. However, scaling qualitative interview analysis is time‑intensive and often obscures analytic decisions about how themes are constructed, validated, and revised. Recent AI-assisted approaches promise to scale qualitative analysis, but many treat large language models (LLMs) as autonomous coders or apply them to isolated analytic tasks, risking automation bias and obscuring interpretive judgment (Lee et al., 2024; Prescott et al., 2024). This study advances an alternative approach grounded in reflexive thematic analysis, where themes are understood as patterns of shared meaning across a dataset rather than frequency-based code aggregation (Braun & Clarke, 2006, 2019, 2022).
Research Questions
This study addresses three research questions: How can LLM-assisted interview analysis be structured to preserve human interpretive authority, transparency, and traceability from evidence to themes? How can sentiment be incorporated without conflating evaluative tone with thematic meaning? How do domain experts evaluate the coherence, distinctiveness, and evidentiary adequacy of themes proposed through an LLM-assisted workflow?
Setting and Data
The empirical illustration draws on semi-structured interviews conducted across multiple implementation sites within a federally funded initiative. Interviews followed a shared protocol while allowing conversational flexibility to capture site-specific conditions (Potter & Hepburn, 2005). Although the substantive context is educational, the analytic challenges mirror those faced in state and local policy research, where consistent protocols intersect with heterogeneous implementation environments.
Research Design
We implement a corpus-first, human-in-the-loop thematic analysis pipeline that separates analytic tasks across distinct stages to reduce automation bias and improve auditability (Naeem et al., 2025). The workflow includes: (a) corpus-level pattern detection and theme construction, (b) structured theme review and consolidation, (c) multi-label response-to-theme assignment, and (d) independent sentiment modeling. Interview transcripts were structurally prepared through segmentation into coherent meaning units, de-identification, normalization, and protocol-aligned mapping to reduce context overload and theme blending (Hayes, 2025; Parfenova et al., 2024). LLM prompts encoded analytic assumptions, protocol intent, and constraints using staged, schema-constrained prompting with required quotation-based evidence for all themes (Schroeder et al., 2024). LLM outputs were reviewed by human analysts at each stage to maintain interpretive control. Sentiment analysis was conducted independently at the response level using a RoBERTa-based model (Liu et al., 2019) and aggregated post hoc to the theme level, ensuring that affective tone informed interpretation without shaping thematic meaning (De Paoli, 2024; Fischer & Biemann, 2024).
Findings/Results
The workflow produced coherent, well-differentiated themes within each protocol domain, each supported by multiple linked excerpts. Separating sentiment modeling from theme construction preserved thematic integrity while providing descriptive context regarding stakeholder experiences. Expert review indicated that the pipeline enhanced analytic transparency, reduced individual analyst burden, and supported systematic comparison across sites. Implications For policy scholars examining variation across states and systems, this approach offers a scalable and transparent method for analyzing interviews while preserving qualitative rigor. As federal authority contracts and state policy divergence grows, such methods are increasingly important for making policy implementation differences visible, comparable, and analytically defensible.