Individual Submission Summary
Share...

Direct link:

How Workflow Design and Political Framing Shape LLM Assessments of Equity in Climate Action Plans

Saturday, November 7, 10:15 to 11:45am, Property: Boston Marriott Copley Place, Floor: 5th Floor, Room: New Hampshire

Abstract

Governments are rapidly adopting large language models (LLMs) to review policy documents, yet whether these systems can support accountable and defensible interpretation in public administration remains unclear. Policy texts are both technical blueprints and political artifacts, and their interpretation carries distributional consequences. Treating LLMs as plug-and-play tools therefore risks embedding unexamined algorithmic discretion into substantive policy decisions. This study asks how workflow design—specifically evidence retrieval, prompt protocol, and political framing—shapes LLM reliability on normatively contested criteria.
    Using commercially available LLMs, we analyze 55 U.S. city climate action plans across six equity criteria, using prior expert coding as the benchmark, and generate 19,800 outputs through experimental variation in retrieval, prompting protocol, and political framing. We compare a standard retrieval pipeline with a customized retrieval system, a baseline zero-shot protocol with a structured evidence-bound protocol, and default, progressive, and conservative framing conditions.
    Results show that LLM performance is not solely an inherent property of the model. Stronger retrieval alone does not improve agreement with expert coding under a low-protocol workflow, and structured prompting yields little benefit when paired with weak or non-diagnostic evidence. Substantial gains emerge only when higher-quality retrieval is paired with a rigorous interpretive protocol that forces the model to ground its judgments in retrieved passages. Under this integrated workflow, agreement with expert coding reaches 0.97, compared with roughly 0.70 to 0.76 in the other conditions. Political framing induces threshold drift: conservative framing disproportionately under-affirms equity commitments (false negatives), while progressive framing skews toward over-affirmation.
    These findings indicate that LLM-based government document analysis is not a neutral plug-and-play utility; rather, it is a workflow outcome shaped by what evidence is surfaced, how interpretation is constrained, and how the task is framed. Achieving trustworthy AI-assisted analysis therefore requires rigorous design attention to retrieval and interpretive protocols, together with proactive controls for framing sensitivity.

Authors