Individual Submission Summary
Share...

Direct link:

Comparing Human and LLM Judgment in Policy Value-Conflict Scenarios

Friday, November 6, 8:30 to 10:00am, Property: Boston Marriott Copley Place, Floor: 4th Floor, Room: Yardmouth

Abstract

Generative AI and large language models (LLMs) are increasingly discussed as potential tools for public administration, from document drafting to decision support. Yet existing research emphasizes that human oversight remains essential when decisions involve competing public values. This raises an important empirical question: how do LLM responses reflect value prioritization in policy scenarios involving value conflict?

This study examines how humans and LLMs respond to policy scenarios involving value conflict. It compares responses generated in users' personalized AI environments with responses generated under controlled conditions designed to reduce personalization effects. A key contribution of this study is that it collects LLM responses directly from participants' own AI accounts rather than relying only on researcher generated outputs. This design allows the study to examine whether interaction history, memory settings, and custom instructions are associated with variation in LLM responses.

The study uses an original online vignette-based survey with three respondent groups: social science graduate students, public sector practitioners, and academic or research personnel. Each respondent evaluates four policy scenarios: pension reform, minimum wage adjustment, infectious disease response, and energy transition. These scenarios capture recurring policy trade-offs that are salient across countries and relevant in the Korean context. Together, they cover both distributive and regulatory value conflicts.

Each respondent completes the survey under three conditions: direct human response without AI assistance, personalized LLM response generated in the respondent's usual AI environment, and controlled-condition LLM response generated under researcher specified instructions designed to reduce personalization effects. The survey collects scaled judgments on value continua, open-ended justifications, compensatory policy proposals, and detailed information on participants' AI environments.

Data collection is ongoing. The study tests whether controlled-condition LLM responses display lower variance than human or personalized LLM responses. It also examines whether personalized LLM responses align more closely with each respondent's value priorities, and whether these patterns differ across respondent groups and policy domains.

By treating personalized AI environments as an empirically observable condition rather than a theoretical assumption, this study contributes to debates on LLM personalization and homogenization, value conflict in AI-assisted policy judgment, and the role of generative AI in policy settings involving conflicting public values.

Authors