Session Submission Summary
Share...

Direct link:

Large Language Model (LLM) Evaluation and Use in AI Policy Research

Saturday, November 7, 8:30 to 10:00am, Property: Boston Marriott Copley Place, Floor: 5th Floor, Room: New Hampshire

Session Submission Type: Panel

Abstract

As large language models (LLMs) rapidly diffuse into research and policy settings, they are transforming not only the objects of governance but also the methods through which policy-relevant knowledge is produced. This panel brings together three complementary papers that examine how LLMs can be evaluated, deployed, and governed in the context of AI policy research. Together, they address a central challenge: how to ensure that LLM-enabled research and decision-making are reliable, interpretable, and aligned with public values.The first paper, Expanding the Scope of Systematic Literature Reviews with AI, develops and validates a scalable framework for integrating LLMs into systematic literature reviews (SLRs). Drawing on an application involving over 150,000 abstracts on AI and labor, the authors propose an ensemble-based screening approach that improves reproducibility and reduces stochastic variation in model judgments. Importantly, the paper advances a “humans-in-the-loop” design that combines LLM scalability with human validation, and provides a transparent, PRISMA-aligned protocol for social science applications. This contribution speaks directly to the growing need for methodological infrastructure in an era of exponential knowledge production.The second paper, Divergence in LLM Ethical Reasoning under the EU AI Act, examines how LLMs generate ethical judgments across governance contexts. Using the EU AI Act as a normative benchmark, the study shows that while LLMs converge in rejecting “unacceptable risk” applications in context-free settings, their reasoning diverges significantly when prompts are framed in U.S. versus Chinese institutional contexts. These findings highlight that LLM outputs are not neutral but reflect context-sensitive value trade-offs, raising important questions about cross-national consistency, regulatory interpretation, and the institutional embedding of AI systems.The third paper, The Missing Prosocial Benchmark, introduces a novel evaluation framework to assess whether LLMs promote public good over self-interest. Using 280 theory-driven scenarios, the authors demonstrate that while current models exhibit a broadly prosocial orientation, there is substantial variation across contexts and systems. By establishing a standardized benchmark, this work provides policymakers and developers with a tool to evaluate AI systems along normative dimensions that are currently under-measured.Taken together, the panel advances a research agenda for LLM evaluation in policy contexts along three dimensions: methodological rigor in AI-assisted research workflows, sensitivity of model outputs to institutional and cultural framing, and the measurement of normative alignment with societal goals. The discussion will highlight implications for evidence synthesis, regulatory design, and the development of evaluation standards, offering actionable insights for scholars and policymakers navigating the integration of LLMs into public policy research. The authors of the three papers are:Daniel Schiff, Purdue University, dschiff@purdue.edu; Luisa Nazareno, Virginia Commonwealth University, nazarenoll@vcu.edu; Lucas Wiese, Purdue University, lwiese@purdue.edu; Zeewan Lee, zeewan.lee@nus.edu.sg; Vasundhara Kaul, Purdue University, kaul9@purdue.eduHaeLee Kim, The George Washington University, haeleekim@gwu.eduJi Ma, IU Bloomington, ma47@iu.edu: Peter Frumkin, Oxford and Penn, peter.frumkin@new.ox.ac.uk The discussant is: John Nelson, Oregon State University, nelsoj34@oregonstate.edu

Policy Area

Chair

Discussants

Organizer

Individual Presentations