Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Browse Artificial Intelligence Presentations
Program Calendar
Sign In
Search Tips
People are increasingly treating AI systems not merely as tools, but as intelligent agents they consult for subjective, value-laden decisions—from resource allocation and charitable giving to workplace and policy trade-offs. Despite rapid advances in benchmarking AI capabilities in domains such as coding, mathematics, and factual recall, there remains no widely accepted framework for evaluating whether these systems promote prosocial outcomes. This gap is especially consequential in light of longstanding work in public administration and social policy—particularly by scholars such as Peter Frumkin—which emphasizes the importance of public value, altruism, and institutional design in shaping socially beneficial outcomes.This paper introduces a novel prosocial benchmark designed to measure how AI systems navigate trade-offs between self-interest and the public good. The benchmark consists of 280 theory-grounded decision scenarios spanning domains such as public goods provision, redistribution, cooperation, and ethical dilemmas. Each scenario is constructed to elicit choices that reveal underlying preferences over individual versus collective welfare, allowing for systematic comparison across models and contexts.We evaluate five frontier LLMs using this benchmark and document three main findings. First, models exhibit a strong prosocial default on average, often favoring collective benefit over narrow self-interest. Second, there is substantial heterogeneity across scenario types, with models behaving differently in contexts involving fairness, efficiency, and strategic interaction. Third, we observe meaningful variation across models, suggesting that prosocial alignment is neither uniform nor guaranteed, even among leading systems.By providing a standardized and extensible instrument, this benchmark enables developers, researchers, and policymakers to assess AI systems along a critical but underexplored dimension: their orientation toward public value. It complements existing capability and safety evaluations and offers a foundation for future work that compares AI behavior to human judgments. Human benchmark data and longitudinal extensions are forthcoming, with implications for AI governance, alignment, and the design of socially beneficial AI systems.