Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Browse Artificial Intelligence Presentations
Program Calendar
Sign In
Search Tips
As large language models (LLMs) are increasingly deployed in policy-relevant contexts, understanding how they generate ethical judgments has become a central concern for AI governance. This study systematically examines how LLMs evaluate AI applications classified as Unacceptable Risk under the European Union’s AI Act (EU AI Act), using the Act’s risk-based framework as a normative benchmark. By comparing context-free, U.S.-framed, and Chinese-framed prompts, the study assesses whether LLM ethical reasoning remains stable across governance contexts or exhibits systematic divergence shaped by regime-specific values and institutional logics.
The analysis evaluates six widely used LLM systems, including ChatGPT, Gemini, and Claude, across seven policy-relevant scenarios drawn from the EU AI Act’s Unacceptable Risk category, including social scoring systems, real-time biometric surveillance, and manipulative AI targeting vulnerable populations. A mixed-methods approach combines Likert-scale evaluations with text-based analyses, including keyword frequency analysis (TF-IDF), sentiment analysis, and topic modeling (LDA), to capture both the direction and substance of ethical justifications across governance contexts.
The findings support two key conclusions. First, in context-free settings, all evaluated LLM systems consistently reject Unacceptable Risk applications, indicating stable baseline ethical judgments broadly aligned with EU AI Act principles. This convergence suggests that shared normative commitments, embedded through training data and alignment processes, produce consistent ethical orientations in the absence of explicit governance framing. Second, when governance context is introduced, ethical evaluations diverge meaningfully in high-salience domains. U.S.-framed prompts emphasize individual rights, privacy, and civil liberties, while Chinese-framed prompts foreground societal harmony, public order, and collective welfare. These divergences are most pronounced in cases such as social scoring systems and real-time biometric surveillance, where underlying value tensions are most acute.
Beyond model behavior, the findings demonstrate that LLM ethical reasoning is both structured and context-sensitive, bounded by shared normative constraints and responsive to institutional framing. This has direct implications for AI governance, as regulatory frameworks such as the EU AI Act may not be interpreted uniformly when operationalized through LLM systems, particularly in cross-national policy environments. The results therefore underscore the need for governance approaches that ensure consistency, accountability, and transparency in AI-assisted decision-making while recognizing variation in institutional contexts.
Finally, the study offers policy-relevant implications for AI human capital development. Divergent ethical orientations across governance contexts correspond to distinct workforce trajectories. The U.S. governance context emphasizes auditing, transparency, and rights-preserving governance capacities, whereas the Chinese governance context prioritizes large-scale deployment, computer vision, and surveillance-related competencies. These differences suggest that AI governance is not only an ethical and regulatory challenge, but also a workforce and institutional capacity issue with long-term global implications.