Individual Submission Summary
Share...

Direct link:

From Prompts to Policy: Advancing Fairness in Educational LLM Evaluation through Difference-Aware Design

Sun, March 29, 4:30 to 5:45pm, Hilton, Floor: Lobby Level - Tower 2, Plaza Room A

Proposal

Research Objectives
Large Language Models (LLMs) are rapidly being integrated into educational systems, supporting feedback generation, assessment, tutoring, and personalized recommendations. While their potential is immense, education presents complex equity challenges that go beyond fairness notions premised on treating all groups identically. In educational contexts, fairness requires equitable recognition of diverse learners’ needs, backgrounds, and cultural contexts. As LLMs increasingly shape decisions, their fairness characteristics demand rigorous evaluation to prevent reproduction of existing inequalities.

This study addresses two urgent challenges in educational AI. First, it investigates which fairness metrics best capture multifaceted equity concerns, particularly where group-specific needs should be acknowledged rather than erased. Second, it evaluates the fairness implications of different prompting strategies. We compare technique-based prompts, emphasizing input semantics and refinement, with process-based prompts, which involve iterative, collaborative construction aligned with learner characteristics. These objectives advance both theoretical insights and practical guidance for using difference-aware prompting to mitigate demographic bias in educational LLMs.


Literature Review
Algorithmic fairness in LLMs has conventionally emphasized a “color-blind” or difference-unaware approach (Wang et al., 2025). Yet ignoring group-specific differences such as linguistic diversity or learning variability risks deepening inequities. Prompts effective for native English speakers may disadvantage bilingual or neurodiverse students. Existing fairness metrics, equalized odds, disparate impact (Gao et al., 2024), capture broad disparities but fail to address context-specific needs. Recent studies on fairness-guided few-shot learning (Ma et al., 2023) and prompt interventions (Atwood et al., 2024) demonstrate promise, but their educational relevance remains underexplored. This project extends these advances by adapting fairness-aware prompt design to education.


Methodology
We employ a benchmark dataset tailored for educational LLM scenarios (Xu et al., 2025). Fairness is evaluated using machine learning metrics, conditional statistical parity, equal opportunity, treatment equality, alongside NLP-specific diagnostics such as DisCo, LPBS, and SEAT. Comparative experiments assess the impact of different prompting strategies, while clustering analyses explore whether combinations of strategies form profiles associated with stronger fairness outcomes.


Preliminary Results
Preliminary findings indicate that difference-aware prompting, particularly process-based strategies, improves fairness for underrepresented subgroups without compromising overall task performance. In contexts involving multilingual learners, difference-aware prompts reduced bias amplification by over 20% relative to baseline technique-based prompts. Results also reveal trade-offs: gains in fairness occasionally come with slight increases in response variability, underscoring the need for context-sensitive evaluation.


Significance and Policy Implications
Beyond technical contributions, this work speaks directly to education policy and governance. As policymakers develop accountability frameworks for AI in schools, our findings highlighted the need for difference-aware standards in LLM evaluation. Policies limited to generic bias audits risk overlooking culturally responsive fairness. We argue for policy guidelines that mandate (1) benchmarking AI systems with education-specific fairness metrics, (2) transparency in prompt design and testing across subpopulations, and (3) accountability mechanisms requiring vendors and districts to demonstrate equity outcomes.

By linking prompt engineering to policy decision-making, this study provides a domain-specific framework for advancing fairness in AI-powered education. It contributes to technical rigor and informs governance strategies that safeguard vulnerable learners in increasingly AI-mediated classrooms. (497 Words)

Authors