Search
Browse By Day
Browse By Time
Browse By Person
Browse By Committee or SIG
Browse By Session Type
Browse By Keywords
Browse By Geographic Descriptor
Search Tips
Personal Schedule
Change Preferences / Time Zone
Sign In
The increasing integration of artificial intelligence (AI) into educational contexts has raised urgent questions about fairness, transparency, and the potential for algorithmic bias (Alrawashdeh & Castillo, 2025; Dwivedi et al., 2023; Farrelly & Baker, 2023). Automated writing assessment tools promise efficiency and objectivity, yet also risk reproducing or amplifying social and linguistic inequalities embedded in training data or activated by identity cues (Blodgett et al., 2020; Liang et al., 2023; Perkins et al., 2021). This study investigates how the structure of user inputs, specifically the inclusion of demographic information and rubric-based scoring criteria, shapes the fairness and consistency of AI-generated writing scores across diverse student populations.
We draw on data from Project Personality, a large randomized controlled trial examining adolescent mental health and mindset interventions (Schleider et al., 2022). The dataset includes free-response writing samples from 652 adolescents aged 13 to 16. Students responded to three open-ended prompts related to personal growth and social-emotional experiences. For this study, the three responses were concatenated into single writing samples per student to create a more robust, general-purpose text appropriate for evaluation. Responses were scored by two publicly available large language models (LLMs): OpenAI’s ChatGPT-4o and DeepSeek-V2. These models were selected based on their accessibility and popularity in classroom and research settings, reflecting real-world usage by non-technical educators.
Scoring was conducted under four input conditions: (1) response-only (no rubric, no demographics), (2) response + demographics, (3) response + rubric, and (4) response + both rubric and demographics. Each condition was operationalized through a standardized prompt presented to the AI model, asking it to assign a score out of 100 and provide justification. The five-criteria, 100-point rubric used in conditions 3 and 4 was adapted from the 6+1 Traits of Writing framework and rubrics published by the National Council of Teachers of English (NCTE). Rubric criteria were scored on a four-point scale, and 4 human raters calibrated a subset of responses (5%) to establish internal rubric reliability (ICC = .87). All AI sessions were run within a fixed scoring window (May 2025), and histories were cleared between sessions to prevent context bleed.
The sample was predominantly assigned female at birth (88.4%) and included substantial representation of transgender and gender-diverse youth (31.5%) and LGBQ+ students (79.2%). The dataset also included racial and ethnic diversity. Because the original study did not collect data on socioeconomic status (SES) or religion, we simulated these variables to enable broader bias testing. SES categories were modeled based on access to community resources (e.g., healthcare, specialty stores), and religious affiliation reflected major world traditions and secular/cultural groups. Simulation is common in algorithmic fairness research when sensitive data are unavailable (Buolamwini & Gebru, 2018; Obermeyer et al., 2019). While simulated demographics cannot substitute for real-world measurement, they allow exploratory, hypothesis-generating insights into potential differential treatment.
Descriptive analysis revealed that average scores were generally high (M ≈ 73–82). ChatGPT-4o produced slightly higher scores than DeepSeek-V2 overall. Pearson correlations showed strong internal consistency across model rounds, with ChatGPT-4o ranging from r = .69–.81 and DeepSeek from r = .61–.76. However, cross-model correlations dropped significantly in Round 4 (r = .38), suggesting that the inclusion of both rubric and demographic inputs introduced score variability and reduced agreement between models. Notably, the highest agreement (r = .76) occurred in the rubric-only condition, indicating that rubrics may serve as a stabilizing force in AI evaluation.
Paired-sample t-tests comparing Rounds 1 and 2 showed that several groups experienced significant score increases when demographic information was included. These included White (Mdiff = 1.73, p < .001), Hispanic/Latinx (Mdiff = 1.75, p = .018), and Mixed-race students (Mdiff = 2.16, p = .005), along with transgender and gender-diverse youth (Mdiff = 2.16, p = .003), and LGBQ+ students (Mdiff = 1.65, p < .001). No significant changes were observed for Asian, Black, or “Other” racial categories, nor for male or nonbinary participants. Score increases were also observed across religious groups (e.g., participants observing Christmas/Easter and Passover/Hanukkah) and SES strata. A linear mixed-effects model confirmed that input structure significantly influenced AI-generated scores (χ²(3) = 414.22, p < .001). Including demographic information predicted significantly higher scores (B = 2.22, p < .001), suggesting an unintended leniency effect when identity cues are present. Conversely, rubric inclusion was associated with lower scores (B = –3.04, p < .001), indicating a more precise and structured evaluative approach. When both types of input were used together, effects offset (B = 0.50, p = .054), revealing no significant difference from the response-only condition. Significant random effects for AI model (χ²(1) = 58.05, p < .001) and student-level interactions (χ²(1) = 20.09, p < .001) underscore the influence of both model architecture and individual variation.
These results suggest a shift in the nature of AI bias. Contrary to prior research that emphasized under-scoring of marginalized identities (Blodgett et al., 2020; Sap et al., 2019), our findings point to a potential overcorrection in newer models. Inclusion of demographic information may result in inflated scoring (a form of leniency bias or “fairness inflation”) that appears equitable on the surface but undermines construct validity. This emerging bias pattern complicates assumptions that algorithmic fairness merely requires reducing negative bias; it raises new concerns about inflated positive bias, especially when scoring outcomes are not fully grounded in the quality of the writing itself. In contrast, rubric-based scoring improved inter-model agreement and appeared to anchor evaluation in more consistent, transparent criteria. Yet rubrics also led to overall score reductions, which may disadvantage students whose writing reflects informal, culturally specific, or identity-expressive styles. These findings align with prior cautionary work on the unintended limitations of standardized rubrics in diverse classroom settings (Baker & Hawn, 2022; Memarian & Doleck, 2023). The divergence between models also highlights the importance of system-level validation. ChatGPT-4o was more sensitive to rubric variation, while DeepSeek-V2 was more reactive to demographic cues. This suggests that model architecture, training data, and alignment methods significantly shape fairness outcomes, an underexplored but critical consideration in educational AI research.
Ultimately, this study challenges overly simplistic narratives of bias reproduction in AI assessment. Rather than repeating past harms, some LLMs may now introduce inflated scoring for marginalized identities, in an effort to appear “fair.” While well-intentioned, such inflation risks distorting student evaluation and masking performance gaps. Our findings call for a more nuanced, evidence-based approach to AI fairness; one that avoids both penalization and ungrounded leniency.
As AI is increasingly integrated into educational systems, with minimal oversight or regulatory clarity, these insights carry significant implications. Teachers, school leaders, policymakers, and AI developers must understand how model inputs shape outputs. Input design matters. Based on this work, we recommend that scoring aids like rubrics be used independently of demographic information to prevent conflicting biases. Moreover, AI-based scoring systems must undergo model-specific validation in equity-sensitive contexts. Future work will extend this analysis through human-AI scoring comparisons and content analysis of model-generated justifications to promote transparency, accountability, and learner-centered innovation in AI-driven education.