Search
Browse By Day
Browse By Time
Browse By Person
Browse By Committee or SIG
Browse By Session Type
Browse By Keywords
Browse By Geographic Descriptor
Search Tips
Personal Schedule
Change Preferences / Time Zone
Sign In
Background
Large Language Models (LLMs) have emerged as powerful tools in student writing assessment, praised for their ability to improve feedback, revision, and engagement (Chiu et al., 2023; Farrokhnia et al., 2024; C. Zhu et al., 2023). These AI systems are often viewed as neutral or efficient, yet writing is a social act, shaped by identity, audience, and culture. This is especially true for adolescents, for whom writing supports both academic growth and self-expression (Burgess & Ivanič, 2010). Critical algorithm scholars argue that AI systems reflect the norms and values embedded in their training data and design (Barocas et al., 2023; Bender et al., 2021). Sociolinguists point out that writing assessment often privileges standardized forms and disregards culturally rooted or expressive styles (Alim, 2011; Lillis, 2013). These insights suggest that AI feedback tools may not just score writing but influence how students’ voices are understood and valued. Despite calls for fairness in AI, much of the focus has been on outcome parity. Mahajan (2025) urges a shift toward ethical clarity, which involves looking at how models explain their judgments. This study takes up that call by analyzing the justifications LLMs give when scoring adolescent writing, with and without identity cues.
Purpose
Drawing on critical algorithm studies and sociolinguistic theory, this study examines how demographic information affects the way LLMs justify their writing evaluations. Rather than comparing only scores, we analyze the model-generated feedback to understand how values shift in response to identity cues. We ask:
• RQ1: How does the inclusion of demographic information reshape the rationale and values guiding AI scoring?
• RQ2: What evidence suggests that demographic data affect how LLMs interpret and evaluate writing?
Methods
We used 652 short writing samples from Project Personality, a self-guided intervention previously tested in a randomized trial (Cohen et al., 2022). Each response averaged 109 words. The same 652 texts were submitted to three LLMs: ChatGPT-4o, Gemini 2.5 Pro Preview, and DeepSeek-V2, under two scoring conditions:
• Condition 1: no demographic information provided
• Condition 2: with demographic cues constructed from available metadata: race, gender identity, sexual orientation, and simulated SES/religion using stratified randomization, as commonly done in fairness research (McDuff et al., 2018; Solnick et al., 2020).
Each model was prompted to assign a score from 0 to 100 and provide a short explanation. No rubric was provided to encourage open-ended reasoning. Model explanations were analyzed using reflexive thematic analysis (Braun & Clarke, 2006). Two coders per condition worked independently, blind to whether identity data had been included. Intercoder-reliability was high (κ = .84 and .76). Disagreements were resolved through consensus. Scores were compared using t-tests and descriptive statistics. Five case responses were analyzed in depth; with 3 selected for deviation from expected topics.
Findings
Overall, scores were higher when identity information was included. Gemini showed the largest shift (effect size = 1.01). Thematic coding (n=407 in Condition 1; 512 in Condition 2) revealed four primary themes in the model justifications:
1. Emotional Awareness/Expression
2. Response Quality/Effectiveness
3. Writing Quality/Structure
4. Interpersonal Growth/Resilience
In Condition 1, models emphasized emotional insight and connection. In Condition 2, the focus shifted toward grammar, structure, and clarity. References to emotional expression declined as attention to form increased. Case analyses revealed how identity information altered interpretation. Quinn, a secular, queer, transgender student, received higher scores with identity visible. Their informal and peer-like tone was praised as authentic and socially engaged. Mara, similarly identified but two years younger and Jewish, saw no change. Her concise and empathetic tone was dismissed as lacking development. DeepSeek cited a need for elaboration. Three outlier responses added further context. A blunt comment was praised for honesty when anonymous, but flagged as disengaged with identity visible. A fanfiction reference was seen as playful without identity, but labeled off-topic once race and religion were disclosed. These cases suggest that identity shapes how tone, humor, and relevance are interpreted.
Discussion
The findings support Yeung’s (2018) view of algorithms as regulatory systems that assign value based on embedded expectations. LLMs adjusted their criteria based on visible identity. When no demographic information was provided, the same writing was more often read as expressive or insightful. Once identity cues were added, the evaluation seemed to prioritize technical writing skills, a focus that likely contributed to the higher scores observed given the group’s general proficiency. This has important implications for adolescent writers, particularly those negotiating identity through literacy. Writing is not only an academic task but also a mode of self-making (Alim, 2011; Cushing, 2021, 2023). Feedback that reinterprets tone, structure, or clarity based on who the student is may cause them to question whether their voice belongs in formal writing spaces. Even when scores remain stable, the explanation can alter the meaning of the feedback. The same rhetorical choices (humor, informality, brevity) are perceived differently depending on identity. This reveals how AI systems may reinforce dominant rhetorical norms even when aiming for fairness. As models become more embedded in classrooms, their interpretive flexibility may become a hidden source of inequity. Rather than pursuing only numeric parity, fairness in writing assessment must entail interpretive consistency/respect for diverse rhetorical styles. If AI systems shift values based on demographic cues, they risk filtering student voices instead of fostering growth.
Significance/Implications
This study expands current fairness research by moving beyond score comparisons to analyze how model-generated explanations shift based on identity cues. The paired-condition design provides a method for detecting subtle interpretive bias in educational AI. Our findings show that LLMs shape more than outcomes; they influence how student voice is understood. When the same writing is evaluated differently depending on demographic information, these shifts carry real consequences for adolescent expression and rhetorical development. Equitable writing assessment requires more than balanced scores. It must include attention to how meaning is framed. Tools intended to support learning must not quietly sort expression through hidden assumptions. To support fairness in practice, we recommend auditing rationales for bias, making model training data and evaluation procedures transparent, involving educators in model review, and designing prompts that reflect diverse rhetorical traditions.