Search
Browse By Day
Browse By Time
Browse By Person
Browse By Committee or SIG
Browse By Session Type
Browse By Keywords
Browse By Geographic Descriptor
Search Tips
Personal Schedule
Change Preferences / Time Zone
Sign In
Introduction
The integration of Artificial Intelligence (AI) into educational technology (EdTech) has accelerated the development of AI tutors for personalized learning (Jurenka et al., 2024). However, a critical limitation persists: many systems provide answers or solution steps directly, bypassing the student’s critical thinking process. Increasing evidence suggests that while AI-assisted learning offers several benefits, it can inadvertently hinder the development of higher-order thinking skills (Zhai et al., 2024). To mitigate this, carefully designed, human-integrated safeguards can be employed to eliminate the negative effects (Bastani et al., 2025). Research indicates that AI interventions are highly effective when they actively engage learners, deliver constructive feedback, and encourage reasoning rather than simply supplying solutions (Vieriu & Petrea, 2023).
Current AI-based assessment tools fall short in several key areas, limiting their efficacy and equity:
1. Limited Pedagogical Depth: Many models perform superficial content extraction, generating trivial questions that lack pedagogical depth and cultural context (Liao & Koh, 2020). This fails to promote critical thinking.
2. Lack of Meaningful Localization: Built for global scalability, AI tools often embed cultural biases favoring western perspectives, marginalizing local languages and contexts (Baker, 2025). They struggle to adapt to hyper-localized curricula and lack the syntactic and semantic nuance for languages in the Global South.
3. Superficial Standards Alignment: Alignment with curricular standards (e.g., CBSE, ICSE) is often reduced to simplistic tagging rather than ensuring genuine integration with intended learning objectives and cognitive levels.
Addressing these gaps requires a fundamental shift toward an ethical, human-centered design. This involves supporting constructivist pedagogies that promote active engagement and knowledge construction, and maintaining a human-in-the-loop model where AI augments, rather than replaces, teacher capabilities (Baker, 2025).
The Anytime Testing Model (ATM): A Participatory Framework
In direct response to these gaps, the Anytime Testing Model (ATM) was developed as an AI-based system to generate pedagogically sound assessments, grade responses, and provide personalized feedback. The core objective of ATM is to accurately diagnose a learner's knowledge state—identifying precise gaps in knowledge, comprehension, and understanding—to inform effective individual learning journeys, while supporting teachers with scalable, accurate evaluation tools.
ATM’s development followed a participatory, iterative design methodology involving in-service teachers, subject matter experts, curriculum designers, field implementation teams, and technology developers. This collaborative process informed rubric development, question framing, grading guidelines, and feedback strategies, ensuring that the system is deeply embedded in educational practice. The system underwent two controlled pilot studies with over 100 learners and facilitators across 2 Indian states (Rajasthan and Madhya Pradesh) for empirical validation.
Key Features of the ATM System
1. Intelligent Question Generation
ATM’s question generation is guided by a hybrid rubric that synthesizes global cognitive frameworks (Bloom’s Taxonomy) with Indian educational standards (NIOS, NCERT, CBSE). Its development involved iterative consultations with SMEs to ensure pedagogical relevance, appropriate cognitive weightage, and contextual alignment for regional and under-resourced settings.
Pilot testing revealed the critical need for linguistic simplification. For example, the question “What was the primary reason for Kondiba moving to Mumbai?” was refined to “Why did Kondiba move to Mumbai?” based on teacher feedback. Such insights directly informed the integration of psycholinguistic parameters—sentence length, vocabulary difficulty, and syntactic complexity—into the algorithm.
The current system generates a spectrum of questions from factual recall to complex analysis, each tagged by difficulty level. It allows educators to dynamically create assessments tailored to specific examination boards (e.g., State Board, NIOS), question types (MCQ, short answer, long answer), and cognitive balance. Trained on past exam papers and blueprints, the model accommodates regional language nuances, ensuring generated questions are linguistically accurate, contextually meaningful, and pedagogically sound.
2. AI-Powered Grading with Human Oversight
ATM employs AI to grade student responses using rubrics co-designed with content experts, field teams and educators. The system emphasizes answer completeness and conceptual relevance while adopting a deliberately lenient approach to spelling and grammatical inaccuracies.
In a pilot comparison of AI-generated and teacher-assigned grades for 30 students, the system initially underscored responses due to overly strict expectations of clarity and articulation. Educator analysis of these discrepancies led to crucial adjustments, retraining the AI to prioritize the presence of essential keywords and ideas over linguistic formality. This iterative, human-centered calibration process ensures the system’s alignment with pedagogical values of fairness and contextual sensitivity, effectively supporting teacher judgment rather than replacing it.
3. Personalized and Actionable Feedback
A cornerstone of ATM is its provision of constructive, student-friendly feedback, highlighting strengths, identifying errors or omissions, and offering actionable suggestions to promote critical thinking rather than merely revealing answers. Feedback is delivered in the student’s mother tongue, using English only when necessary, and structured into three concise sections: what worked, what could be improved, and actionable next steps.
Early iterations (~150 words) were too lengthy, prompting refinement to 60–80 words for clarity and readability. Feedback is sensitive to learners from under-resourced communities, supporting mix-language guidance that is clear, positive, and motivational. A human-in-the-loop element ensures teachers review and approve both grades and feedback before sharing with students, maintaining pedagogical alignment, contextual relevance, and reliability.
Challenges and Conclusion
Implementing ATM has highlighted several technical and operational challenges. On the technical side, achieving high OCR accuracy for regional languages remains a significant hurdle, affecting the system’s ability to accurately process handwritten responses. Operationally, implementation is constrained by factors such as blurred or poorly captured images, limited digital infrastructure, and low technological proficiency among students and teachers, which can impede smooth deployment.
Despite these challenges, ATM represents a significant advancement in AI-driven educational assessment. By combining global pedagogical frameworks with deep local contextualization through a participatory design, it demonstrates a viable model for generating equitable, high-quality assessments. Early pilots with 100 learners and ongoing testing with 4,000 students indicate positive reception from both educators and learners. By rigorously adhering to a human-in-the-loop, ethically designed approach, ATM aims to make quality, personalized learning support universally accessible, even in the most under-resourced and multilingual classrooms.