Individual Submission Summary
Share...

Direct link:

Digitization of early grade reading assessments in Tanzania using voice AI

Mon, March 30, 4:30 to 5:45pm, Hilton, Floor: Fourth Floor - Tower 3, Union Square 5&6

Proposal

Accurate, timely data on children’s reading skills is essential to tackling the global learning crisis. Yet, in low- and middle-income countries (LMICs), particularly in Sub-Saharan Africa and South Asia, collecting such data remains resource-intensive. Oral, one-on-one early grade reading assessments are an effective way for understanding foundational literacy, but their scale is constrained by costs, logistics, and the need for trained assessors. As a result, they are conducted infrequently, limiting opportunities for timely feedback and targeted instructional support.

Advances in voice Artificial Intelligence (AI), particularly Automatic Speech Recognition (ASR), offer a promising solution. By capturing and analysing children’s oral reading attempts in real time, ASR can automate parts of the assessment process, enabling more frequent feedback for teachers and systems. However, most existing ASR models are trained on adult speech in high-resource languages, optimized for conversations, meaning they smooth over mistakes to interpret what the reader was trying to say. In contrast, in the context of early grade reading assessments, the priority is to recognize exactly what a child says - often in noisy classroom conditions and without reliable internet access.

To meet this challenge, ASR models must be trained not only on the target language, but also on children’s voices, fine-tuned for assessment tasks, and optimised for low-resourced classrooms. Adapting voice AI in this way is essential to reliably and fairly evaluate foundational literacy across varied languages and classrooms, supporting inclusive education systems that can help close learning gaps.

This presentation reports on our recent efforts to advance voice AI for early literacy assessment by developing and fine-tuning ASR models using a newly collected Kiswahili child voice dataset from Tanzania. The work focuses on some of the most technically demanding and least addressed reading subtasks (letter sound identification, syllable reading, and nonword reading), where precise capture of children’s speech is critical. The dataset was used to fine-tune ASR models optimised for short, task-specific utterances and for operation in noisy classroom environments.

Navigating informed consent and data protection processes presents significant barriers that limit AI innovation in many LMICs and risk deepening educational inequalities. Working with local and international experts, we ensured compliance with national and international laws and ethical standards. Beyond the technical outcomes, the presentation shares lessons from navigating these ethical, legal, and operational challenges - insights that are critical to advancing inclusive, robust voice AI in education and to establishing responsible practices for its safe and equitable deployment.

Author