Search
Browse By Day
Browse By Time
Browse By Person
Browse By Committee or SIG
Browse By Session Type
Browse By Keywords
Browse By Geographic Descriptor
Search Tips
Personal Schedule
Change Preferences / Time Zone
Sign In
Quality Assurance in AI-EdTech is multi-faceted, with overlapping concerns and mixed terminologies. Part of this stems from the cross-disciplinary nature of work. We use the following terminology:
Standard: formally agreed requirements any product or process must meet (e.g., ISO 42001) – sets a common minimum bar across companies and jurisdictions.
Framework: structured map of domains to examine (pedagogy, data privacy, accessibility) – guides reviewers on where to look and how to score.
Benchmark: use-case-agnostic test probing general model capability (e.g., pedagogy tests) – enables neutral comparison across models.
Evaluation/Evals: task-specific assessment of performance on a concrete job (e.g., generate Grade-3 reading passage) – Tells decision-makers whether a tool works for their exact need.
Rapid-Cycle Evaluation: short, low-cost study (8-10 weeks) that still measures learning outcomes – this aligns evidence generation with the fast release cadence of AI models.
Having many standards is a feature, not a bug of global systems. Global standards are easiest to set up when there is a single task to be done with low diversity of interests and/or strong network effects, such as with shipping containers.
However, education has multiple objectives – learning, inclusion, teacher workload, contextual relevance; as well as diverse contexts with strong sovereign views. This fragmentation can be treated as a bug to fix, or a natural feature of different demands from stakeholders. As AI-Edtech is widely applicable across different use cases and geographies, it is likely to be a persistent feature.
In this context, in order to best support developers, funders, and governments may require a more practical middle-ground – with different layers of standards that could take the form of:
A “core” layer – such as data privacy, safety, ethics, shared agreed standards and benchmarks
An “application” layer – tailored to a specific application – such as tutoring specific benchmarks; benchmarks for teaching and learning; user experience.
Contextual/local layers – such as alignment with curriculum; specific interest groups (e.g. SEND) – added when relevant.
Where the overall balance results in sufficient agreement to embed the resulting standards and benchmarks as industry standards that can be widely adopted by developers, funders, and governments for procurement, investment, and policy decisions.
This involves both a top-down approach—engaging major funders, multilaterals, and national governments to promote uptake and alignment—and a bottom-up approach, focused on understanding what tooling, guidance, and support developers and implementers need to apply the standards in practice. We present initial insights and mapping from this process.
Overall, we think that having a range of components to develop different standards is acceptable; but what is most important is that those components are shared and accessible, enabling agreement on the most cross-cutting aspects. The presentations in this panel are great examples of that.