Paper Summary
Share...

Direct link:

Calibrating an Automated Rater for Use With a Validated Observational Measure of Teacher Questioning (Poster 4)

Thu, April 21, 8:00 to 9:30am PDT (8:00 to 9:30am PDT), Marriott Marquis San Diego Marina, Floor: North Tower, Ground Level, Pacific Ballroom 18

Abstract

The quality of teacher questions is critical for fostering rich and rigorous mathematical discourse (Ball, 1993; Michaels et al., 2008) and supporting more equitable access to mathematical learning (Aguirre et al., 2013; Boaler & Staples, 2008). Within ambitious mathematics teaching contexts, validated instruments have been used to provide teachers with feedback that has helped them improve the quality of the questions they ask (Boston & Candela, 2018). However, the use of these instruments is time-intensive and difficult to scale. This study examines the extent to which a tool developed to automatically classify the quality of teacher questions can be used to provide validated and more accessible feedback opportunities for teachers and their coaches.
Providing feedback to teachers using validated observational measures of their teaching is an empirically grounded practice for improving the quality of classroom teaching (Pianta & Hamre, 2009). One implementation challenge is ensuring reliability across multiple raters, but with training and calibration the consistency of rater scores can be improved (Cash, Hamre, Pianta, & Meyers, 2012). Automated raters offer the potential for consistent scoring at scale, but algorithmic biases mean that calibration is needed when converting their observations into feedback for improving classroom teaching.
This study leverages an artificially intelligent teacher question classifier as an automated rater. It categorizes teacher utterances as either: probing or exploring questions, factual or recall questions, expository or cueing statements, or other. The classifier was trained using a database of teacher utterances recorded during ambitious mathematics teaching in authentic classrooms. The classifier performs well (F1=0.71) and its performance is comparable to similar classifiers of teacher utterances (e.g., Suresh et al., 2019).
This study compared the scores of two human raters and an automated rater. The raters scored 16 transcripts of authentic classroom teaching, each covering approximately 15 minutes of classroom discourse between high school mathematics teachers and small groups of students (Bywater et al., 2019). Scores followed the teacher questioning component of the Instructional Quality Assessment (IQA; Boston & Candela, 2018) that rewards the number of probing questions teachers ask. The human raters independently read each transcript and assigned a score. The automated rater categorized each teacher utterance and reported the number and percentage of probing questions. These were converted to an IQA score, first, by following the IQA rubric rules for the number of questions asked, and second, by following a calibrated version of the IQA rubric rules that relied on the percentage of questions asked.
Results found agreement between the two human raters was “moderate” (K=0.50). The agreement between the two human raters and the automated rater was only “fair” (K=0.22) when using the number of probing questions but returned to “moderate” (K=0.55) when using a calibrated percentage.
This study highlights the importance of calibrating automated raters with human raters. Just as training and calibrating human raters is important when implementing validated instruments, both training and calibration of automated raters is important. While much research has focused on training automated classifiers, calibrating them may also be critical.

Authors