Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Browse By Descriptor
Search Tips
Annual Meeting Theme
Exhibitors
About Philadelphia
About AERA
Personal Schedule
Sign In
X (Twitter)
Despite initial studies showing that scores from CVA assessments are reliable and predict instructional quality and student learning, wider use of the assessments has been impractical because teachers’ written responses have to be scored by trained human raters (Kersting et al., 2010; Kersting et al., 2012), which is a time and cost intensive process. To improve scoring of the CVA scales we report in this paper on our efforts to develop computer algorithms that might effectively reproduce scores assigned by trained human raters.
Considerable progress has been made in recent years in the application of automated text analysis techniques for a wide range of text classification problems. Depending on the kind of text and suitability of the rubrics, agreements between human and computer generated scores have ranged from .70 to .93, indicating that in some cases agreement between computer and human generated scores is higher than agreement between any pair of raters (Streeter, Bernstein, Foltz, & DeLand, 2011).
Using human-scored responses from three different CVA assessments (on fractions, on ratio and proportions, and on variables, expressions and equations), we developed automated classifiers, specifically a kind of Naïve Bayes classifier, to computer-score teacher responses and compared the scores produced by these classifiers to those assigned by human raters. To be able to test the performance of our scoring algorithms, we only used a subset of human scored responses in the algorithm development and used the remaining responses for testing. We then used cross-validation techniques in which we randomly varied the set of manually scored responses used in the algorithm development phase varied 50 times to explore the generalizability of our algorithms’ performance. Finally, we explored the validity of the computer-generated scores by relating them to scores from another measure of teacher knowledge, the Mathematics Knowledge for Teaching (MKT) instrument and compared the strength of association produced by human-generated and computer-generated scores.
Our study produced interesting results. First, we found that total scores and subscores, based on computer-generated scores correlated highly with total and subscores computed using human assigned scores. Across all three CVA scales correlations exceeded .85 for total scores and around .80 for three of the four rubric subscores, suggesting that teacher rank-ordering was fairly stable regardless of the scoring method and that computer-generated total scores might be considered reliable.
We also that found that average percent agreement and kappas across clips and across the 50 separate runs were lower, ranging from .49 to .75 for percent agreement depending on clip, while kappa values varied from .40 and .56.
Finally, we found that the strength of correlations between teachers scores on the MKT and human-generated and computer-generated scores were very close, providing some evidence for the validity of computer-generated scores.
Overall, our results show that scoring of teacher responses on CVA scales can be automated with a reasonable degree of accuracy for some video clips. A closer examination of clips for which automated scoring worked less well revealed that clip content and strength of the clip stimulus affected algorithm performance.
Nicole B. Kersting, The University of Arizona
Bruce Sherin, Northwestern University
James W. Stigler, University of California - Los Angeles