Paper Summary
Share...

Direct link:

Evaluation of Automated Scoring Methods for Students' Claim, Evidence, and Reasoning Responses in Science (Poster 5)

Sat, April 15, 2:50 to 4:20pm CDT (2:50 to 4:20pm CDT), Radisson Blu Aqua Hotel, Chicago, Floor: 1st Floor, Atlantic E

Abstract

Introduction
The Next Generation Science Standards (NGSS Lead States, 2013; NRC, 2012) emphasize constructing explanations as a key science inquiry practice; however, it can be challenging and time-consuming for teachers to develop scoring rubrics and grade students’ written science inquiry explanations. An Intelligent Tutoring System [ITS, name blinded] (Authors, 2013) addresses this challenge by using natural language processing (NLP) techniques to automatically assess students’ science explanations written in a Claim, Evidence, Reasoning (CER) format (Authors, 2017a,b; McNeill et al., 2006), with machine-learned and knowledge-engineered algorithms for other inquiry competencies. In this study, we expand on prior work evaluating the auto-scoring of science explanations (Li et al, 2017a,b) by evaluating two different auto-scoring methods for the future implementation of automated scoring method and the real-time feedback and scaffolding in the ITS [name blinded].

Methods
We first obtained students’ CER responses in the ITS [name blinded] for four activity topics (Table 8) and then selected a random sample of 100 response sets from each topic to manually hand-score.
A trained human rater completed the hand-scoring using rubrics (Table 9) adapted from previously developed rubrics for CER science explanations (c.f., Authors, 2017a,b). However, with respect to the algorithms used to automatically score students’ responses, there were two different methods for implementing the scoring rules: a RegEx method and a WordDistance method.
The RegEx method uses regular expressions (Thompson, 1968) to match patterns of words associated with the scoring component (example). This method has been described and used in prior ITS [name blinded] work (Authors, 2017a,b). The WordDistance method is an alternative method in which the terms associated with the key concepts in an activity are linked to an identifier in a dictionary, with all similar terms (e.g., “affects”, “impacts”) using the same identifier. These key terms can then be combined with others to form larger concepts, such as the IV relationship (e.g., “the mass of the green ball increases”). Components are then scored based on the presence of those terms, or concepts, occurring within a specified distance from each other. For both methods, the developer creating the automated scoring used both the rubric and previously hand-scored student data as a reference when developing the scoring patterns.

Results and Discussion
The average inter-rater agreement, as measured by Cohen’s quadratic weighted kappa (Cohen, 1968), was 88.80% for written claims, 85.95% for written evidence, 85.80% for written reasoning between RegEx method scores and human scores, and 85.90% for written claims, 89.10% for written evidence, 88.50% for written reasoning between WordDistance method scores and human scores (Table 10). These results indicate that the algorithms are performing moderately well in scoring students’ CER responses. However, further analyses revealed that, for three activities, there was low agreement between the hand-scoring and auto-scoring for the “Theory” sub-component of the Reasoning portion for both algorithms. Future work will explore how to operationalize students’ written scientific theories across the different science domains to improve automated scoring in this area.

Authors