Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
The Next Generation Science Standards’ (NGSS) practices call for students to be able to argue using evidence-based reasoning in addition to demonstrating conceptual understandings (NGSS Lead States, 2013; NRC, 2012). Rubrics that capture the argumentative components of scientific explanations, CER, in the open response format have been developed based on a modified version of Toulmin’s (1958) model of argumentation (McNeill et al., 2006; Gotwals & Songer, 2010). The use of these rubrics, however, requires human scoring, which can be time consuming (Wainer & Thissen, 1993) and subject to error (Bejar, 2012; Myford & Wolfe, 2009). An effective and efficient solution is the use of automated scoring programs. These programs use natural language processing (NLP) techniques, such as regular expressions (RegEX), and machine learning algorithms to identify the presence of key concepts in students’ scientific explanations (Liu et al., 2014; Liu et al., 2016), but have not previously been able to capture students’ CER. Moreover, previous studies on automated assessment of essays (McNamara et al., 2015) and summaries (Li, Cai, & Graesser, 2016) showed that language features could better predict human scores. However, no studies have used language features to automatically score open response items.
The present study compared the performance of two automated scoring methods. The concept-based automated method used regular expressions that were generated based on key terms extracted from a rubric (see Table 1) used by human coders and an if-then algorithm (Li, Shubeck, & Graesser, 2016) to determine the presence of particular words or phrases. The language-form-based method adopted features frequently used in automated assessments including: word counts, narrative genre, word concreteness, syntactic simplicity, referential cohesion, and deep cohesion. This study is the first one to compare scientific concept and linguistic representation methods for scoring open responses. The findings will not only facilitate the development of automated techniques, but will also promote consideration of both concept and linguistic representations in writing scientific explanations.
293 middle school students first investigated how the shape of a container impacted the density of a liquid in a Density virtual lab in an intelligent tutoring system (ITS), Inq-ITS (Gobert et al., 2013), and then wrote a claim, evidence to support their claim, and reasoning for how their evidence supported their claim. Two expert human raters scored students responses using the fine-grained rubric that operationalized subskills of CER with extremely high interrater-reliabilities (Li et al., 2017a).
Pearson correlations between automated scores and human scores were .98, .98, and .95 for CER, respectively. Results of linear regressions indicated that the linguistic features could predict human scores with a moderate or good level, r = .70, .59, .61 for CER, respectively.
The findings of this study indicate that automated scoring using regular expressions and if-then algorithms based on the fine-grained CER rubric is an effective method for autoscoring students’ scientific explanations. This has great implications for providing automated, real-time scores and feedback. However, language form could still moderately predict human scores, even though the rubric did not include specific criteria for scientific writing.
Haiying Li, Rutgers University
Janice Gobert, Rutgers University
Rachel Fallon Dickler, Rutgers University