Paper Summary
Share...

Direct link:

Using Automated Analysis to Assess Middle School Students' Competence With Scientific Argumentation

Sat, April 18, 10:35am to 12:05pm, Virtual Room

Abstract

Objectives
This paper describes the use of computerized tools to a) identify concepts in written responses to argumentation assessment items, b) allow iterative development of emergent coding schemes, and c) create statistical scoring models of text. The goal is to develop accurate and reliable models that can score written responses at levels equal to human expert scorers, and to accurately place students on a learning progression for argumentation.
Theoretical Framework
Argumentation is fundamental to both science and science education. This perspective is reflected in international reforms, as well as the Next Generation Science Standards (NGSS; Lead States, 2013) where argumentation is presented as one of eight science and engineering practices. However, these new standards will only have a meaningful impact if they are accompanied by high quality assessments. Such assessments demand a move away from the efficiency of multiple-choice items, to the use of more authentic tasks. These tasks involve significant student written work, which is expensive to score. Therefore, our research question is: Can we develop automated computer scoring models of students’ argumentation responses that accurately predict expert human coding?
Method and Data Sources
Our work is based on the work of Osborne et al., 2013, who have developed assessments for argumentation and validated a learning progression. In the first phase of the project we used machine learning techniques to explore how computers could be trained to predict expert scores, and to learn how different item types lend themselves to this approach. We then engaged in new data collection (1000 5-6 grade students from the California Bay Area), human scoring, and machine classification, to gain greater sensitivity of the measure, and to use the scores to reliably place students along the learning progression.
Results
As an example of our work, we report findings from one item from one context: sugar dissolving in water, which targets level 1c (Constructing a complete argument) of the learning progression. Scoring for this item requires identification of all three elements of scientific argumentation: claim, warrant and evidence. Table 1 shows the levels of agreement for the scored components of this item.
Overall, we have good model performance across all analytic components. Each component had accuracy ranges above 0.9 and Cohen’s kappa above 0.7. Models for the Claim component and one of the Reasoning components are the lowest performers, although the models still have reliability measures greater than 0.7. We then combined these analytic component scores into a single, four-level, holistic score, then used these holistic scores to generate a predictive model. The holistic model for this item had an accuracy of 0.75 and a Cohen’s kappa of 0.562. The holistic model tends to under-predict students’ argumentation ability and has the most difficulty predicting a complete argument. This is likely due to the lower model performance on the Claim analytic component.
Significance
By developing computer scoring models that accurately replace the time-consuming process of expert human scoring, the resulting instruments can provide rapid scores for formative feedback and research purposes.

Authors