Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Personal Schedule
Sign In
X (Twitter)
Objectives
This paper describes the use of computerized tools to a) identify concepts in written responses to argumentation assessment items, b) allow iterative development of emergent coding schemes, and c) create statistical scoring models of text. The goal is to develop accurate and reliable models that can score written responses at levels equal to human expert scorers, and to accurately place students on a learning progression for argumentation.
Theoretical Framework
Argumentation is fundamental to both science and science education. This perspective is reflected in international reforms, as well as the Next Generation Science Standards (NGSS; Lead States, 2013) where argumentation is presented as one of eight science and engineering practices. However, these new standards will only have a meaningful impact if they are accompanied by high quality assessments. Such assessments demand a move away from the efficiency of multiple-choice items, to the use of more authentic tasks. These tasks involve significant student written work, which is expensive to score. Therefore, our research question is: Can we develop automated computer scoring models of students’ argumentation responses that accurately predict expert human coding?
Method and Data Sources
Our work is based on the work of Osborne et al., 2013, who have developed assessments for argumentation and validated a learning progression. In the first phase of the project we used machine learning techniques to explore how computers could be trained to predict expert scores, and to learn how different item types lend themselves to this approach. We then engaged in new data collection (1000 5-6 grade students from the California Bay Area), human scoring, and machine classification, to gain greater sensitivity of the measure, and to use the scores to reliably place students along the learning progression.
Results
As an example of our work, we report findings from one item from one context: sugar dissolving in water, which targets level 1c (Constructing a complete argument) of the learning progression. Scoring for this item requires identification of all three elements of scientific argumentation: claim, warrant and evidence. Table 1 shows the levels of agreement for the scored components of this item.
Overall, we have good model performance across all analytic components. Each component had accuracy ranges above 0.9 and Cohen’s kappa above 0.7. Models for the Claim component and one of the Reasoning components are the lowest performers, although the models still have reliability measures greater than 0.7. We then combined these analytic component scores into a single, four-level, holistic score, then used these holistic scores to generate a predictive model. The holistic model for this item had an accuracy of 0.75 and a Cohen’s kappa of 0.562. The holistic model tends to under-predict students’ argumentation ability and has the most difficulty predicting a complete argument. This is likely due to the lower model performance on the Claim analytic component.
Significance
By developing computer scoring models that accurately replace the time-consuming process of expert human scoring, the resulting instruments can provide rapid scores for formative feedback and research purposes.
Christopher D. Wilson, Biological Sciences Curriculum Study
Molly A.M. Stuhlsatz, BSCS Science Learning
Brian Matthew Donovan, BSCS Science Learning
Zoe Bracey, BSCS Science Learning
April Lynn Gardner, Biological Sciences Curriculum Study
Jonathan F. Osborne, Stanford University
Tina Cheuk, California Polytechnic State University - San Luis Obispo
Kevin Haudek, Michigan State University
Marisol Mercado Santiago