Paper Summary
Share...

Direct link:

Machine Learning Analysis of Student Responses to Carbon: Transformations in Matter and Energy (Carbon TIME) Learning Progression Items

Tue, April 12, 8:15 to 9:45am, Marriott Marquis, Floor: Level Four, Independence Salon G

Abstract

Machine learning (ML) based evaluation of written responses has the potential to open a wider variety of item types for learning progression (LP) assessment at a larger scale. Large-scale LP assessments often rely on forced-choice (FC) formats due to ease of scoring. While constructed-response (CR) formats are highly desirable, they are very expensive to employ on a large scale if humans must evaluate each and every response. ML algorithms use patterns in student responses to predict human scores or classifications, and can flexibly combine data from multiple FC and CR questions to predict the LP diagnosis a human expert would make. Furthermore, because these methods make the scoring process explicit, they can pinpoint the specific reason why a student received a particular classification. In this presentation we discuss our experience using machine learning to evaluate written responses to Carbon TIME items.
During the first wave of Carbon Time scoring we developed machine scoring models for a total of 18 items. Most of these items collected a mixture of FC and CR responses that human raters evaluated as a package to arrive at an LP classification. Most items had between 1,000 and 2,000 responses. After being spell-corrected, responses were broken down into component words, phrases, parts of speech, and other features. Features with no relationship to a particular LP level were filtered out. Models were fit to the remaining features using two different ML methods: logistic regression and decision trees. Model accuracy was evaluated by comparing computer and human classifications using quadratic weighted kappa (QWK; Cohen, 1968; Sim & Wright, 2005). Values above .70 indicate satisfactory agreement (Williamson, Xi, & Breyer, 2012). Models for 14 of the initial 18 items analyzed (78%) achieved a QWK above .70, and 9 achieved a QWK above .80. Of the 4 items with scoring problems, 3 were caused by insufficient (less than 50) responses at Level 4 of the Carbon Time learning progression.
In addition to evaluating the LP level of an item response, we also investigated whether the ML models could reliably distinguish different subcategories of responses within a level. In the process of human coding, raters also coded for the specific type of reasoning displayed by the student. This amounted to a sub-classification that formed the basis for the LP level diagnosis. These within-level distinctions are not intended to affect a student’s overall score in, but they can be used to provide additional feedback about the nature of reasoning used in addition to the LP level description. As a somewhat surprising side benefit, we found that fitting ML models to within-level subcategories actually improved overall scoring accuracy. For example, the one item with scoring difficulties that could not be explained by inadequate sample size initially achieved a QWK of .68 when classifying responses based on overall LP level. This value improved to .77 when we began by obtaining sub-classifications and then aggregated those into LP levels.

Authors