Paper Summary
Share...

Direct link:

Integrating Human and Automated Essay Scoring in a Scenario-Based Formative and Summative Assessment

Fri, April 4, 12:25 to 1:55pm, Convention Center, Floor: 100 Level, 121C

Abstract

Research and development in automated essay scoring has made machine scoring a realistic solution in providing fast score turnaround and reducing the expense and effort in human scoring. Machine scoring might be best viewed as complementary to human scoring in that the former can offer advantages that the latter can’t, and vice versa. For example, most automated scoring systems (e.g., e-rater®) can offer instantaneous performance feedback and suggestions on a few writing dimensions (e.g., grammar, word choice, sentence structure), for which would be difficult for human graders to provide such instant feedback. On the other hand, human raters can evaluate higher-order thinking skills (e.g., the quality of the argumentation), but automated scoring had not been able to meaningfully measure cognitively-demanding skills.
Automated scoring has been considered as a critical capability in CBAL™ from the birth of the research program because it allows for scaling up the kinds of task design that are used frequently in CBAL™ and that require complex constructed responses. In this poster, we intend to highlight our research findings on the feasibility and challenges (from both theoretical and measurement points of view) to integrate human and automated scoring methods in such technology-enabled assessment as CBAL™.
The use of automated essay scoring in CBAL™ separates itself from the use in other operational testing programs in at least two aspects. One, the essay writing task is (the last) part of scenario-based writing assessments. Unlike the essay items in traditional standardized tests, when being presented the writing prompt, students (supposedly) have already gained certain level of understandings and formulated some thoughts on the issue at hand. Therefore, their writings tend to be rather content rich and topic specific. Two, when the test is used for formative or summative assessment purposes, as in CBAL™ assessments, offering feedback to the teachers and students becomes more important than for admission or placement purposes.
Both unique aspects as mentioned above have implications for automated essay scoring. For example, “writing from sources” as typical feature in scenario-based writing assessment usually leads students to “borrowing” from the source texts. Automated scoring system therefore should be responsive to this unique feature, and incorporate mechanisms that can detect the nature and extent of use from source texts and potential plagiarism, in order maintain the validity of the automated scores at optimal level. As for giving feedback, we have been able to develop an Essay Feedback Engine that can generate real-time feedback while an essay is still being written, which is available per writer’s request. The feedback engine uses both process and product measures that are generated from the automated scoring system, and is also designed to allow for configurable comments appropriate for a variety of teaching and learning conditions.
In this poster presentation, we plan to share the lessons we have learned so far in incorporating automated essay scoring system in a technology-enabled testing program, as well as offer guidance on how to consider and design machine scoring that is most appropriate for scenario-based formative and summative assessments.

Authors