Paper Summary
Share...

Direct link:

Navigating the Challenges of Open-Ended Response Scores

Tue, April 12, 8:15 to 9:45am, Marriott Marquis, Floor: Level Four, Independence Salon G

Abstract

Open-ended response items are a staple of assessment design. Proponents of this format argue that they provide more valid means of eliciting students’ actual understanding, as it forces students to articulate their own ideas rather than match their thinking to a pre-determined response (Linn, Baker & Dunbar, 1991). Open-ended responses also allow for a range of knowledge and understanding to be elicited in fewer items. This is particularly valuable in learning progression research, where the goal for assessment is to map this range of understandings with respect to the proposed LP. However, this format also has its know challenges specifically in terms of consistent and valid scoring, which become particularly challenging when considered in relation to LPs. As the goal is to measure to what extent or in what ways students understand a concept, scoring open-ended responses must be done with respect to an outcome space (Wilson, 2004) that provides guidance in assigning value to specific kinds of answers that map to particular levels of the construct.
Our project explored the progression of student reasoning of large-scale Earth science phenomena with dynamic models. We developed three aligned assessments on different topics, administered to 357 8th/9th grade students in a pre/post-instruction format. The majority of items in each assessment were in short constructed-response format. Although we eventually achieved high inter-rater reliability and consistency across assessments (kappa=0.84-0.90), we faced numerous challenges in mapping between actual student responses and the guiding outcome space, particularly how to walk the line between giving students as much benefit of the doubt as possible, while still scoring in a way that was reliable and consistent with our proposed LP construct. The challenges tended to fall into three general areas: a) item prompts eliciting greater, or lesser, range of student thinking than anticipated; b) how much to interpret students’ inexact use of science vocabulary; and c) consistency in coding across variations both within and between items.
For example, we designed an item where students observed a dynamic model of moon phases, and were asked to list three things about the model that were different from the real Earth-sun-moon system. Our goal for this item was find out what aspects students saw as problematic in the model for helping them understand the causes for moon phases. However, when coding responses, numerous issues came up, such as: did students list only surface characteristics (color, materials), because that is what they understood or were those the easiest answers to get it “right”? When students used “rotate” and “revolve” inaccurately, should they get credit or not? If three characteristics were very different, should they each be counted or only the most sophisticated response on our proposed construct? What to do with the six responses (out of over 1000) that were coded at the highest level of the construct but therefor were not statistically valid? In the presentation, more detailed examples and elaborations of these challenges and our approaches to them will be shared and discussed.

Author