Paper Summary
Share...

Direct link:

Developing Measurement Models for Complex Scenario-Based Assessment Tasks

Fri, April 4, 8:15 to 9:45am, Convention Center, Floor: 100 Level, 111A

Abstract

Purpose
This paper discusses how the use of an Evidence-Centered Design (ECD) (Mislevy & Riconscente, 2006) approach for assessment development can aid in the identification of appropriate measurement models for complexly-structured performance tasks. In an ECD approach the connections between the student model (what it is we want to say about the student), the evidence model (which includes the measurement model), and the task model (features of the task that provide evidence) are made explicit. The development of these connections helps aid in the identification of appropriate measurement models.

Theoretical Perspective
Recent developments in assessment have seen a shift from traditional standardized testing to include the use of technology-enhanced scenario-based assessment tasks (Quellmalz & Pellegrino, 2009). These assessments present opportunities, in which students can engage in complex tasks such as designing investigations and manipulating representations of real-world tools. The use of these assessment tasks also provides challenges to the assessment designer when determining how the tasks should be scored. Traditional methods of scoring such as IRT are often not appropriate as the local independence assumption of items, and the unidimensionality assumption are violated by the presentation of an overarching scenario and multiple constructs being measured. In addition, advances in technology have led to new and innovative item types, and one task might present several different types of items which might not be scored the same. This paper presents some of the issues when it comes to the development of the measurement model for scoring these complex scenario-based assessment tasks.

Methods
This paper presents three measurement models that can be applied. One measurement model is an extension to IRT that allows for multiple dimensions to be measured, the MRCML model (Briggs & Wilson, 2003). Another method that will be presented is diagnostic classification models (DCMs) (Henson, Templin, & Willse, 2009). A final measurement model that will be discussed is Bayesian Networks (Almond, DiBello, Moulder, & Zapata-Rivera, 2007). All of these models allow for the incorporation of multiple dimensions into the measurement model and provide benefits when it comes to scoring. The paper will provide background on each of the models, and discuss benefits and drawbacks to each method based on the literature. It will also demonstrate the application of these models to an example of a scenario-based assessment task.

Results
While there is not one measurement model that would apply in all situations, the identification of the evidence needed to make valid inferences about the student, and the evidence that can be accumulated from the task can provide information that can be used to identify the appropriate measurement model for that task.

Significance
With more and more complex tasks being developed and used, the identification of a measurement model that can best leverage the information provided by the task will aid assessment designers in their use and development of these types of tasks.
 

Authors