Paper Summary
Share...

Direct link:

Addressing Practical Challenges in the International Performance Assessment of Critical Thinking

Sun, April 24, 4:15 to 5:45pm PDT (4:15 to 5:45pm PDT), Division Virtual Rooms, Division D - Section 1: Educational Measurement, Psychometrics, and Assessment Virtual Paper Session Room

Abstract

1. Objectives or Purpose
An assessment that claims to measure critical thinking (CT) must include one or more performance tasks (PT) to elicit evidence with respect to the CT facets that are not measurable using simpler task types (e.g., multiple choice) (Davey et al., 2015). PTs are extended, scenario-based constructed response items. Two major challenges face the use of extended PT in instructional and/or (standardized) assessment contexts: lengthy time requirements (~90 min) and concerns about reliability with only one PT (Davey et al., 2015). Partitioning an extended PT into “mini-clones” or segments of the original extended PT might provide one possible solution. In this presentation, we explore PT design and construction considerations within the iPAL Assessment Framework. We conclude with a discussion of the validation strategies to be followed in the coming years.
2. Theoretical Framework
The iPAL framework (Braun et al., 2020) for designing PT comprises four main aspects: (1) The storyline describes a carefully curated version of a complex, real-world situation. (2) The challenge frames the task to be accomplished, with reference to the documents provided and with varying degrees of scaffolding. (3) A portfolio of documents in a range of formats (e.g., reports, charts, blogs, twitter threads) is drawn from multiple sources chosen to reflect different levels of relevance, trustworthiness, and susceptibility to bias. (4) The scoring rubric comprises a set of scales, each linked to a facet of the CT construct.
Following principles of evidence-centered design (Mislevy & Haertel, 2006) and argument-based approaches to validity (Kane, 2013), interpretability of responses depends on strict coherence between the framework, the design of the mini-clones, and the corresponding scoring rubrics. Efficiency requires that each PT give rise to several mini-clones with overlapping document sets.
3 Methods
We employ a systematic design approach to mini-clones that isolates a subset of CT facets (e.g., evaluating the relevance and trustworthiness of information), then modifies the challenge and document set to be consistent with the measurement of that subset, along with appropriate changes to the scoring rubric. Validation will be carried out through expert reviews and cognitive laboratories.
4 Data Sources
The primary data sources are the PT materials and the extant mini-clones that have been constructed in our study. Additional validity evidence from expert reviews and cognitive laboratories will be collected.
5 Results
This study evaluates the feasibility of constructing sets of mini-clones from PT assessing CT in a range of contexts. Preliminary data indicates that time requirements are substantially reduced while maintaining construct validity.
6 Scholarly Significance
Multifaceted skills such as CT pose challenges in balancing competing psychometric and practical considerations. This “mini-clone” study examines one possible way to achieve this balance, while also addressing concerns about the use of performance assessment for teaching purposes (developing student skills) and accountability purposes (increasing reliability, validity, and utility of PT). This work sets the stage for further validation studies.

Authors