Paper Summary
Share...

Direct link:

Developing Assessments for College-Ready Critical Reasoning in Statistics

Fri, April 17, 2:15 to 3:45pm, Virtual Room

Abstract

Objectives and framework(s)
This presentation will describe the design of the Critical Reasoning for College Readiness (CR4CR) Assessment using the BEAR Assessment System (BAS). Specifically, it will focus on the use of Wright Maps, exploratory graphical devices that plot item and student locations along a construct, as validity evidence. We had two main priorities for the empirical validation work using our field test data: (1) validate the alignment of item responses to construct levels and (2) compare the empirical behavior of selected and constructed response item formats. Wright Maps can speak to both priorities.
Drawing on a literature review on statistical curricula, we developed an instrument to measure three constructs—Linking Data to a Claim (LDC), Meta-Representational Competence (MRC), and Formal Inference (FoI)—each defined as a collection of ordered and qualitatively distinct levels of sophistication. We used the BEAR Assessment System (BAS), a principled approach to instrument design to develop our constructs and test items (Author_D, 2005). In this approach, construct development precedes item development and all item responses are aligned to a level of the construct. The aligned construct level is what gives the numerical scores assigned to item responses meaning.
Methods and data sources
Response data from a Fall 2018 administration of the CR4CR Assessment to 248 high school and early college students was analyzed using the Partial Credit Model (Masters, 1982). Estimates of item step difficulties were nonlinearly transformed to step thresholds as they have a more straightforward interpretation than step difficulties. These step thresholds are then plotted along the vertical axis of a Wright Map (see Figure 3) and are organized horizontally by construct level. The distribution of student ability is also shown in the display, here, as an on-its-side histogram. From these displays, we can investigate the comparative relationships among response categories and students. Specifically, for the comparison of selected response items to their open-ended counterparts, we can see whether they have similar difficulty within each construct level, as we intended, or if the selected response items were easier than constructed response, as existing literature may suggest.
Conclusions
The banding apparent in our Wright Maps for FoI and LDC serve as supporting evidence that the items on the CR4CR Assessment can distinguish and sort respondents into the levels of our constructs, as we defined them. This means that student scores can be interpreted meaningfully in terms of our construct levels, providing teachers with valuable diagnostic or formative information. Results for the MRC construct, however, show that more work needs to be done, especially at the lower levels of the construct, to write items that distinguish among those levels.
Further, with strategic form design and assignment, we were able to compare the empirical behavior of the selected and constructed response item counterparts. We found that though existing literature suggests that selected response item types are easier than constructed response item types, this may not be the case for all constructs. We found that, for the LDC and FoI constructs, in many cases multi-select multiple choice items behaved similarly to their constructed response counterparts.

Authors