Paper Summary
Share...

Direct link:

Validating Assessments for a Genetics Learning Progression

Tue, April 12, 8:15 to 9:45am, Marriott Marquis, Floor: Level Four, Independence Salon G

Abstract

In our research we study a learning progression in genetics that spans grades 5-10 and includes multiple constructs with four levels each (Authors, 2014; Authors, 2009). Assessing a multitude of constructs and levels requires a large set of assessment items. Given the constraints of time and our study sample size we had limitations on the number of items we could administer and therefore decided to use ordered-multiple choice (OMC) items in order to obtain more information from each item. In OMC items different response options are linked to levels of conceptual understanding (Briggs, et al., 2006; Briggs & Alonzo, 2012); different items responses are given partial credit thus providing more information about students’ level of reasoning compared to multiple-choice items that are scores as correct or incorrect. We developed an item pool of over 80 OMC items with most of the items including response options that mapped onto 2-4 levels of a particular construct in the progression. Through an iterative process of pilot testing we ultimately selected a working pool of 31 items that we used in the final study. We conducted four pilot tests of varied samples both in size (from~100 to over 1000), student ability levels, and settings (school districts).
The polytomous item response data were analyzed using multidimensional random coefficient multinomial logit model (Adams, Wilson, & Wang, 1997). We first developed general item quality indices by combining information from classical item analyses (discrimination indices and point-biserial correlations) and item response modeling. We used these indices to identify problematic items that did not discriminate between levels of student ability as expected. We then examined student choices and the item parameters to characterize the exact nature of the problem in discrimination and implement revisions in terms of changing the item itself, its scoring, or the description of the level expectation (and evidence statements) in the progression itself.
In our presentation we will describe this process, and the revision strategies we employed through illustrative cases of problematic items. Importantly we wish to highlight the difficulty in operationalizing interpretations of statistical analyses in actual changes to items and scoring schemes. We offer a brief example of an item (humangrowth) designed to discriminate between four levels of students’ understanding of the construct in the progression that deals with the critical role of proteins in genetic phenomena. Higher levels of ability entail understanding that proteins are involved in all cell and body functions, whereas lower level of ability assume a more constrained set of protein functions. The item before and after revision is shown in Figure 1. The item quality measures are provided in Figure 2. Statistical analysis showed that the level 3 response option (A) was too easy. We hypothesized this was due to unbalanced response options in terms of amount of text and clarity of differences between the options. We revised the item to balance out the other options and clearly delineate which functions are not carried out by proteins. Subsequent pilot tests showed improvement in this item’s functioning.

Authors