Paper Summary

Using the Many-Facet Rasch Model to Inform Standard-Setting Procedures: Setting performance standards for Advanced Placement examinations

Sun, April 15, 12:25 to 1:55pm, Marriott Pinnacle, Floor: Third Level, Pinnacle I

Abstract

Objectives

The purpose of this study is to utilize the Many-Facet Rasch (MFR) model to evaluate the quality of judgments provided by panelists who participated in a Modified Angoff standard setting for the 2011 Advanced Placement Environmental Science (APES) exam. Specifically, this application of the MFR model will focus on three facets: the panelist facet, the judged item facet, and the performance standard facet across two rounds of ratings.

Theoretical Framework

The conceptual framework presented by Engelhard (2009) for using the MFR model to evaluate the judgments from standard setting panelists will guide this research. This framework is grounded in the Rasch Measurement Theory criteria for evaluating scores on rater-mediated assessments (e.g., Engelhard, 2002). For the current study, these criteria are interpreted using standard setting as the context as opposed to the scores from rater-mediated assessments. Previous studies have employed this model to evaluate panelist ratings from the Michigan Educational Assessment Program (MEAP; Engelhard, in press).

Methods

The APES exam is comprised of 100 multiple choice (MC) items and 4 free response (FR) items. Examinees receive an overall AP score of 1, 2, 3, 4, or 5; thus, 4 cut-scores are set for AP exams. The data source for this study is the item-level ratings provided by the 15 panelists who participated in the Modified Angoff standard setting for APES. The MFR model will be applied to the ratings, and various fit indices and displays (e.g., variable maps) will be examined to answer the questions stated above.

Results/Conclusions

Table 1 in the uploaded document presents the demographic information for the 15 panelists. At this point, there are no results from the MFR analysis to report; however, all data have been collected, and analyses will be conducted and completed by the end of 2011, allowing for time to complete the paper before the AERA conference.

Scientific Significance

Standard setting procedures produce subjective judgments that inform the cut-scores for many large-scale assessments (e.g., AP, NAEP, MEAP; Engelhard, in press; Peterson et al., 2011). Given that these subjective judgments directly impact the decisions about students from high-stakes assessments, it is imperative to evaluate the quality of these judgments. The application of the MFR model demonstrated in this paper provides testing programs with an example of how to evaluate the quality of judgments from a standard setting so that they are equipped to use the results from standard setting procedures in a psychometrically appropriate manner. For example, these results can be used to provide descriptive information about the standard setting results that can inform the acceptable ranges of cut scores. Also, these results can be used to identify panelists whose ratings should be considered for exclusion in final cut score computations. Finally, a variable map presents a holistic illustration of how all facets are related to one another and can be used to inform cut scores.

Authors