Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
Objectives
The purpose of this study is to examine the quality of ratings on free-response questions in large-scale assessments. Specifically, this study will focus on indices of rater errors and systematic biases in large-scale rater-mediated assessments, as well as their implications for the validity of the assessment system. This study is guided by the following research questions:
Do the raters vary in severity?
Is the structure of the rating scale comparable across the free-response questions?
Does the Many-Facet Rasch Model provide a useful framework for analyzing the AP Statistics data?
Theoretical Framework
Item Response Theory is the general theoretical framework for this study. Specifically, the Many-Facet Rasch Model is used to guide our analyses and interpretation of these data. The guidelines recommended by Linacre (2010) are used to explore the psychometric quality of the rating scales for each of the six essays.
Methods
This study focuses on the 2009 administration of the AP Statistics Examination (N ≈ 350 students, N ≈ 90 raters). The AP Statistics exam consists of 40 multiple-choice questions and six free-response questions with four rating categories. A Many-Facet Rasch Model with four facets (students, modes (essays & multiple-choice items), items and raters) is used to model the data. A partial-credit model is used to model the multiple-choice items and the six free-response questions.
Results/Conclusions
Preliminary analyses have been conducted. Table 1 (please see the uploaded summary document) presents a summary of findings from the Facets analyses for students, modes, items and raters, indicating high reliability of separation for each variable of interest. Raters vary significantly in severity. The overall model-data fit appears to be quite good. The final analyses will present a detailed look at the rating scale structure for each free-response question.
Scientific Significance
The College Board currently examines rating quality for free-response questions with Reader Reliability Studies based on Generalizability Theory (Miao & Odumade, 2011). This study augments the current research on reader reliability with indicators of rating quality based on Rasch measurement theory.