Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Help
About Vancouver
Personal Schedule
Sign In
Results on reading tests are typically reported on scales composed of levels, each giving a statement of student achievement or proficiency. In PISA, the largest international survey of reading literacy, scale descriptions at each level are derived from item difficulty, following analysis of international main survey data, from 67 countries and 46 languages. The PISA reading framework describes features of reading that drive difficulty, derived from empirical and theoretical work conducted by Kirsch and colleagues (e.g. Kirsch, et al., 2002; Kirsch & Mosenthal, 1990), but it has been unclear how useful these features are in accurately predicting the difficulty of individual reading items in PISA. The PISA reading scales provide broad descriptions of skill levels associated with reading items, intended to communicate to policy makers and teachers about the reading proficiency of students at different levels. However, the described scales are not explicitly tied to features that predict difficulty. Difficulty is thus treated as an empirical issue, using a post hoc solution, while a priori estimates of item difficulty have tended to be unreliable. Understanding features influencing the difficulty of reading tasks has the potential to help test developers, teachers and researchers interested in understanding the construct of reading.
Work directed towards this goal for the PISA 2000 cycle derived features of reading presumed to drive difficulty, from PISA items, related to the framework variables of aspect (reading process) and text format. The work was extended during the PISA 2009 cycle, providing a single set of variables, for a set of 84 PISA items. Further refinement, using both content analysis and data analysis, led to a set of ten variables, each described at four steps of difficulty. Negotiation of a common interpretation of the descriptors for each step allowed a consensus to be reached by a group of three expert raters on the ratings for each of the ten variables, for the same items.
Correlations between the consensus ratings and item difficulty were calculated. A regression analysis identified that three of the variables with the highest correlations (all above 0.55) and one additional variable together predicted 56% of the variance in difficulty for these items. A principal components analysis showed that a higher level of variance still could be explained with factors relating to the aspect or skill involved, features of the text, features of the task, and the task format (restricted choice versus constructed response of various kinds).
An extension of the study is currently underway, involving independent rating of PISA items by raters trained to use the rating scheme developed here. The results of this work will be included in the presentation.
The process described in this study has shown that a theoretically motivated scheme, drawing on multiple perspectives, and refined in the light of empirical data, has the potential to advance significantly our understanding of the features that contribute to the difficulty of diverse reading comprehension test items. It is likely to have applicability in contexts well beyond PISA.