Paper Summary

Some Drivers of Test Item Difficulty in Mathematics: An Item Analysis Rubric

Mon, April 16, 8:15 to 10:15am, Marriott Pinnacle, Floor: Third Level, Pinnacle II

Abstract

The purpose of this component of the symposium is to present a rubric used in a programme of ongoing research into some key drivers of the difficulty of mathematics test items. A set of mathematical competencies that are derived from the work of Niss (2003) is at the heart of this research, but similar competencies are defined and used in mathematics curriculum documents of many educational systems around the world, and they have been a central component of the PISA mathematics frameworks used since the inception of that survey (for example, see OECD, 1999). In addition, the competencies have heavily informed the descriptions of increasing levels of mathematical literacy that form a central plank of the OECD’s reporting of PISA outcomes (OECD, 2005).
This paper will focus firstly on the operational definition of six mathematical competencies used in this study (labelled as communication, devising a strategy, mathematisation, representation, using symbolic, formal and technical language and operations, and reasoning and argument) and on a set of descriptions of increasing levels of demand for activation of each of the competencies. The paper will also present a procedure that allows trained experts to evaluate mathematics survey items according to the extent to which solution of each item demands activation of each of the six mathematical competencies.
Previous analysis of expert ratings of mathematics test items developed for the PISA 2003 survey had shown that the item analysis rubric can be used to predict up to 70% of the variability in item difficulty for those items (Turner, Dossey, Blum and Niss, forthcoming). That analysis led to some refinement of the scheme for use in the present study.
The current research has focussed on test items developed for use in the PISA 2012 survey administration. A group of mathematics educators used the scheme to rate the items before any data on relative item difficulty were available.

Author