Paper Summary
Share...

Direct link:

Models for Examining Gender Bias in the Marking of Writing Assessments

Tue, April 17, 10:35am to 12:05pm, Vancouver Convention Centre, Floor: Second Level, East Room 13

Abstract

Do teachers mark writing assignments differently for boys and girls? Is there a natural tendency to look at writing from boys and girls through different lenses? In particular, might it be something about the nature of the writing itself that leads to these differences, as opposed to simply knowing that the writer is a boy or a girl?

According to Engelhard (in press), the measurement of persons must be independent of the particular raters that happen to be used for the measuring as well as characteristics of the individuals being rated. The degree to which a particular assessment instrument meets or fails to meet these requirements can be used to determine the instrument’s level of invariance across a student population. When raters are introduced into the measurement process, this requirement is particularly relevant for addressing issues related to the quality of ratings in writing assessment (Engelhard, Gordon, & Gabrielson, 1992).

This study examines measures of rater agreement, errors, and systematic biases, as well as the ability to make inferences regarding the validity of ratings of writing ability by practicing school teachers on a large sample of students from a national educational monitoring project in New Zealand. In this study, models based on principles of invariant measurement from item response theory will be used along with Brunswick’s (1952) lens model for examining perceptions of reality to compare the implications of evaluative criteria for rating quality for male and female students. Examining the consistency of marking behavior separately for boys and girls, the findings have implications for the validity of the ratings that may well be generalizable to classroom practice.

The data set for this study comes from New Zealand’s National Education Monitoring Project (NEMP). In NEMP, students are assessed on a variety of writing tasks which are then scored by teams of trained teacher assessors. The most recent administration of NEMP in writing (2010) has shown substantial discrepancies between boys’ and girls’ performance at both Year 4 (grade 3) and Year 8 (grade 7) on randomly-drawn, nationally representative samples of students (over 1400 students at each level). The question arises as to whether the differences seen in these performances might be attributable to different standards of marking for boys and girls. There are several aspects of the NEMP approach that make it particularly appealing here. First, authentic writing tasks are used. Second, there is a large sample of students at Year 4 and Year 8. Third, the same tasks are used at both years. And fourth, and critically, the writing tasks are marked by practicing elementary school teachers. Thus, there is the ability to use psychometric approaches requiring larger samples on a data set of natural school-like tasks, that are marked by teachers. The many-faceted Rasch Model, and the Brunswick Lens model are applied to this robust data set to look for consistency and inconsistency that might provide basic information about how teachers look at writing from boys and girls. All data for this study have been collected.

Author