Paper Summary
Share...

Direct link:

The Role of Essay Length in Foreign Language Assessment: Appropriate Heuristic or Disruptive Factor?

Fri, April 5, 12:00 to 1:30pm, Metro Toronto Convention Centre, Floor: 200 Level, Room 201A

Abstract

In the context of high-stakes assessments of (foreign) language proficiency, prompt-based essay tasks are often used to measure writing skills (Guo, Crossley & McNamara, 2013). Still, adequate and reliable scoring constitutes a major challenge because different (text-inherent) influence factors on human ratings lower the judgement accuracy (e.g. Wolfe, Song & Jiao, 2016). One of those factors is essay length (EL). Exemplarily, in a study by Kobrin, Deng & Shaw (2011), EL and human scores were significantly positively related (.66) - even when controlling for multiple-choice measures of reading and writing skills (.59). This correlation is not an unusual finding and emphasizes a serious difficulty in writing assessment: Since text quality is mostly operationalized via scores assigned by trained human raters, it is unclear whether the correlation between EL and human scores reflects a relation between EL and text quality or whether it stems from judgement biases.
The increasing use of automated essay scoring stresses the importance of this issue. Corresponding systems offer a promising way to complement human ratings (Deane, 2013). However, it has been criticized that at this point they mainly count words when computing a score (Perelman, 2014). Thus, it is essential to investigate whether the consideration of EL is an appropriate heuristic when humans or automated systems rate essays. As assigned scores are supposed to reflect writers’ productive foreign language proficiency, the appropriate heuristic assumption could be supported by two findings: First, EL should be strongly associated with other variables of students’ English proficiency. Second and consequently, EL should not explain incremental variance above and beyond writers’ proficiency in human respectively automated essay scores. The aim of the current study was to investigate this empirically.
Methods
The sample consisted of N=1,867 upper secondary students (11th grade, 59.8% female, M=17.62 years) in Germany and Switzerland, answering an independent prompt of the Test of English as a Foreign Language (TOEFL iBT; ETS). The essays were scored on a 6-point rating scale at ETS. ETS also provided the trained human raters (two ratings per essay) and the automated scoring system (e-rater®). EL was operationalized via word count. The latest English grade, measures of English listening and reading skills as well as estimates of English self-concept constituted the students’ English proficiency variables.

Results and Discussion
As expected, EL correlated strongly with the two human ratings as well as with the e-rater® score (see Table 1). Comparatively, associations between EL and proficiency variables were moderate.
The results of the stepwise regression used to answer the second part of our research question are provided in Tables 2-4. When controlling for proficiency, EL still accounted for 11-15% of the variance in human essay scores. Regarding the e-rater® score, EL raised the explained score variance to the extent of one third (41 to 61%).
Taken together, these findings indicate that the associations between EL and assigned scores might not simply reflect a relation between EL and text quality. Rather, they suggest that EL might actually be a disruptive factor in human and automated essay scoring.

Authors