Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
In the context of high-stakes assessments of (foreign) language proficiency, prompt-based essay tasks are often used to measure writing skills (Guo, Crossley & McNamara, 2013). Still, adequate and reliable scoring constitutes a major challenge because different (text-inherent) influence factors on human ratings lower the judgement accuracy (e.g. Wolfe, Song & Jiao, 2016). One of those factors is essay length (EL). Exemplarily, in a study by Kobrin, Deng & Shaw (2011), EL and human scores were significantly positively related (.66) - even when controlling for multiple-choice measures of reading and writing skills (.59). This correlation is not an unusual finding and emphasizes a serious difficulty in writing assessment: Since text quality is mostly operationalized via scores assigned by trained human raters, it is unclear whether the correlation between EL and human scores reflects a relation between EL and text quality or whether it stems from judgement biases.
The increasing use of automated essay scoring stresses the importance of this issue. Corresponding systems offer a promising way to complement human ratings (Deane, 2013). However, it has been criticized that at this point they mainly count words when computing a score (Perelman, 2014). Thus, it is essential to investigate whether the consideration of EL is an appropriate heuristic when humans or automated systems rate essays. As assigned scores are supposed to reflect writers’ productive foreign language proficiency, the appropriate heuristic assumption could be supported by two findings: First, EL should be strongly associated with other variables of students’ English proficiency. Second and consequently, EL should not explain incremental variance above and beyond writers’ proficiency in human respectively automated essay scores. The aim of the current study was to investigate this empirically.
Methods
The sample consisted of N=1,867 upper secondary students (11th grade, 59.8% female, M=17.62 years) in Germany and Switzerland, answering an independent prompt of the Test of English as a Foreign Language (TOEFL iBT; ETS). The essays were scored on a 6-point rating scale at ETS. ETS also provided the trained human raters (two ratings per essay) and the automated scoring system (e-rater®). EL was operationalized via word count. The latest English grade, measures of English listening and reading skills as well as estimates of English self-concept constituted the students’ English proficiency variables.
Results and Discussion
As expected, EL correlated strongly with the two human ratings as well as with the e-rater® score (see Table 1). Comparatively, associations between EL and proficiency variables were moderate.
The results of the stepwise regression used to answer the second part of our research question are provided in Tables 2-4. When controlling for proficiency, EL still accounted for 11-15% of the variance in human essay scores. Regarding the e-rater® score, EL raised the explained score variance to the extent of one third (41 to 61%).
Taken together, these findings indicate that the associations between EL and assigned scores might not simply reflect a relation between EL and text quality. Rather, they suggest that EL might actually be a disruptive factor in human and automated essay scoring.
Anna Lara Paeske, Leibniz Institute for Science and Mathematics Education
Jennifer Meyer, Leibniz Institute for Science and Mathematics Education
Thorben Jansen, University of Kiel
Johanna Fleckenstein, Pädagogische Hochschule FHNW
Stefan Daniel Keller, University of Applied Sciences of Northwestern Switzerland
Olaf Koeller, Leibniz Institute for Science and Math Education