Paper Summary
Share...

Direct link:

Comparing Preservice and In-Service Teachers' Judgments of EFL Student Essays With Expert Ratings

Fri, April 5, 12:00 to 1:30pm, Metro Toronto Convention Centre, Floor: 200 Level, Room 201A

Abstract

The assessment of students’ written performance is one central aspect of diagnostic competence as it directly influences the self-assessment of students (Möller & Trautwein, 2015) and their course achievement (Hamp-Lyons, 2016; Weigle, 2002). Research has shown substantial difficulties for teachers to judge students’ texts objectively (Ingenkamp & Lissmann, 2007; Kaiser et al., 2016). Previous research suggests that teachers’ knowledge and experience have positive effects on judgment accuracy (Baumert & Kunter, 2006; Klieme & Leutner, 2006). Yet, empirical field studies find no effects of teachers’ knowledge and experience (Meadows & Billington, 2010; Royal‐Dawson & Baird, 2009) or negative effects of teaching experience (Ready & Wright, 2011). However, in these studies teacher experience and knowledge were confounded, limiting the validity of results. To distinguish the two variables, we conducted an experimental study to investigate if in-service teachers judge upper-secondary English as a foreign language (EFL) essays more accurate than pre-service teachers. Additionally, we conducted a knowledge test (content knowledge and pedagogical content knowledge). The following research questions were addressed:
1. Do in-service teachers judge student essays more accurately than pre-service teachers?
2. Controlling for teachers’ knowledge, do in-service teacher judge student essays more accurately than pre-service teachers?
3. Do in-service teachers score higher in the knowledge test than pre-service teachers?
Methods
Sample
The sample consisted of 47 pre-service teachers (72.3% female, mean age 24.89 years [SD=2.04]) and 36 in-service teachers (75.0% female, mean age 46.86 years [SD=11.72], mean teaching experience 16.06 years [SD=10.05]).
Essays and expert ratings
The participants judged four upper-secondary EFL essays chosen out of 906 essays from the project Measuring English Writing at Secondary Level (Keller, 2016). For every essay, expert ratings were assigned by the Educational Testing Service® (ETS) an automatic scoring algorithm (e-rater®) and two trained raters.
Measures
Teachers judged the essays on the six-point holistic NAEP rating scale (Driscoll et al., 2010) and on four-point analytic rating scales with seven dimensions adapted from the 6+1 trait model (Culham, 2003). The knowledge test (=.79) was based on the Cambridge English Certificate of Proficiency in English (UCLES, 2015) and Cambridge Teaching Knowledge Test (UCLES, 2007).
Results and Discussion
Descriptive results showed that pre-service teachers’ holistic (Figure 1) and analytic assessment (Figure 2) was closer to expert ratings than in-service teachers’ assessment. A multivariate ANOVA showed significant multivariate difference between in-service and pre-service teachers (F(8,74) = 5.13, p<.001), even when we controlled for teacher knowledge (F(8,73)=3.89, p<.001). Post-hoc-Tests showed significant differences for all scales except grammar (mean d=1.01). As expected, in-service teachers (M=23.53, SD=5.57) scored higher (t(81)=-5.69, p<.001) in the knowledge test than pre-service teachers (M=17.09, SD =4.74).
Findings replicate previous research showing that knowledge (rknowledge.accuracy=-.228) and teaching experience do not improve accuracy of teacher judgments. Considering implications for practice, these findings suggest that difficulties which teachers show in assessing essays should be addressed directly, since they will not simply disappear by doing the job. Further investigations analyzing other factors affecting judgment accuracy are needed in order to shed light on the complex process of assessing student writing.

Authors