Paper Summary
Share...

Direct link:

A 12-Trait Model of Variation Among Student Essays

Mon, April 25, 2:30 to 4:00pm PDT (2:30 to 4:00pm PDT), Marriott Marquis San Diego Marina, Floor: South Building, Level 1, Leucadia

Abstract

This paper presents a confirmatory factor analysis that combines features from three engines: the e-rater engine for automated writing evaluation, the Text Evaluator engine for readability and genre analysis, and the WAVES engine that underlies Writing Mentor, a writing feedback tool. We identified twelve dimensions of variation, based upon features from these engines: aspects of text structure (organization and text cohesion), aspects of linguistic sophistication (vocabulary length and difficulty, sentence length and complexity, and conventions), and dimensions of genre variation (academic language, concreteness of language, oral/interactive style, argumentative stance taking, and narrative contextualization).

We developed and validated this model using a corpus of 1.37 million essay submissions collected during 14 years of operation by the Criterion digital writing tool, and further validated it using additional essay corpora: a corpus of >20,000 scored essays used to create e-rater models for Criterion, a corpus of >20,000 mostly middle school essays collected by ETS research, a corpus of >33,000 essays submitted as part of a high school equivalency test, a corpus of EFL essays submitted as part of a test preparation MOOC, and a corpus of student college papers collected by Burstein et al. (2020). The traits underlying the model are consistent with the literature, and fit our study datasets at an acceptable level, with RMSEA values below .07 and CFIs around .85.

We then examined the relation between trait scores and external variables. In the Criterion data, trait scores were strongly predictive of human score and showed clear longitudinal trends, with most traits increasing significantly within the school year and between grades. There were strong genre differences, especially on traits like stance-taking and contextualization, which derived from features used to differentiate genres in a Biber-style factor analyses. Some traits (especially the conventions trait) showed significant increases after revision. We were able to link the Criterion data with IES school-level demographics. Analysis of demographic effects showed commonly observed patterns of female advantage, and low-SES disadvantage, on many writing traits, but also unexpected areas of strength or emphasis; for example, schools with higher levels of male enrollment tended to have higher average stance-taking scores, while schools with higher levels of female enrollment tended to have higher contextualization (narrativity) scores, regardless of genre. When we examined other corpora, the model continued to fit acceptably well, and we continued to see strong relationships with score, clear genre differences, and some evidence of developmental effects. Interestingly, the model fit acceptably well even in the college-level data analyzed by Burstein (2020), which contained longer forms, including research papers and literary analyses. In the college data, we were able to examine the relation between trait scores and external measures, such as SAT scores and class grades. The language traits, which are the most likely to be stable measures of linguistic ability, as opposed to measures of genre variation, had weak but significant relations to these measures. These results suggest the trait model could be used to track student growth and identify strengths and weaknesses in a classroom context.

Authors