Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Visiting Washington, D.C.
Personal Schedule
Sign In
X (Twitter)
Objective: To review early studies of grading reliability. The earliest investigation of which the authors are aware is a statistical study published in the Journal of the Royal Statistical Society. Edgeworth (1888) applied normal curve theory to the case of grading examinations.
Method and Data Sources. Electronic searches using the terms “teachers’ marks (or marking)” and “teachers’ grades (or grading)” were conducted in the following databases: ERIC; the Journal of Educational Measurement (JEM) and Educational Measurement: Issues and Practice (EMIP), ProQuest’s Periodicals Index Online, and the Journal of Educational Research (JER). Abstracts of identified studies were read to determine if the research addressed individual pieces of student work (usually examinations) and not composite report card grades. The reference sections of qualifying studies were searched to identify additional relevant studies. This resulted in a pool of 15 studies published between 1910 and 1935 and one study published in 2011 that replicated a 1912 study. In addition, two early reviews of studies (Kelly, 1914; Rugg, 1918) summarized 21 other relevant studies.
Results and Conclusions. These early studies of grading focused primarily on reliability and not on other aspects of validity. The main finding about the reliability of teachers’ grades was that teachers vary greatly in their assignment of grades. Three studies (Bolton, 1927; Jacoby, 1910; Shriner, 1930) argued against this conclusion, however, contending that teacher variability in grading was not as great as commonly supposed. Many of the early studies of grading unreliability present a “what’s wrong with teachers” undertone that today would likely be seen as researcher bias. Nonetheless, taken as a whole this collection of studies allows discussion of several sources of variation.
Early researchers attributed sources of variation in teachers’ grades to one or more of the following sources: criteria (Ashbaugh, 1924; Brimi, 2011; Healy, 1935; Silberstein, 1922; Sims, 1933, Starch, 1915; Starch & Elliott, 1913a,b), students’ work quality (Bolton, 1927; Healy, 1935; Jacoby, 1910; Lauterbach, 1928; Shriner, 1930; Sims, 1933), teacher severity/leniency (Shriner, 1930; Silberstein, 1922; Sims, 1933; Starch, 1915; Starch & Elliott, 1913b), task (Silberstein, 1922; Starch & Elliott, 1913a), scale (Ashbaugh, 1924; Sims, 1933; Starch 1913, 1915), and teacher error (Brimi, 2011; Eells, 1930; Hulten, 1925; Lauterbach, 1928, Silberstein, 1922; Starch & Elliott, 1912, 1913a,b). Starch (1913, Starch & Elliott 1913b) found that teacher error and emphasizing different criteria were the two largest sources of variation.
Regarding criteria, most of the studies investigating teachers’ grading reliability sent examination papers to teachers without specific grading criteria and simply asked teachers to assign grades. Today this would be considered a fault in assessment design. Most of these early studies of grading confound teachers’ inability to judge student work consistently and random error, considering both “teacher error.” Rater training offers a modern solution to the first problem (Myford, 2012). Working toward clearer criteria and involving students in the development of those criteria appear to be promising approaches to enhancing grading reliability.