Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Visiting Washington, D.C.
Personal Schedule
Sign In
X (Twitter)
For the past century, standardized history exams in the United States have emphasized factual recall. In recent years, there have been renewed calls for tests that assess more complex aspects of historical thinking (VanSledright, 2014; Wineburg, 2004). In response, we developed a series of short, document-based tasks that target specific aspects of historical thinking. We call these items History Assessments of Thinking (HATs).
This presentation will detail research that considers whether HATs measure the historical thinking constructs they were designed to measure. Our study explored aspects of cognitive validity, or the relationship between the kind of thinking the items were designed to elicit and the thinking they actually do elicit (Linn, Baker, & Dunbar, 1991; Pellegrino, Chudowsky, & Glaser, 2001). More specifically, we used think-aloud interviews to compare the thinking elicited by HATs with the thinking elicited by multiple-choice versions of HATs.
The study materials were eight HATs and eight parallel multiple-choice items. The multiple-choice parallels used the same documents as the parallel HATs and were designed to measure the same construct. The data for this study includes interview protocols collected as students completed the selected HATs and parallel multiple-choice questions. High school students (n = 26) were randomly assigned to one of two groups, each of which received a different test form that included four HATs and four multiple-choice versions of HATs. HATs and their multiple-choice parallels appeared on different forms. Thus, the multiple-choice items on the first form were parallels to HATs from the second form.
Student interview data was used to conduct two analyses. The first examined whether selected HATs tap the historical thinking constructs they were designed to measure. Student transcripts were coded for the alignment between the thinking processes elicited by the items and the constructs that the items were designed to measure. The second analysis compared the thinking elicited by HATs with that elicited by the multiple-choice versions of HATs. Protocols were coded for the degree of alignment between the processes used to arrive at the correct answer and the constructs targeted by each assessment. The processes elicited by each HAT were then compared to the processes elicited by its multiple-choice parallel to determine if HATs are better indicators of students’ historical thinking than multiple-choice questions.
Results suggest that HATs do measure the targeted aspects of historical thinking. Results also suggest that student performance on HATs better reflect student proficiency in historical thinking than the multiple-choice versions. The think-aloud data also showed that item format mattered—and mattered a lot. With multiple-choice questions, some students were able to arrive at the correct answer using test-taking strategies unrelated to the targeted construct and other students showed some understanding of the targeted skill but ultimately selected the wrong answer.
This study suggests that new forms of history assessment can be used to assess aspects of historical thinking and that these items present advantages over multiple-choice questions for measuring higher-order skills. These findings could influence both classroom assessment and the re-design of new summative history exams.
Joel Breakstone, Stanford University
Sam Wineburg, Stanford University
Mark D. Smith, Stanford University