Paper Summary

Sampling Issues for Using Science Notebooks as Assessment Tools

Sun, April 15, 10:35am to 12:05pm, Pan Pacific, Floor: Lobby Level, Oceanview 1&2

Abstract

Among numerous efforts to improve students’ science learning in schools, student notebooks have been proposed and investigated as an instructional, unobtrusive assessment tool (e.g., Aschbacher & Alonzo, 2004; Ruiz-Primo & Li, 2004; Shepardson & Britsch, 1997). This paper explores the sampling issues of scoring notebooks so that notebook scores can be reliable interpretation of student learning at individual level and class level.

Science notebooks are a written account, with more or less detail, and of diverse quality, of what students do and learn in their science class (Ruiz-Primo & Li, 2004). Unlike other external assessments, notebooks require no additional time to administer and less likely involve negative consequences such as teaching to the test because they are exactly embedded in the instruction. Using notebooks as an assessment tool requires a reliable, efficient scoring method to generate scores to index the student performance level. However, characteristics of notebook entries vary greatly from entry to entry, reflecting the diverse set of activities implemented across different lessons, which presents a huge challenge for drawing reliable inferences about student understanding. We apply the Generalizability (G) theory to explore sampling facets of lesson, part within lesson, rater, and student within class, and provide empirical evidence on the effects of these facets and their interactions on measuring student learning.

We applied G theory to estimate the amount of score variance associated with lesson, part within lesson, rater, and the combinations of these facets. Using the student x rater x (part : lesson) design and other related designs, we conducted G studies to address whether: (1) well-trained raters can reliably score students’ notebooks entries; and (2) student performance was consistent across lessons and parts. We also estimated the optimal numbers of raters, lessons, and parts needed to maintain an acceptable reliability. Lastly, we considered the unit analysis as student within class instead of student to explore the sampling for drawing inferences of the class level performance index.

We conducted the analysis with two data sources: notebook scores of 10 elementary classes (6 students per class) on two science units and of 12 middle school classes (9 students per class) on one science unit. Both studies involved two raters of scoring student notebooks on 4 to 6 lessons.

We observed the main source of measurement error was lesson rather than rater or part. The result suggests at least two lessons need to be sampled when scoring student notebooks while only one rater and one part is needed. The findings are consistent across the grades. Regarding the sampling design for the class level interpretation, fewer students need to be sampled for elementary level than secondary level, indicating less variation within class than between classes for elementary classes.

This paper explores technical issues about sampling designs for using notebooks as assessment tools. Findings will provide useful information to promote the use of notebooks to address the tension between external testing and classroom assessment and to maximize the assessment value of notebooks.

Authors