Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Housing and Travel
Personal Schedule
Sign In
X (Twitter)
There has been an increasing amount of interest in what happens inside K-12 classrooms during instruction. This interest largely stems from a need to examine relationships between instructional practices and student learning outcomes (Ing et al., 2015). To accurately capture instructional practices and thereby investigate these relationships, measures are needed that generate information from which valid interpretations can be drawn. However, unlike traditional assessments such as achievement tests, collecting validity evidence for measures of instructional practices is relatively new and less straightforward (Bell et al., 2012). The goal of this paper presentation is to share the overlaying of two common frameworks for developing and evaluating validity arguments for two different measures of mathematics teaching practices, an observational protocol (OP) and an instructional log (IL). The frameworks are the Standards for Educational and Psychological Testing (AERA et al., 2014) and Kane’s (2013) interpretation-use argument (IUA).
In our work, we have utilized Kane’s framework for laying out the key inferences (scoring, generalization, extrapolation, and implications) and supporting assumptions that need to be evaluated. To evaluate the inferences, we outline the types of evidence from the Standards (e.g., evidence based on response processes; reliability) that are needed. By overlaying the two frameworks, the proposed interpretations and uses remain central as Kane’s framework emphasizes, but the evidence types provide a tangible way to organize the argument.
Here, for example purposes, we compare this approach for the generalization inference for the OP versus the IL, which measure nine and five dimensions (e.g., representations, problem solving) of mathematics instruction, respectively. The OP is utilized on video-recorded mathematics lessons by trained coders who are external to the classroom while the IL is completed by teachers daily about the mathematics lesson that they taught. For the generalization inference, there are two assumptions for each measure, both focused on reliability as described in the Standards because “the level of reliability/precision in scores has implications for validity” (AERA et al., 2014; p. 34). We assume that each measure can reliably measure the given dimensions of mathematics teaching practices and that unexplained error is minimized. For the latter assumption, generalizability studies examine the sources of variation to ensure sufficient accounting for unexpected error, but these sources vary by measure (e.g., OM: teacher, coder, rubric; IL: teacher, item). For the first assumption, decision studies for the OM determine how many video-recorded lessons are needed per teacher and how many coders are needed per lesson to obtain a reliable estimate of a teacher’s mathematics teaching practices. In contrast, a decision study for the IL focuses on how many days of logging are needed for a reliable estimate.
The endeavor of measuring mathematics teaching practices is complex, but one that should be taken seriously. Identifying variation in mathematics teaching practices has implications for allocating resources and focusing instructional improvements. However, without data that produces valid inferences, measurement efforts become pointless. Attending to validity by overlaying Kane’s framework with the Standards provides a rigorous, careful, and tangible approach.