Paper Summary
Share...

Direct link:

Spring or Fall Annual Tests? Implications for Value-Added Models

Mon, April 16, 2:15 to 3:45pm, New York Hilton Midtown, Floor: Second Floor, Madison

Abstract

1. Objectives
The proposed study fills a gap in the value-added measure (VAM) literature by directly addressing the question of whether fall or spring annual assessments yield more accurate VAMs. The concern is that annual, “cross-year” VAMs necessarily mis-attribute summer learning to the teacher or school being rated by the VAM. We address the proposed research question by comparing the accuracy of fall-to-fall and spring-to-spring “cross-year” VAMs relative to the arguably more valid fall-to-spring “within-year” VAMs.


2. Theoretical Framework
Socioeconomic gradients in summer learning rates are well documented (Cooper et al. 1996). It is concerning, then, that most school districts rely on standardized tests that are administered once per academic year, usually in the spring, to evaluate teacher and school effectiveness. This approach is potentially problematic, as students’ summer learning is incorrectly attributed to students’ current teachers and schools. The resultant potential bias is well documented (Downey et al., 2008; Gershenson and Hayes, 2016; McEachin and Atteberry, 2017). Accordingly, if high-stakes tests are to be administered only once per year, it is vital for school administrators to know whether implementing those assessments in spring or fall yields reduces more credible VAMs.

3. Methods
Following Gershenson and Hayes (2016), we make two sets of comparisons. First, we compare VAM-based rankings generated by fall-to-spring (within-year) achievement gains to the arguably less-valid rankings generated by spring-to-spring (cross-year) gains. Second, we compare VAM-based rankings of fall-to-spring (within-year) achievement gains to the arguably less-valid rankings generated by spring-to-spring (cross-year) gains. Whichever of the two cross-year VAM rankings more closely resembles the within-year VAM rankings is arguably more accurate, and thus the preferable testing date for schools that are limited to having one assessment per year.

4. Data
We use two complementary, student-level datasets. The first is the Early Childhood Longitudinal Study, Kindergarten Class of 2010-11 (ECLS-K). The second is the Northwest Evaluation Association’s (NWEA) Measure of Academic Progress (MAP) assessment. Both are commonly used, well documented, publicly available datasets that the authors have used in previous research (e.g., Gershenson and Hayes, 2016; McEachin and Atteberry, 2017). The unique feature of these data sets that enables the current research design is the presence of both fall and spring tests in consecutive years: we observe a given student’s test scores in the spring of grade g, the fall of grade g+1, and spring of grade g+1, and the fall of grade g+2.


5. Preliminary Results
Table 1 reports Spearman rank correlation coefficients for three comparisons made in the ECLS-K analysis. They suggest that estimated classroom effects from VAMs using spring-to-spring achievement gains for both subjects are significantly more robust to test timing than similar VAMs using fall-to-fall achievement gains.


6. Policy Implications
Preliminary results suggest that spring-to-spring “cross-year” VAMs are more valid than fall-to-fall “cross-year” VAMs. If only one assessment per year is feasible, spring assessments are preferred to fall assessments, at least when the objective is to obtain valid estimates of school or teacher effectiveness.

Authors