Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Register for SRCD21
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Lord’s paradox describes two ways that analysts can treat prior scores when trying to estimate the effect of a variable on final scores: ANCOVA and gain scores. The choice between ANCOVA and gain scores can have important consequences when trying to reach causal conclusions. The focus here is on using student test scores as part of the procedure to estimate the effectiveness of schools. In this case students’ prior scores are used when predicting scores after some amount of schooling in order to estimate the differential effect of each school on student performance.
Because students are not randomly allocated to schools, data models are developed for how different variables may influence which school a student attends. Consider two models depicted in Figure 1.
In panel A what school a student attends relates to the geographical and demographic groups the student is in, and these will be correlated with ability. Typically, where a student lives is the greatest influence on which school they attend. In Panel B, the prior score is used to stream students into different schools. This can occur to some extent with SAT/ACT scores and university admissions but is less common for PK-12.
Data were simulated for these two models. A multilevel gain score model (using the conditional models for school effectiveness estimates) provided good estimates for Panel A, but poor estimates for Panel B. The reason is that if students are chosen for a school because of high scores on a prior assessment, their scores are likely to regress towards the mean. A multilevel ANCOVA model (the value-added model, or VAM) performed poorly for Panel A, but relatively well for Panel B. Because Panel A is more common for PK-12 school allocation and popular methods for assessing schools (e.g., VAMs, student growth percentile models) are ANCOVA based, this is a concern for the fairness of using these models for estimating school effectiveness. The VAM and other ANCOVA procedures produce lower scores for schools that serve historically low performing groups than is appropriate. This raises ethical questions for recruiting principals and teachers to these schools and questions about why a method that is predicted from Galton’s 1886 paper to be biased (and easily shown to be biased) was adopted by some states.
Other more involved data models are considered, and data simulated for these. Situations are shown where the VAMs/ANCOVAs perform poorly (i.e., where the correlation between the estimated school effectiveness and true effectiveness used to create the data can be negative). The author is currently working with state officials to develop protocols to simulate data models using proposed statistical procedures in order to evaluate their reliability, fairness, and validity before they are used in high-stakes decision making. Progress in this collaboration will be discussed.