Paper Summary

Science Assessment and Causal Inference: A Simulation Study on Regression Discontinuity and Propensity Score Matching

Sun, April 15, 10:35am to 12:05pm, Pan Pacific, Floor: Lobby Level, Oceanview 1&2

Abstract

Large-scale assessment, as a natural experiment (Shadish & Cook, 2009), commonly uses a non-randomized design that is realized in the evidence-based studies of science education. In a natural experiment, the treatment effect is not estimatable because the units in treatment have not been assessed in control. Rubin's Causal Model (Holland, 1986) estimates the treatment effect based on a counterfactual situation. That is, examinees exposed to treatment would have been also assessed in control. Propensity score (PS) matching (Rosenbaum & Rubin, 1983) realizes the causal modeling through a post hoc sampling that creates statistically comparable treatment-control group units. Regression discontinuity (RD) analysis, as a post-hoc analytical approach, can also estimate the treatment effect at the threshold/cutoff point. This study will use simulated data to examine and compare two approaches in the context of science proficiency assessment.

Treatment/control group status in RD analysis is determined by a cutoff score on the selection measure (Schenider, et al, 2007). Participants that are close (just above or below) to the cutoff have comparability “on a set of unobserved variables that predict scores on the test determining the merit of the selection measure” (p. 52). These participants have been assigned, in a random sense, into treatment group (just above the cutoff) and control group (just below the cutoff). Shadish and Cook (2009) pointed out that participants are assigned into the treatment group (indicated by 1) or the control group (indicated by 0) through the logical arithmetic using a latent continuum (indicated by L) and a specific cutoff point (indicated by l). That is: Group = 1 if L >= l and Group = 0 otherwise. This procedure approximates a probability-based (i.e., random) assignment. The estimate of regression model on the cutoff point produces the unbiased estimator of treatment effect.

Traditionally, RD analysis restricts the participants used for comparison to those who are JUST below or above the cutoff point. Thus, the treatment and control groups don’t have overlap (i.e., common support) on the selection measure. This will make the PS matching impossible. In order to compare RD and PS, this study focuses on multi-dimension large scale assessment in science education. A dimension represents a subject domain of knowledge. TIMMS is an example of such assessment, which focuses on science and mathematics. RD has been realized in multi-dimension assessment in English language and mathematics education (Robinson, 2010). However, Robinson (2010) did not compare RD with PS.

The simulation study compares RD analysis and PS matching by computing the propensity score that is the probability of a participant being assigned into the treatment group. The PS score model bases on the selection measure and the cutoff point of RD analysis. The RD analysis and PS matching follow a latent variable modeling (LVM) framework to account for measurement error to achieve robust treatment effect estimate (Wang, Maier & Houang, 2011). The proposed study innovatively links large scale assessment and causal inference in program evaluation in science education.

Author