Paper Summary

Empirical Investigation of the Construct in a Brief Electricity and Magnetism Assessment

Sun, April 15, 10:35am to 12:05pm, Pan Pacific, Floor: Lobby Level, Oceanview 1&2

Abstract

This paper seeks to explore the construct validity of a Brief Electricity and Magnetism Assessment (BEMA) (Ding, Chabay, Sherwood & Beichner, 2006). BEMA is a 30-item multiple-choice test designed to measure student conceptual understanding of basic electricity and magnetism topics at the introductory physics level. BEMA mainly targets comprehension and application of these topics as opposed to simple recall of facts.

Differing from regular concept inventories which focus on one topic in each assessment (Wallace & Bailey, 2010), BEMA covers a broad range of E&M concepts. This raises a question: do individual BEMA items form a cohesive construct? This study addresses this unanswered question.

We use Rasch model to fit student data on BEMA. The Rasch model is a stochastic theory. Through examining the individual items of an assessment, the Rasch modeling generates both item difficulty and person ability estimates at the interval scale, independently of each other (Linacre, 2011). An important assumption of Rasch modeling is that the individual items should measure the same construct (ability or trait), namely unidimensionality (Linacre, 2011). In the case of this assumption being violated, Rasch model will fail to fit the data. Consequently, results from the analysis are likely to show a noticeable deviation from acceptable ranges.

Based on the above framework, we use Rasch modeling to analyze individual BEAM items and the whole instrument. Specifically, we look into model-estimated reliability (person reliability and item reliability), person-item map, and item fit statistics (infit and outfit mean squares). Any items that do not fit under the unidimensionality display unsatisfactory fit statistics, resulting in a poor model fit with low reliability and/or person-item mismatch. The Winsteps program is used for Rasch analysis.

We administered BEAM and collected 648 data points from over a dozen calculus-based introductory E&M classes at the college level. The participants were science and engineering major students who took E&M as a mandatory course. We first screened the data to eliminate those who left blank on more than 60% of the questions before performing Rasch analysis.

The Rasch analysis of BEMA data shows satisfactory person and item reliability values. Also, a person-item map displays a good match between model-estimated item difficulty and student ability levels. Moreover, item fit statistics show that all except four items have acceptable infit and outfit mean squares. Only one item has both infit and outfit statistics out of the acceptable range. These results suggest that albeit the individual BEMA items may seemingly measure different topics, they collectively form a cohesive construct.

BEMA has been broadly adopted for use in physics classes. However, no evidence was shown it is an assessment with construct validity. This study fills the gap and empirically confirms that BEMA measures what it purports to measure, which is student understanding of basic E&M topics. Also, the methodology and techniques used in the study can be extended to the development and analysis of other assessments in science education, especially those with broad content coverage.

Author