Paper Summary

Using Experiments to Optimize the Validity of a Practical Measure

Sat, April 14, 12:25 to 1:55pm, Vancouver Convention Centre, Floor: Second Level, West Room 208&209

Abstract

In the first two experiments, we test one of the most widely used forms of survey questions: questions that offer response options on a scale ranging from “agree” or “strongly agree” to “disagree” or “strongly disagree.” Agree / Disagree questions have several widely documented problems (e.g., Saris, Revilla, Krosnick & Shaeffer, 2010). They are subject to acquiescence response bias, which means that people tend to simply “agree” with statements due to processes of informational influence (producing the well-documented result that reverse-valenced items often have lower factor loadings on aggregate scales). Such bias decreases validity. They also have the disadvantage of being vague: quite often, such questions are not directly tied to a single behavior or attitude, and it is not clear why the respondent disagreed or agreed. As a result, they may be less directly useful for improvement research--when the goal is to maximize face validity of each item so as to directly inform improvements.

Hence, in two experiments we wrote alternative versions of sub-scales measuring (a) math anxiety and (b) study skills. Respondents were randomly assigned to receive either the Agree / Disagree format or in what we call a “Construct Specific” format. Construct specific questions directly asked respondents for their opinions. For instance, an Agree/Disagree question stating “I feel anxious the moment before a test” could be re-written as “How anxious do you usually feel the moment before a test,” and provide responses from “Not at all anxious” to “Extremely anxious.” Our analytic method is to compare the correlation between math test performance the Agree/Disagree version of the scale to the Construct Specific version of the scale. Analyses of this data are ongoing.

Next, we tested the question of whether motivation and engagement items are more highly predictive of grades when they have been answered after completing a math test as compared to before. We did this by randomly assigning half of respondents in our pilot to complete the math test as the first set of 20 questions in the survey or as the last set of 20 questions. We then compared the criterion validity of, for instance, math anxiety predicting overall math grades in each order.

Interaction effect regression analyses showed that, in multiple comparisons, motivation and engagement measures were significantly more predictive of grades when students had completed math items first. Hence, by leading students to experience test-taking, perhaps their true attitudes were more vivid and therefore they reported them more accurately. However, completing the math items first had some costs. Almost 25% of respondents in the math first condition dropped out of the survey, compared to only 7% in the math last condition. In addition, participants performed better when the math items were the last questions on the survey. As a result, this experiment provided a real tension between maximizing concurrent validity and increasing the amount of data that one collects and can use for improvement. We will discuss this result in light of competing needs of practical improvement research and theory-testing and theory-building research.

Author