Paper Summary
Share...

Direct link:

The Psychometrics of Situational Judgment Tests

Sun, April 6, 12:25 to 1:55pm, Convention Center, Floor: 100 Level, 112A

Abstract

This study examines current approaches to psychometric evaluation of SJT items, and will present them in the context of several key aspects of the evaluation process with real data: (1) item- and test-level validity, (2) internal consistency, and (3) scoring. While many approaches are familiar to those who work with SJTs, newer or potential methods for analyzing SJTs are seemingly unknown to most. By way of comparison to older and more familiar methods, we hope to shed light on alternatives to enhance evaluation options for SJTs and strengthen the conclusions researchers and practitioners draw from their use.

Employers and educators have become increasingly interested in the set of traits and characteristics often tied to interpersonal skills. Knowledge of a person’s interpersonal skills can provide an understanding of behaviors by identifying qualities like their ability to work with others and their approach to trying situations. Situational judgment tests (SJTs) are regularly used to measure these skills, but have been used to address a number of other constructs as well. Research has shown SJTs exhibit greater criterion-related validity relative to other types of measures (e.g., most personality test adopting traditional Likert-type responses), and have demonstrated greater resistance to biases like social desirability. As well, due to the structure and format of SJTs, they are not as susceptible to measurement confounds linked to demographic factors like gender and ethnicity because of less cognitive demand.

Despite these positives, SJTs generally have poor internal consistency, and can seem somewhat unsatisfactory when evaluated using methods like structural factor analysis . Many researchers note complications with these indices of internal stability can be attributed to the multidimensional nature of most SJTs, often considered a positive aspect of SJTs. SJTs allow for the measurement of several constructs simultaneously, but tend not to be as transparent as other types of tests where items are more clearly tied to particular dimensions. This creates an interesting problem for SJTs though, because opaque measurement can make it more difficult to determine if the constructs of interest are truly being measured.

Finally, there is evidence that different scoring methods can lead to different conclusions. For instance, approaches based in IRT, which remove some residual noise, may have more criterion-based validity than scores based in classical test theory (e.g., simple composite sums, or consensus scores).

Using data collected from Amazon Mechanical Turk, we will evaluate the psychometric properties of 25 teamwork SJT items and 58 interpersonal skills SJT items. Analyses include estimates of criterion and incremental validity, Cronbach’s alpha, exploratory and structural factor analysis, unidimensional and multidimensional item response theory models, and structural equation models. The study findings would not only provide useful information about the two measures of interest, but could help clarify when a particular set of methods, including newer or to-be-explored methods, would be appropriate for their SJT data.

Authors