Individual Submission Summary
Share...

Direct link:

Improving Measures of Skin Color

Sun, September 1, 10:00 to 11:30am, Marriott, Virginia B

Abstract

Many social scientists and, increasingly, government statistical agencies, are collecting measures of skin color for the purpose of studying discrimination and racial / ethnic disparities in socioeconomic outcomes. One popular measure is the “Massey-Martin Scale,” a Likert-type measure based on an interviewer’s rating of the subject against a skin-color palette which the interviewer has committed to memory. The present study uses MTurk data to investigate problems with the Massey-Martin Scale, and to explore alternatives.

We have three specific objectives. First, we seek to replicate and confirm previous findings that skin-color ratings using the Massey Martin scale vary with the race of the coder. Second, we test Massey-Martin ratings for another potential problem, one overlooked in the literature to date: image-spillover effects. We hypothesize that persons or photos rated following a sequence of “dark” subjects will be rated lighter than if the same person or photo had been rated following a run of “light” subjects. Such sequencing or spillover effects could result in systematic biases if—as seems likely—a black person who lives in a heavily black geographic region is more likely to be rated following other black subjects than is a black person who lives in a mostly white geographic region. Third, we compare and evaluate several strategies for dealing with race-of-coder and spillover problems: (1) averaging multiple observations of the same subject; (2) allowing the coder to view the Massey Martin palette while she evaluates and records the skin color of the person she is coding; and (3) eliciting pairwise comparisons of skin color (rather than using the Massey Martin scale), and then inferring skin-color rankings using Bayesian item-response models. To evaluate the sensitivity of different coding methods to variation in the race of the coder, and to variation in the race of persons coded prior to the target person, we construct an experiment in which we randomize the assignment of coders to
various coding conditions, and then analyze subsets of the data that have been generated through different coding procedures.

Authors