Paper Summary
Share...

Direct link:

Development of the Measure of Original Thinking in Elementary Students: Preliminary Evidence of Computational Psychometric Properties

Fri, April 22, 2:30 to 4:00pm PDT (2:30 to 4:00pm PDT), Marriott Marquis San Diego Marina, Floor: South Bulding, Level 3, Balboa

Abstract

Divergent thinking (DT) tasks have been often used for creativity assessment (Acar & Runco, 2019) and gifted identification (Kim, 2006). However, their scoring has been labor-intensive and expensive. Recently, researchers have developed novel scoring methods for originality that rely on text-mining methods. These methods score the originality of a response based on its computed semantic distance from a DT prompt (Dumas et al., 2020; Forster & Dunbar, 2009), where closely related prompt-response pairs score lower on originality. Past research typically used Alternate Uses Tasks with college students using general language corpora (e.g., GLoVe, Pennington et al., 2014). Acknowledging that measuring DT in children requires different considerations, we developed an instrument called Measuring Original Thinking in Elementary Students (MOTES), with a new corpus (i.e., MOTES corpus) that is reflective of children’s language (Organisciak et al., 2021). In this study, we present evidence on the psychometric properties of the MOTES instrument and corpus based on a pilot of 35 third through fifth grade children.
The preliminary version of MOTES instrument comprised four sets of verbal DT tasks: Uses, Instances, Consequences, and Scenarios. Following expert validation with 32 experts (i.e., elementary teachers and creativity experts), the final set included 28 prompts. To avoid the fluency confound (Forthmann et al., 2020), participants were instructed to give a single response for each prompt and were asked, specifically, to think of a surprising response. The instructions were presented in a short, animated video where a creative character modeled this original response style for the respondents. In a cognitive pilot, children (N = 35) typed in their responses via Chromebook in a one-on-one session with an adult expert. The preliminary data from the pilots were collected without a time limit to determine the optimal time needed.
We first examined if MOTES corpus could successfully reflect the semantic structure of children’s language. Thirty-two expert judges rank ordered five words in terms of their relation to a target word (e.g., spy: investigators, review, pictures, espionage, fictional), which were compared to the text model. Over 10 sets of words, normalized discounted cumulative gain--a measure of rank quality comparing the computer model’s ranks with expert truth–provided a value of 0.975 (perfect score is 1).
Evaluated against the pilot children’s data, we compared the MOTES and GloVe models for DT scoring, using the Open Creativity Scoring platform (https://openscoring.du.edu). To compare if the corpus used (MOTES vs GLoVe) makes any difference, we examined the correlations with human-rated originality of the same responses (Mdn ICC (1,2) = .75). MOTES had a stronger correlation with human raters than GloVe across all items (r = .31 vs r = .26) as well when a subset of prompts that are most suitable for automated scoring were considered (r = .37 vs r = .32). Those findings indicate that automated scoring of DT responses given by children can be better captured when children-specific corpus such as MOTES was used.

Authors