Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Register for SRCD21
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Verbal children communicate how objects, people, and spaces are represented by using spatial language – words that describe features, shapes, dimensions, locations, or orientations (e.g., “turn that flat block around,” “that’s a tall man”). Spatial language has demonstrated relations with children’s individual differences in mental transformation skills, spatial mapping, math knowledge, and more (Loewenstein & Gentner, 2005; Pruden et al., 2011; Purpura et al., 2017). However, spatial language research is notoriously time consuming to conduct. Analysis of domain-specific spatial language often requires researchers to examine the context of each utterance before classifying words as spatial because many potentially spatial words (e.g., “on the left side”) can be used non-spatially (e.g., “he left yesterday”). To date, only one naturalistic, longitudinal sample has been analyzed for spatial language (Pruden et al., 2011; Pruden & Levine, 2017), yet longitudinal studies are vital for examining spatial development. A freely available machine learning algorithm that analyzes language data in a fraction of the time it takes using traditional methods may encourage investigators to conduct longitudinal studies and further elucidate developmental trajectories of children’s spatial language production and related spatial abilities.
The Naïve Bayes algorithm is a supervised machine learning approach used in classification tasks such as spam-email filtering. We investigated whether this algorithm could analyze transcripts for spatial language such that it accurately classifies utterances as spatial or non-spatial using R software (2019). The dataset used to train the algorithm included a total sample of 3,913 utterances (1,531 spatial; 2,382 non-spatial) datamined from the CHILDES TalkBank system (Demetras et al., 1986; Dickinson & Tabors, 2001; MacWhinney, 2000). Utterances were dichotomously coded as spatial (1) or nonspatial (0; Cannon et al., 2007) by the first author and two research assistants trained to 80% agreement. Preliminary tests of the algorithm were performed on four publicly available transcripts – each from a different study in which parents and children engaged in a variety of activities such as blocks, puzzles, pretend play, play dough, general conversation, and Piagetian conservation tasks (Bates et al., 1988; Braunwald, 1971; Nicholas & Geers, 1997; VanKleeck, n.d.). The four test transcripts contained a total of 2,229 utterances but varied greatly between the studies (M=557, SD=294, Range=266-938).
Test statistics indicated that on average, the algorithm performed well (Table 1) with .653 Kappa (SD= 0.113, Range=.59-.82), 88% Accuracy (SD= 0.029,Range=.85-.92), 81% Sensitivity (SD= 0.065,Range=.72-.88) and 90% Specificity (SD=0.031,Range=.87-.94). Confusion matrices of false positives and false negatives were produced and further support this interpretation (Figure 1). Between the four transcripts, only 201 of 1,750 (11.5%) non-spatial utterances were incorrectly labeled as spatial and only 96 of 479 (20%) spatial utterances were incorrectly labeled as non-spatial.
These preliminary findings are encouraging as they provide a proof of concept that this supervised algorithm can successfully classify utterances as spatial or non-spatial. Importantly, it accurately classified utterances produced in a wide variety of contexts. Future research will improve the training data and optimize it for more fine-grained spatial language analysis. By April 2021, the training dataset will include over 30,000 coded utterances.
LaTreese Hall, Florida International University
Presenting Author
Melanie Rengel, Florida International University
Non-Presenting Author
Hannah Marie Bowley, Florida International University
Non-Presenting Author
Denise Hernandez, Florida International University
Non-Presenting Author
LaShonica L. Giles, Tuskegee University
Non-Presenting Author
Shannon Marie Pruden, Florida International University
Non-Presenting Author