Individual Submission Summary
Share...

Direct link:

Modeling Early Word Learning from the Infant’s Point of View

Thu, April 8, 10:00 to 11:30am EDT (10:00 to 11:30am EDT), Virtual

Abstract

Computational modeling approaches have been used in developmental science to reveal the underlying learning processes that children may employ to extract statistical regularities from the sea of data. Different from previous modeling studies, the present work used egocentric video and gaze data collected from infant learners during naturalistic toy play to examine how they learn the associations between object names and visual objects from inherently ambiguous experiences. Our approach captured the learning environment from the learner’s own point of view to closely approximate the learning input perceived by young children.

Thirty-four child-parent dyads participated in our study. The children ranged in age from 15.2 to 24.2 months (M=19.4 months, SD=2.2 months). Each dyad was told to play with a set of 24 toys scattered on the floor, without more specific instructions. We collected egocentric video and eye tracking data during free-flowing joint play. The final dataset contains 212 minutes of video, with each dyad contributing different amounts of data ranging from 3.4 minutes to 11.6 minutes. The final training dataset consists of all the naming instances in parent-child joint play, with each instance containing a target name and a set of 90 image frames from the child’s first-person camera that co-occur with the naming utterance. As shown in Figure 1, each image typically contains multiple visual objects and the named object may or may not be in view.

We used a state-of-the-art Convolutional Neural Networks (CNN) model, ResNet50, trained with stochastic gradient descent. In Study 1, we examined the effects of different attentional strategies by assigning each naming event into one of two categories: sustained attention if the infant attended to a single object for more than 60% of the frames in the naming event, and distributed attention otherwise. The results reveal that the model trained with sustained attention events (Msustained=30:53%; SEsustained=2:08%) outperformed the model trained with distributed attention events (Mdistributed=23:26%; SEdistributed=1:78%; β=0:20; t=2:65; p<0:005), suggesting that sustained attention on a single object while hearing a name leads to better learning. In Study 2, we examined the effects of visual properties of attended objects. The naming events were grouped into two subsets by a median split of object size. The large subset contains naming instances in which named objects are larger than the median size (6%) whereas the small subset contains naming instances in which named objects are smaller than the median. The same model was separately trained on the large set and the small set. We found the model trained with large objects achieved significantly higher accuracy than that trained with small objects (Mlarge=30:50%; SElarge=2:20%; Msmall=18:81%; SEsmall=1:81%; β=0:29; t=4:12; p<0:001).

Our findings provide a sensory account of the role of sustained attention in early word learning. More generally, our work highlights the importance of moving toward examining realistic data in developmental research because learning outcomes depend on not only internal learning processes but also the data on which those processes operate.

Authors