Individual Submission Summary
Share...

Direct link:

Active viewing in infants facilitates visual object learning: A computational modeling approach

Fri, April 9, 11:35am to 1:05pm EDT (11:35am to 1:05pm EDT), Virtual

Abstract

Early object recognition in a visual world full of clutter is a complicated task at which infants are incredibly efficient. By two years old, infants recognize roughly 300 object categories and can generalize a newly learned label to novel visual instances. Our hypothesis on how infants become efficient learners of visual objects is that they actively create many views of the same object by actively manipulating objects in everyday contexts, and this active viewing creates high-quality data for visual recognition. We tested this idea by extracting and quantifying visual properties of object instances perceived by infant learners in toy play and using state-of-the-art machine learning models (Convolutional Neural Networks, CNNs) to simulate how the visual information is processed.

Method. Twenty-six infant-parent (M=19.4mo, SD=2.2mo) dyads participated in our study. Each dyad was brought into a room with a set of 24 toys on the floor. The infants and parents were instructed to play with the toys, without other instructions. To approximate visual data perceived and processed by young learners, we used head-mounted eye trackers to record egocentric video (Figure 1, left) of both the infant’s and parent’s views and detected which visual objects were in view (Figure 2, middle). With gaze data from eye trackers, we also detected which objects were visually attended by infants and parents moment by moment. Figure 1 (right) shows a collage of visual instances of one particular toy (football helmet) attended by a young learner in a toy play session.

Results. As shown in Figure 2A, we computed pairwise distances across all instances within each object and projected them into a 2D space using multidimensional scaling (MDS). Figure 2B shows that object instances in the infant’s view contain a mix of similar and dissimilar instances while most instances in the parent’s view are similar to each other. We trained the CNN with either visual instances from the infant view or visual instances from the parent view, and found the same model trained on the infant data (Minfant=60.23%, SDinfant=2.63%) outperformed the model trained based on the parent data (Minfant=41.72%, SDinfant=1.85%). To understand how manual actions from both participants create visual data for infant learning, we divided the infant data into three subsets: 1) a child-held subset containing the object instances held by infants; 2) a parent-held subset containing the instances held by parents; and 3) a no-holding subset containing the object instances not held by anyone. Training the same model using these three subsets, we found that the model trained by both infant-held and parent-held sets outperformed the model trained by the no-holding set (Minfant-held=44.45%, SDinfant-held=3.28%; Mparent-held=46.82%, SDparent-held=2.17%; Mno-holding=32.81%, SDno-holding=1.26%).

Conclusions. The relevant data for learning are not the statistics of the physical and social world but only the samples that emerge within the learners’ own experiences. The sampling process is implemented through both infants’ and parents’ actions in everyday social interactions, such as free-flowing toy play. Active viewing and exploration of infants in such naturalistic contexts create high-quality training data for object recognition.

Authors