Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Search Tips
Register for SRCD21
Personal Schedule
Change Preferences / Time Zone
Sign In
X (Twitter)
Young children have wide-ranging, sophisticated knowledge about the world around them, yet the origin of early knowledge is often unclear. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and deep learning.
In this talk, I will discuss recent progress on these questions using machine learning applied to baby headcam videos. We applied contemporary self-supervised learning algorithms to videos recorded from the perspective of a young child (Sullivan et al., 2020). The algorithms aim to learn visual features that are invariant in the video frames across time, without accessing any category labels during training. Thus, the algorithms learn only from a strict subset of one baby's experience without labels (~1 week of video spread over 6 to 32 months).
We find that useful, high-level visual representations emerge from self-supervised learning through the eyes of one child. To analyze the trained model, we examine whether baby-appropriate visual categories can be discriminated using only a linear classifier trained on top of the fixed, self-supervised features. The classifiers perform well on a range of tasks, in some cases outperforming strong models trained with supervised learning on large sets of labeled images (ImageNet). The classifiers can also confidently discriminate categories such as car vs. road, door vs. window, foot vs. hand, etc. from a model trained through the eyes of one baby. Attention maps reveal that the networks often look in the right places when making classification decisions (see figure).
This work shows that high-level visual representations are learnable from a small subset of actual experience from one baby. I will discuss open questions regarding how far we can push generic learning mechanisms, through even larger and richer data sources, to learn mental representations ever closer to those acquired by children early in their development.