Individual Submission Summary
Share...

Direct link:

Computational Account for the Development of Top-Down Visual Attention

Sat, March 23, 12:45 to 2:15pm, Hilton Baltimore, Floor: Level 1, Peale A

Integrative Statement

This study investigates infant attention during a Solo context of face scanning using computational modeling and, as Paper 3, pursues the development of infant attention. Visual attention is controlled by two types of processing: bottom-up processing determines where to attend based on sensory stimuli, while top-down processing is driven by intention and/or the person’s prior knowledge. Although previous studies revealed that younger infants rely mainly on bottom-up attention, how they develop top-down attention remains a mystery.
We hypothesize that top-down attention is acquired through the maximization of information gain. In order for infants to recognize visual input, they have to optimize their attention so as to gain more important information from the input. We employ a recurrent neural network to verify our hypothesis. The network shown in Figure 1 (modified from (Mnih et al., 2014)) receives an image and the attention location at the previous time step, and learns to decide where to attend at the next time step. The network also learns to recognize the input image and updates the connecting weights based on a reward given to the recognition task. A positive reward “1” is given for a correct recognition while “0” is given for an incorrect recognition. Owing to the recurrent connection in the hidden layer, the network can learn to maximize the information gain based on the history of attended information.
We verified our hypothesis in a scenario of face scanning. A face image conveys multiple information such as the identity of a person, the emotion expressed by the person, the direction of the person’s attention, etc. We chose two recognition tasks: estimation of emotion and estimation of head orientation, and examined how they guide the development of attention. It is known that when observing a person’s face, scan paths of young infants are scattered over a stimulus, whereas older infants repeatedly gaze at the stimulus’ eyes and mouth (Kato and Konishi, 2013). We compared visual attention generated by the network with that observed in infants.
Our experiments demonstrated that the recurrent neural network could reproduce infant-like visual attention when it was trained for both emotion and head orientation estimation. Figure 2(a) depicts the developmental change of attention location (top) and attention shift (bottom) over learning. The attention of the network gradually converged to the eyes and mouth regions as learning progressed. Both the larger circles corresponding to the eyes and mouth areas, and the thicker lines connecting these areas indicate more frequent attention to these regions. In contrast, when the network was trained only for emotion or head orientation estimation, it resulted in different behaviors (Figure 2(b)). The network exhibited stronger attention only to the mouth when it was trained for emotion estimation (top), whereas it exhibited more frequent attention shift between the eyes when it was trained for estimation of head orientation (bottom). These results suggest that the two recognition tasks play complementary roles in guiding attention and that a recurrent neural network can reproduce infant-like development only when complementally tasks are given.

Author