Individual Submission Summary
Share...

Direct link:

Acoustic Properties of Infant Directed Speech are Similar in Video Chat Versus In-Person Interactions

Fri, April 9, 3:15 to 4:15pm EDT (3:15 to 4:15pm EDT), Virtual

Abstract

Infant-directed speech (IDS) differs from adult-directed speech (ADS) in acoustic properties of high and variable frequency (mean-F0 and F0 range, respectively), short utterances, and vowel alterations, which collectively contribute to the prosody of speech (Fernald et al., 1989; Kitamura et al., 2001). Salient acoustic characteristics of IDS have been linked to processes important for language acquisition, such as speech segmentation (Golinkoff et al., 2015). Additionally, IDS prosody plays a role in communication of caregivers’ positive emotions and regulation of infants’ emotional states during the first year (Fernald, 1992; Bryant & Barrett, 2007). For example, infants respond with greater positive affect to IDS than ADS (Werker and McLeod, 1989) and infants’ preference for the emotional content of IDS differs across development (Kitamura el al., 2009), demonstrating caregiver-infant emotional communication in IDS.

In this age of technology, heightened by the COVID-19 pandemic, family members often communicate to infants over video chat. Research has not investigated whether IDS prosody varies significantly in face-to-face versus virtual interactions. The current study examines the differences in acoustic properties of IDS over Skype versus in-person play. Specifically, we analyzed the mean-F0 and F0 range (minimum-F0 subtracted from maximum-F0) of 29 mothers in three conditions: (1) a 3-min free-play interaction with their 5-month-old infant, (2) a Skype session with their infant, and (3) a Skype conversation with an adult researcher. In Skype situations mothers were instructed to discuss their weekly schedule.

Videos of mother-infant free-play and Skype interactions were converted into WAV files using VLC Media Player. A 10-s warm-up period was given at the beginning of each recording to allow mothers to adjust to the lab environment. The next 10-s of uninterrupted vocalizations from the mother were analyzed in Praat version 6.0.48 according to the criteria established in Broesch and Bryant (2013).

The mean-F0 and F0 range that typically distinguish IDS from ADS were consistent with those previously reported (Fernald, et al., 1989). Separate within-subject Anovas comparing mean-F0 and F0 range across the three situations resulted in a significant main effect for mean-F0 (F(1.38,38.86) = 25.70, p<.001, Greenhouse-Geisser correction) and for F0 range (F(2,56)=7.14, p=.002). Post hoc pairwise comparisons revealed significant differences for mean-F0 between free-play (M=239.87, SD=30.09) and ADS Skype (M=200.40, SD=13.76) (p<.001) and between IDS Skype (M=248.98, SD=29.99) and ADS Skype (p<.001), but not between free-play and IDS Skype (p=1.0). The same group differences were found for F0 range; free-play (M=345.76, SD=74.78) and IDS Skype (M=357.48, SD=69.84) did not differ (p=1.0), while both differed from ADS Skype (M=301.79, SD=66.92) (free-play vs. ADS p=.027, IDS Skype vs. ADS p=.001). These results suggest that adults’ speech to infants over video chat conveys some of the salient prosodic properties of IDS that communicate and regulate emotion. They also extend knowledge of infant word learning from video chat interactions to the prosodic properties of the speech itself (Roseberry et al., 2014). Future directions should investigate whether other acoustic properties of IDS, such as slowed speaking rate and vowel hyper-articulation, are present in video chat interactions.

Authors