Individual Submission Summary
Share...

Direct link:

Poster #68 - A validation of LENA speaker classification in Ghana

Fri, March 24, 11:30am to 12:15pm, Salt Palace Convention Center, Floor: 1, Hall A-B

Abstract

Early verbal interaction with children has been found to play a pivotal role in early language development. Recent technological advances have made conducting studies of child-directed speech in the home easier than ever. LENA is a combined recording device and analysis software that can automatically identify and quantify speech across daylong audio samples (Gilkerson et al., 2017). These estimates have been found to correlate with human transcription across multiple languages (e.g., Chinese: Gilkerson et al., 2015; French: Canault et al., 2016; Hebrew/Arabic: Levin-Asher, Segal, & Kishon-Rabin, 2022). However, while LENA has been used in African communities to study language development (see Weber, Fernald, & Diop, 2017), we are not aware of any validation studies focused on these regions. There are also some methodological issues raised by previous validation studies. Recent work has found that LENA performs less favorably when audio is sampled randomly (rather than selected from speech dense segments) or when overlapping speakers are included (Cristia et al., 2021). In our study, we attempt to address these concerns by examining speech throughout the day in Ghana.

We collected 52 hours of audio from a small sample of 4 children (15.5 to 41mos) living in Accra, Ghana. The main language of all the children’s households was Twi. Each child’s primary caregiver was asked to place the device in the child’s shirt pocket at the beginning of the day and to leave it in until the next day. We randomly selected 2 minutes of audio from every hour of recording and broke each sample down into 10ms frames. We used the ACLEW coding scheme (Bergelson et al., 2017) to classify each frame by speaker identity (e.g, Female Adult, Target Child), marking multiple speakers as Overlap. We then conducted a comparison of our hand-coded samples to the automatic speaker diarization produced by the LENA software. For each speaker identity category, we produced a measure of Recall (i.e., what percentage of speech coded by the human annotator did the LENA find) and Precision (i.e., what percentage of speech did the LENA correctly classify).

We found relatively low rates of Recall across all speakers (28% of Female Adult; 26% of Male Adult; 31% of Other Child and 33% of Target Child), suggesting LENA failed to detect large amounts of speech (Figure 1). We found higher rates of Precision for Female Adult (45%) and Target Child (56%), but not for Male Adult (32%) or Non-Target Children (13%) (Figure 2). We then compared our results to those found by another group (Cristia et al., 2021) who conducted a similar evaluation of LENA speech classification using recordings drawn from middle-class families in the U.S. and the U.K., as well as recordings drawn from Tsimane families, a foraging society in Bolivia. We found comparable rates of Recall in the two evaluations, but lower Precision overall in the Ghanaian sample, especially for Female Adult. Our results suggest that LENA accuracy may be lower in noisier settings, but it may capture comparable amounts of speech to other similar studies.

Authors