Individual Submission Summary
Share...

Direct link:

Asking to the crowd: large-scale validation of Zooniverse labels on the complexity of child vocalizations

Wed, April 7, 2:45 to 4:15pm EDT (2:45 to 4:15pm EDT), Virtual

Abstract

Childhood language disorders carry important costs for both the individual and the society, and early diagnosis and intervention could greatly reduce these costs. Recent work has attempted to find early markers of language disorders from day-long audio-recordings gathered with wearables (e.g., LENATM recorders, [1]), but automatic analyses are errorful ([2]) and human annotation is laborious and expensive. This study's objective was to investigate whether we can rely on collaboration with citizen scientists through the Zooniverse platform ([3]) to accurately quantify the maturity of child vocalizations, such as the production of canonical syllables (i.e., consonant-vowel or vowel-consonant sequences). We hypothesized that indices derived from citizen scientists' classifications would correlate with those based on laboratory annotations.

Methods. Our dataset included speech samples LENA identified as having been produced by the key child. A total of 2,000 samples (100 per child) were drawn from daylong recordings of 10 children (6 males, 4 females; age range 11-53 months, mean=41.5 months) diagnosed with Angelman syndrome (AS), a neurogenetic syndrome characterized by severe language impairments, and 10 low-risk (LR) control children (6 males, 4 females; age range 4-18 months, mean=11.7 months). These speech samples were annotated by highly trained (“gold standard”) annotators in the laboratory and by citizen scientists on Zooniverse. Annotators assigned each clip to one of five categories: Canonical, Non-Canonical, Crying, Laughing, and Junk.
Feasibility: All data have been collected and coded thus will be available for presentation.

Analyses and Results. First, we examined the correspondence between citizen scientist and gold standard annotations using a confusion matrix. We found that performance was moderate to good, with 65% accuracy (95% CI .64-.66), and a kappa Gwet’s A1C coefficient of .58 (95% CI .57,.59). The Non-Canonical class showed the highest recall (79%), and it was also the most common class in the dataset. The other classes had fewer clips assigned to them, but the Laughing and Canonical categories were well detected (Recall 62% and 58% respectively), with other classes having more variable assignments (49% for Crying and 39% for Junk).
Next, we analyzed descriptors at the level of individual children, to investigate whether citizen scientist labels are able to capture individual variation. We found strong positive correlations in two derived measures that have been found to relate to age and/or diagnosis. The correlation across methods was r=.83 for linguistic proportion ([LP=Canonical+Non-Canonical)/All vocalizations], and r=.93 for canonical proportion [CP = Canonical/(Canonical+Non-Canonical)], with small differences across groups (LPLR r=.82; LPAS r=.87; CPLR r=.96; CPAS r=.85).

Discussion: Our results demonstrate that indices derived from classifications obtained through a citizen science platform correlate highly with those derived from the more time- and cost-intensive laboratory annotations. Our findings are particularly exciting in an era when wearables open novel avenues for research in language development. With further testing, the use of crowd-sourcing may pave the way to the creation of large, high-quality datasets that will drive the future of developmental research in new and ecological ways.

Authors