Search
Browse By Day
Browse By Time
Browse By Panel
Browse By Session Type
Browse By Topic Area
Browse Posters
Search Tips
Register for SRCD23
Personal Schedule
Welcome Letter
Program Guide
Change Preferences / Time Zone
Sign In
In preschool classrooms, children’s vocalizations are associated with children’s assessed language abilities, an association that extends to children with disabilities such as hearing loss. (Fasano, et al., 2021; Mitsven, et al., 2022; Perry et al., 2022). Investigators have defined vocalizations within social contact based on children’s patterns of co-location and orientation with peers and teachers. However, the factors that predict children’s classroom vocalization in real-time are essentially unknown. Machine learning models based on deep learning (DL) appear to be an effective tool to model complex spatiotemporal data (Wang, 2020). In this study, we investigate a DL model to predict the amounts of vocalization of the children in preschool inclusion classrooms based on their own history of vocalizations in response to the location, orientation, and vocalizations of their local partners.
We observed 63 preschoolers in two oral language inclusion classrooms for children with hearing loss (HL) who use hearing aids or cochlear implants and their typically hearing (TH) peers (Table 1). In each physical classroom, one yearly cohort was used to train a model for preschooler speech prediction, and the successive cohort was used to test the model. Data came from automated measurements of children and teachers in both cohorts of both classrooms over a total of 36 multi-hour observations (Table 1). For all children and teachers (2-3 per classroom), an ultrawideband radio frequency identification tag worn over the left and right hips tracked position and orientation (Messinger et al., 2019). LENA audio recorders worn by all children and teachers were diarized with LENA algorithms (Dykstra et al., 2013; Perry et al., 2018).
To model the relationship between vocalization and movement of the target child, peers, and teachers, we apply deep learning where a 3-minute history of egocentric vocalization and position/orientation data are used as inputs to predict the amounts of vocalizations of the target child in successive one-minute intervals. A key feature of the model is that all data are represented egocentrically around the target child, which emphasizes the relative distance and interpersonal orientations of the partners to the target child. The model is trained on observed data in which the constellations of interpersonal distance/orientation features and partners’ vocalizations occurred with or without the target child vocalizing. Thus, the model learns whether the target child vocalized given the position, orientation, and vocalization of the partners surrounding them.
The mean average error (MAE) of the predicted vocalization duration per minute was 0.282 (observed 2.332) and 0.481 seconds (observed 2.549), respectively. The relative prediction error (RPE) (the MAE divided by the observed duration) was 11.06% and 20.97%.
The results demonstrate the predictive power of predicting child vocalizations from the child’s egocentric perspective in which they are the center of the interaction. Child speech is both an index and a vehicle of language development. Interpretive ML is currently in the process of better identifying those features of classroom interaction that best predict children’s speech.