Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
This presentation will include an overview of speech data collection and speech processing procedures and methods, addressing the question, Q5. What technologies are available to design and evaluate CRI? It will discuss advances and issues surrounding automated speech recognition, speaker diarization (trying to distinguish between multiple speakers in one audio stream), and microphone options.
Speech-related technologies that most consumers are familiar with (year and making their way into more and more household devices (e.g., Alexa, Siri, Hey Google, etc.) involve many different technologies stacked together to produce a service, like a personalized assistant. Developing interactions and rules for a robot to listen and/or talk to a child involves many of these same technologies, the most important of which is probably automated speech recognition (ASR). While algorithms and advances in ASR for adult speech have progressed significantly over the past decade, these advances have not translated to successful ASR for younger speakers (Russell, 2007; Gerosa 2009). Young children have vocal tracts that are significantly smaller than those of adults, causing speech to occur at higher frequencies and children’s speech contains much higher variability than adult speech (Russell, 2007; Gerosa, 2009). Due to the dramatic differences in their speech, when recognition systems trained on adult speech are applied to speech from younger children, error rates can easily double or triple (Russell, 2007).
ASR is an important part of many systems that are attempting to determine how well children can read or trying to engage in conversations, but there are other types of analyses from speech data that are relevant for other kinds of applications. Speech contains more information than just what words might have been said. Speech features such as prosody (tonality), rate of speech, loudness contours, and intonation patterns (e.g., questioning vs. statement), as well as more complex utterance styles such as sarcasm, reflect more than just propositional meaning. Also, patterns of interaction between or among speakers—such as who speaks more, latencies between conversational turns, disfluencies and discourse markers, or how often individual speakers overlap with each other or negotiate who gets the floor—provide rich information about individual speakers and about groups as a whole (e.g., Ford & Couper-Kuhlen, 2004; Sacks, Schegloff, & Jefferson, 1974). Many of these types of speech analyses are more successful with younger children since they do not require ASR techniques, however they are still not up to the same accuracy levels as adult speakers.
The presentation will also discuss different projects that have made advances in conversational analysis or tools that use speech processing and related technologies. There are many different applications for speech technologies within conversational agents. These systems are text-based, like DialPort and SlugBot that can help answer specific questions. Nonetheless, interactions between humans and robots is one of the main areas where much research is focused right now.