Introduction
摘要
Spontaneous speech, characterized by its fluidity and adaptability, is undoubtedly the most natural form of human communication. It mirrors the complexities of human thought, emotions, and social dynamics, enabling people to express themselves authentically. However, the realm of Conversational AI often leans towards prescriptive language due to technological limitations. Models are trained on written language corpora, widening the gap between natural and spoken language processing. Most recent advances in Conversational AI aim at the facilitation of human-machine interactions, such as issuing voice commands to phones, cars, and home assistants, or interacting with voicebots. Truly spontaneous human-to-human speech remains a challenge for current systems. In this chapter, we introduce the problems associated with the processing of spontaneous human speech. We present an outline of the history of Conversational AI and advocate for the emergence of a new moniker: spoken language processing, a holistic term to describe the set of technologies that allow machines to consume and act upon input expressed as spontaneous human speech.