Automatic speech recognition (ASR), a speech technology also referred to as Speech-to-Text, converts an acoustic signal into textual sequences of words using probability estimates based on language and acoustic modelling. Although programmes differ in accuracy, research has shown that various ASR systems, including Google’s Voice Typing, reach similar levels of accuracy and correlate with human listeners’ transcription accuracy. Researchers have looked into ASR programmes for supporting language learning across a range of skills, but the most promising arena has been pronunciation. Research has shown that both ASR-dictation practice and ASR-integrated Computer-Assisted Pronunciation Training (CAPT) practice can support improvements in segmental (vowel and consonant) accuracy. Although CAPT programmes, which offer explicit training and feedback on mis-pronunciations, encourage greater learner improvement than indirect feedback provided by ASR-dictation, CAPT programmes often control learner utterances in order to compare the learners’ pronunciation to an expected model, limiting creation of meaningful utterances. Yet, the future of ASR CAPT is promising as researchers explore ways to integrate ASR into more communicative applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Speech Recognition

  • Shannon McCrocklin

摘要

Automatic speech recognition (ASR), a speech technology also referred to as Speech-to-Text, converts an acoustic signal into textual sequences of words using probability estimates based on language and acoustic modelling. Although programmes differ in accuracy, research has shown that various ASR systems, including Google’s Voice Typing, reach similar levels of accuracy and correlate with human listeners’ transcription accuracy. Researchers have looked into ASR programmes for supporting language learning across a range of skills, but the most promising arena has been pronunciation. Research has shown that both ASR-dictation practice and ASR-integrated Computer-Assisted Pronunciation Training (CAPT) practice can support improvements in segmental (vowel and consonant) accuracy. Although CAPT programmes, which offer explicit training and feedback on mis-pronunciations, encourage greater learner improvement than indirect feedback provided by ASR-dictation, CAPT programmes often control learner utterances in order to compare the learners’ pronunciation to an expected model, limiting creation of meaningful utterances. Yet, the future of ASR CAPT is promising as researchers explore ways to integrate ASR into more communicative applications.