Automated Transcription of Unfamiliar Words to Improve Recognition of Terms in the Subject Domain
摘要
One of the popular methods of improving the recognition quality of words unfamiliar to the model (out-of-vocabulary) is the vocabulary expansion method. Out-of-vocabulary words are particularly common in subject domains. These are terms and term-like collocations. Many of them in Russian speech have recently been borrowed from English and do not yet have an accepted spelling in Russian. This hampers adding such words to the model’s dictionary for non-multilingual models reducing the quality of their audio performance in subject domains. This paper presents a method for improving the quality of recognition of speech containing such terms based on an algorithm of automatic construction of so-called “Russian transcriptions” for arbitrary English words. Russian transcription means a similar-sounding word written in Russian letters, which could replace the original term when added to the model’s vocabulary and thus be recognized correctly. The results of the experiments that we conducted suggest that the algorithm described here can be used to improve the quality of speech recognition systems.