Design and Implementation of English Multimedia Corpus
摘要
In order to expand the resources of speech corpora and enhance the construction of working methods of speech and communication, a design and implementation method for English multimedia corpus is proposed. We designed the module of the voice acquisition system according to actual needs, using technology-based development tools and platforms that use databases and patterns. We designed each module in detail, mainly including analyzing data objects, structures, and implementation of access schemes, as well as completing the design of the voice database. Continuously improve, modify, and test the program during the development process, and finally conduct partial recording testing. After using the original corpus downloaded from the Internet and performing preliminary processing on the original corpus, an algorithm based on high-frequency word lists and three tone sub lists is used to select the original corpus. 2335 out of 2500 commonly used words and 3683 out of the highest frequency of 4000 words have coverage rates of 93.40 and 92.08% for commonly used words, respectively. For commonly used words that cannot be covered by the corpus, we will automatically generate corpus texts for these words and store them in our corpus form.