错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a Statistical Model for Automated Ground Truth Generation in Low-Resource Languages

  • Sanchali Das

摘要

We presents a computational model that can develop a ground truth for audio classification for any new and poorly resourced language. Here, we used one of the fewer resource languages of northeast India, which is Kokborok (Tribal) music. We have created a significant ground truth set of 259 songs out of 316 random audio files by combining the process of machine learning techniques and statistical analysis. The experimental results found by computational modelling of a classification system show that the proposed approach can significantly improve the database as compared to traditional methods. Traditionally the ground truth set of any database is created by annotators only to annotate those audio files. Our approach is statistically justified and can produce the ground truth set for any new poorly resourced language where limited audio data and very ill-conditioned data are present. This dataset is used to develop an music recommendation system for any new language with the maximum accuracy achieved 63%.