Santali, an indigenous language, is spoken in diverse regions across India. It challenges speech recognition due to limited resources and linguistic diversity. This paper comprehensively investigates Santali vowel recognition by integrating acoustic phonetics and machine learning (ML) techniques. The primary objective is to explore the recognition of Santali vowels from spoken utterances. To tackle the difficulties arising from the limited resources available for the language, various approaches encompassing feature extraction methods, model architectures, and training strategies are investigated. This study’s acoustic phonetics approach incorporates formant frequencies, short-time energy (STE), and short-time zero crossing rate (ST-ZCR). Meanwhile, the machine learning-based approach involves feature extraction techniques like MFCC, Chroma, and Mel-spectrogram, coupled with various classifiers such as SVM, Random Forest, Gradient Boosting Machines, and XGBoost. The results of this study showcase promising outcomes, underscoring the potential for advancements in Santali speech processing. Additionally, this research contributes to preserving and analysing the under-resourced tribal language, supporting efforts to bridge the gap in Santali language technology and promoting its sustainable development.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Santali Vowel Recognition: An Under-Resourced Tribal Language

  • Sandip Jana,
  • Joyanta Basu,
  • Tapan Kumar Basu,
  • Amiya Karmakar

摘要

Santali, an indigenous language, is spoken in diverse regions across India. It challenges speech recognition due to limited resources and linguistic diversity. This paper comprehensively investigates Santali vowel recognition by integrating acoustic phonetics and machine learning (ML) techniques. The primary objective is to explore the recognition of Santali vowels from spoken utterances. To tackle the difficulties arising from the limited resources available for the language, various approaches encompassing feature extraction methods, model architectures, and training strategies are investigated. This study’s acoustic phonetics approach incorporates formant frequencies, short-time energy (STE), and short-time zero crossing rate (ST-ZCR). Meanwhile, the machine learning-based approach involves feature extraction techniques like MFCC, Chroma, and Mel-spectrogram, coupled with various classifiers such as SVM, Random Forest, Gradient Boosting Machines, and XGBoost. The results of this study showcase promising outcomes, underscoring the potential for advancements in Santali speech processing. Additionally, this research contributes to preserving and analysing the under-resourced tribal language, supporting efforts to bridge the gap in Santali language technology and promoting its sustainable development.