This study presents an innovative method to determine tongue movement during speech. This is accomplished by using ultrasound tongue imaging (UTI) from the TaL80 dataset. Our methodology uses advanced deep learning and machine learning techniques. First, a feature extractor is built by using transfer learning using ResNet50 to extract features. Moreover, we use the Extra Tree Classifier to select features effectively. Finally, we use Support Vector Machines (SVMs) to perform classification tasks. Our experiments indicate the effectiveness of the recommended technique in estimating tongue movements with a high accuracy of 100%, and also the validity of our results is verified by employing evaluation measures such as F1-score, recall, and precision. These measures achieved a 100% success rate by examining multiple cases, including one without feature selection and SMOTE technique, and another with feature selection and SMOTE technique. This allowed us to select the most optimal solution. The proposed technique shows great potential for use in speech therapy, language learning, assistive communication devices, and other related domains, opening up new possibilities for future research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Estimating Tongue Movements from Speech Using Deep Learning and Machine Learning Techniques

  • Safa Emad Sabri,
  • Jamal Mustafa Al-Tuwaijari

摘要

This study presents an innovative method to determine tongue movement during speech. This is accomplished by using ultrasound tongue imaging (UTI) from the TaL80 dataset. Our methodology uses advanced deep learning and machine learning techniques. First, a feature extractor is built by using transfer learning using ResNet50 to extract features. Moreover, we use the Extra Tree Classifier to select features effectively. Finally, we use Support Vector Machines (SVMs) to perform classification tasks. Our experiments indicate the effectiveness of the recommended technique in estimating tongue movements with a high accuracy of 100%, and also the validity of our results is verified by employing evaluation measures such as F1-score, recall, and precision. These measures achieved a 100% success rate by examining multiple cases, including one without feature selection and SMOTE technique, and another with feature selection and SMOTE technique. This allowed us to select the most optimal solution. The proposed technique shows great potential for use in speech therapy, language learning, assistive communication devices, and other related domains, opening up new possibilities for future research.