Speech is a means by which an individual can interact with the society in which he or she is inserted, but its weakness can lead to social exclusion and pathological stigma, so it is necessary to understand the speech process in a systematized and specified way. In this paper, this understanding is proposed through the detection of vocal tract using the YOLO v8 framework in a vocal tract represented by a magnetic res-onance image, to observe the accuracy and performance of the model in order to help the specialist recognize a pattern in the individual’s speech and consequently the maturity and evolution of the applied framework. The model’s results were stable, but in detecting the tongue, the accuracy dropped to 0.72 due to the complexity and variability of the movements of this vocal tract object, while the epiglottis had an average accuracy of 0.48, due to the loss of gradient and low resolution.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Vocal Tract Detection to Aid Speech Pattern Recognition

  • Luã C. S. Saunders,
  • Haroldo G. B. Filho,
  • Clebson M. M. Silva,
  • Daniel L. S. Aires,
  • Idemilson D. A. Silva,
  • Judson R. C. Filho,
  • Leonardo V. S. S. Menez,
  • Luiz F. S. Santos,
  • Nubia V. A. Freitas,
  • Paulo G. S. Gomes

摘要

Speech is a means by which an individual can interact with the society in which he or she is inserted, but its weakness can lead to social exclusion and pathological stigma, so it is necessary to understand the speech process in a systematized and specified way. In this paper, this understanding is proposed through the detection of vocal tract using the YOLO v8 framework in a vocal tract represented by a magnetic res-onance image, to observe the accuracy and performance of the model in order to help the specialist recognize a pattern in the individual’s speech and consequently the maturity and evolution of the applied framework. The model’s results were stable, but in detecting the tongue, the accuracy dropped to 0.72 due to the complexity and variability of the movements of this vocal tract object, while the epiglottis had an average accuracy of 0.48, due to the loss of gradient and low resolution.