A framework for tongue model visualization based on ultrasound images
摘要
Ultrasound, as the most promising means of obtaining tongue movement information, has increasingly attracted researchers’ attention in the field of speech visualization, particularly in the study of driving three-dimensional tongue models using ultrasound. Currently, most methods that use ultrasound to drive models mimic the EMA point-driven approach by using the motion trajectories of a limited number of contour points to drive the tongue model. However, the points extracted from ultrasound images lack consistency compared to those obtained from EMA, which may lead to pathological shapes that do not correspond to phonetic shapes. To address the issues present in current ultrasound-driven tongue models, this paper proposes a visualization framework that utilizes the limited tongue contour information to drive the tongue model. The entire framework is based on neural networks, achieving full automation from input ultrasound images to output three-dimensional tongue models while avoiding the occurrence of pathological shapes. The framework first uses convolutional neural networks to extract the tongue contour from ultrasound images, then employs fully connected networks to establish the mapping relationship between the entire tongue contour and model control parameters, and finally drives the three-dimensional tongue model using the synthesized model control parameters. To validate the rationality of the driving effect, mean square error and curve similarity were used for error analysis of the model control parameters and to assess the similarity between the ultrasound tongue contour and the model shape. The results show that the average mean square error of the control parameters for all phonemes is within 3%, and the average similarity of the contour curves between the tongue model and the ultrasound tongue contour in the test set is greater than 88%. The results indicate that the method of driving the tongue model using the limited tongue contour is feasible, effectively obtaining a 3D tongue model that matches the 2D ultrasound images, avoiding the issue of pathological shapes in the driven tongue model, while achieving full automation of the ultrasound-driven tongue model.