Application of the fast diffusion model ST-diffusion in multi-modal lip-to-speech generation
摘要
With the continuous advancement in lip-to-speech generation technology, diffusion-based models have become a leading methodology for synthesizing speech from silent talking-face videos by extracting and aligning orofacial movements. This capability provides practical solutions for assistive technologies and audio restoration in cases where the original speech signal is degraded or entirely absent. However, most current diffusion-based approaches exhibit two major shortcomings: first, they frequently lack optimization for computational efficiency, leading to prolonged training convergence; second, their multi-step denoising process does not incorporate a recursive feedback mechanism, failing to exploit previous intermediate outputs as conditional cues for subsequent generation. This limitation restricts the potential for further enhancing perceptual and temporal accuracy. To overcome these challenges, we propose ST-diffusion (Self-condition and token drop diffusion), a novel framework that synergistically integrates two complementary strategies to improve the performance of diffusion models. The first is a self-conditioning strategy that recycles intermediate results from earlier denoising steps to inform and guide the current prediction, thereby improving the intelligibility and naturalness of the synthesized speech. The second is a token drop strategy, which exploits the intrinsic redundancy of audio representations. During training, this method randomly masks a subset of input tokens, encouraging the model to learn more robust and essential features. This not only maintains semantic consistency but also significantly accelerates model convergence. Experiments on the GRID dataset, recorded under controlled conditions, and the unconstrained LRS2 benchmark show that ST-Diffusion achieves competitive results: the word error rate (WER) drops to 16.5% on LRS2 and 1.74% on GRID, surpassing several existing lip-to-speech systems. In addition, the proposed framework attains approximately twice the training convergence speed of conventional diffusion-based baselines. These findings confirm that ST-diffusion achieves an effective balance among speech quality, lip-speech synchronization accuracy, and training efficiency.