错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Two-Stage GAN for Field-of-View Standardisation and Tongue Region Enhancement in Ultrasound for Cleft Palate Speech Pattern Analysis

  • Saja Al Ani,
  • Joanne Cleland,
  • Ahmed Zoha

摘要

Ultrasound tongue imaging (UTI) offers a non-invasive, real-time view of tongue motion, but variability in the field of view (FoV) and distracting background anatomy hinder automated analysis. We present a two-stage generative pipeline based on a Pix2Pix conditional generative adversarial network (cGAN) to standardise and refine UTI frames. Stage 1 harmonises FoV by converting wide-angle frames to a uniform 97° view, and Stage 2 enhances the region of interest (ROI) by isolating tongue structures. Stage 1 achieved near-lossless fidelity (SSIM = 0.96, PSNR = 88.5 dB), while Stage 2 preserved structure (SSIM = 0.91, PSNR = 33.2 dB). In binary classification of cleft palate ± lip versus typically developing (TD) children, FoV standardisation reached 98.8% accuracy with mixed real–synthetic training, and ROI refinement increased recall to 97.8%. Cross-domain tests showed Stage 1 outputs were indistinguishable from real data, and mixed training mitigated Stage 2 shifts. The framework enhances data consistency, interpretability, and robustness, demonstrating the potential of generative AI for clinical UTI analysis.