XR-Diffusion: A Robust Framework for High-Quality Chest X-ray Synthesis
摘要
To address the limitations of traditional diffusion models in generating high-quality, diverse, and realistic medical images, we propose the XR-Diffusion model. This advanced latent diffusion model integrates the Med-BERT text encoder with a visual-language pretraining (VLP) filtering strategy, including classification and image-text retrieval filters, ensuring precise text alignment. By reducing hallucinations, the dual-filtering mechanism significantly enhances image fidelity, diversity, and clinical reliability. XR-Diffusion is ideal for radiology education, AI model training with limited annotated data, and clinical decision support. Evaluations on the MIMIC-CXR and CheXpert datasets show the model achieves FID scores of 12.17 and 17.15, and MS-SSIM scores of 0.12 ± 0.07 and 0.18 ± 0.09, demonstrating exceptional performance.