<b>Purpose</b> <p>Surgical simulation offers a promising addition to conventional surgical training. However, available simulation tools lack photorealism and rely on hard-coded behaviour. Denoising Diffusion Models are a promising alternative for high-fidelity image synthesis, but existing state-of-the-art conditioning methods fall short in providing precise control or interactivity over the generated scenes.</p> <b>Methods</b> <p>We introduce SurGrID, a Scene Graph to Image Diffusion Model, allowing for controllable surgical scene synthesis by leveraging Scene Graphs. These graphs encode a surgical scene’s components’ spatial and semantic information, which are then translated into an intermediate representation using our novel pre-training step that explicitly captures local and global information.</p> <b>Results</b> <p>Our proposed method improves the fidelity of generated images and their coherence with the graph input over the state of the art. Further, we demonstrate the simulation’s realism and controllability in a user assessment study involving clinical experts.</p> <b>Conclusion</b> <p>Scene Graphs can be effectively used for precise and interactive conditioning of Denoising Diffusion Models for simulating surgical scenes, enabling high-fidelity and interactive control over the generated content.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SurGrID: controllable surgical simulation via Scene Graph to Image Diffusion

  • Yannik Frisch,
  • Ssharvien Kumar Sivakumar,
  • Çağhan Köksal,
  • Elsa Böhm,
  • Felix Wagner,
  • Adrian Gericke,
  • Ghazal Ghazaei,
  • Anirban Mukhopadhyay

摘要

Purpose

Surgical simulation offers a promising addition to conventional surgical training. However, available simulation tools lack photorealism and rely on hard-coded behaviour. Denoising Diffusion Models are a promising alternative for high-fidelity image synthesis, but existing state-of-the-art conditioning methods fall short in providing precise control or interactivity over the generated scenes.

Methods

We introduce SurGrID, a Scene Graph to Image Diffusion Model, allowing for controllable surgical scene synthesis by leveraging Scene Graphs. These graphs encode a surgical scene’s components’ spatial and semantic information, which are then translated into an intermediate representation using our novel pre-training step that explicitly captures local and global information.

Results

Our proposed method improves the fidelity of generated images and their coherence with the graph input over the state of the art. Further, we demonstrate the simulation’s realism and controllability in a user assessment study involving clinical experts.

Conclusion

Scene Graphs can be effectively used for precise and interactive conditioning of Denoising Diffusion Models for simulating surgical scenes, enabling high-fidelity and interactive control over the generated content.