<p>We present AuDiffusion, a diffusion framework that introduces a multi-agent design to improve controllability, semantic alignment, and efficiency in text-to-image generation. The system comprises three cooperating agents responsible for enriching textual input, selecting suitable structural constraints, and performing image synthesis with an enhanced diffusion backbone. This modular design provides more explicit structural control and adaptive decision-making compared with conventional monolithic pipelines, while retaining strong global context modeling. We evaluate AuDiffusion on the ImageNet 256 <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 256 and 512 <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 512 benchmarks. At 256 <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 256 resolution, AuDiffusion achieves a FID of 2.21, IS of 274.12, Precision of 0.85, and Recall of 0.59; at 512 <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> 512, it attains a FID of 3.02 and IS of 268.31 while remaining computationally efficient. These results indicate that a multi-agent diffusion framework can improve controllability and image quality without incurring prohibitive computational overhead, making AuDiffusion a practical candidate for applications such as visual prototyping in creative workflows.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AuDiffusion: multi-agent controlled text-to-image generation with attention-enhanced mamba blocks

  • Dezhi An,
  • Wanyao Zhang,
  • Shengcai Zhang,
  • Jun Lu

摘要

We present AuDiffusion, a diffusion framework that introduces a multi-agent design to improve controllability, semantic alignment, and efficiency in text-to-image generation. The system comprises three cooperating agents responsible for enriching textual input, selecting suitable structural constraints, and performing image synthesis with an enhanced diffusion backbone. This modular design provides more explicit structural control and adaptive decision-making compared with conventional monolithic pipelines, while retaining strong global context modeling. We evaluate AuDiffusion on the ImageNet 256 \(\times \) × 256 and 512 \(\times \) × 512 benchmarks. At 256 \(\times \) × 256 resolution, AuDiffusion achieves a FID of 2.21, IS of 274.12, Precision of 0.85, and Recall of 0.59; at 512 \(\times \) × 512, it attains a FID of 3.02 and IS of 268.31 while remaining computationally efficient. These results indicate that a multi-agent diffusion framework can improve controllability and image quality without incurring prohibitive computational overhead, making AuDiffusion a practical candidate for applications such as visual prototyping in creative workflows.