DiffMoCa: Diffusion Model Based Multi-modality Cut and Paste
摘要
The Multi-mOdality Cut and pAste (MoCa) method cuts data from other frames and pastes it onto the current training data frame to increase the number of training object samples. However, the samples used by MoCa are all derived from the original dataset, which limits its ability to enhance object diversity. Recently, diffusion models have achieved remarkable results in the field of image generation, where simple prompts can enable the model to create entirely different paintings. In this paper, we propose DiffMoCa, which leverages the powerful creative capabilities of diffusion models to redraw the images cut by MoCa, thereby increasing the diversity of the objects and enhancing the generalization ability of the model. DiffMoCa demonstrates its capabilities in extensive experiments, wherein it surpasses MoCa by 2.2% in mAP on the KITTI dataset under moderate conditions.