Traffic Controller Data Synthesis for Autonomous Vehicles: Joint Baton-conditioned Diffusion Model and Scene-aware Scaling
摘要
Traffic controllers play an important role in managing traffic flow at construction and accident sites. However, they are rarely encountered in typical driving environments, meaning most existing road scene datasets contain very few traffic controllers. These individuals are usually labeled as ordinary pedestrians rather than as a separate class. As a result, autonomous driving systems trained on such datasets cannot distinguish traffic controllers from regular pedestrians, and are unable to respond appropriately in these special situations. To address this data scarcity, we propose a two-stage data augmentation framework. In the first stage, we introduce a baton-aware pose-guided diffusion model (B-PDM), which generates diverse images of traffic controllers by controlling the pose of both the body and the baton. In the second stage, we propose a human height estimation module (HHEM) that predicts the appropriate height for each traffic controller based on the scene context. Using these predictions, we realistically paste the generated controllers onto various road background images. Our framework enables the effective creation of large-scale, realistic training data for traffic controllers, as validated by our experiments. Our generated dataset of 10,000 images achieves a Fréchet inception distance (FID) of 38.4 and a learned perceptual image patch similarity (LPIPS) of 0.04 when compared to the original Cityscapes dataset, confirming the data’s realism. Moreover, a detection model trained with our synthetic data achieves an average precision (AP50) of 85.9% for traffic controller recognition, showcasing its practical applicability for downstream autonomous driving tasks.