Scene image visual layout based on deep encoder–decoder network and visual image attention model
摘要
Generating complex scene layouts faces challenges such as cross-modal semantic alignment bias and low efficiency in modeling dynamic spatiotemporal relationships. Existing methods have limitations on physical constraint awareness and real-time performance, which hinder the application of virtual reality and smart cities. To address the three core issues of cross-modal semantic alignment bias, low efficiency in modeling dynamic spatiotemporal relationships, and insufficient physical constraint awareness, this study proposes a framework based on deep encoding-decoding networks and visual graph attention. It introduces overall nested edge detection to optimize multi-scale feature fusion and designs an explicit edge feature modeling strategy to enhance physical constraint awareness. The framework is compared with benchmark methods such as LayoutTransformer and SceneGraphNet on the COCO-Stuff dataset (172,000 annotated images). Intersection-over-Union (IoU) ratio, feature coverage, and energy consumption are compared. The results demonstrate that the model achieved an IoU ratio of 0.82 and a key feature coverage rate of 94.6% in common object context datasets, with a training efficiency improvement of 38% compared with the benchmark method. The energy consumption during the inference phase was controlled at 54.9Wh, and the memory usage was reduced by 19.5%. Real-world scene testing showed that the IoU difference between the generated layout and the preset plan was less than 0.03, with an average score of 9.39 points for manual evaluation (9.43 points for professional designers), and a layout overlap rate below the 4% threshold in dynamic scenes. The model demonstrated near professional design capabilities in 30 subjective evaluations while maintaining a real-time inference speed of 26.94ms, validating the comprehensive advantages in cross-modal alignment, physical constraint perception, and computational efficiency. In summary, the research model optimizes cross-modal semantic consistency, multi-scale dynamic interaction relationship modeling, and complex spatial constraint perception capabilities, providing high-precision and efficient layout generation technology support for virtual reality scene construction and digital design of smart cities.