<p>Accurate semantic segmentation of LiDAR point clouds demands substantial computational resources due to massive data volumes and complex feature extraction, particularly in autonomous driving scenarios. To mitigate these computational bottlenecks, this paper proposes MSF-CSCNet, an advanced 3D semantic segmentation framework explicitly optimized for parallel and distributed processing. MSF-CSCNet integrates multi-scale voxel feature fusion and enhanced channel context modeling, enabling efficient utilization of parallel computing environments. The multi-scale spatial fusion module (MSF) effectively addresses spatial density variations through efficient cross-distance feature fusion in cylindrical coordinates, thereby enabling parallel execution. Additionally, the channel context and spatial decomposition module (CSC) dynamically reorganizes feature channels and implements dual-path attention, balancing computational loads across multiple GPUs. Furthermore, the Pinwheel3DConv module, employing asymmetric convolution kernels, significantly enhances segmentation granularity while maintaining computational efficiency. Experimental evaluations conducted on a high-performance computing cluster equipped with NVIDIA RTX 3090 GPUs demonstrate near-linear scalability for MSF-CSCNet across multiple GPUs, achieving end-to-end throughput exceeding seven frames per second. Evaluations on the SemanticKITTI dataset further validate MSF-CSCNet’s superior accuracy (65.81% mIoU, 91.55% Acc, and 71.44% Acc_cls), confirming the method’s capability to effectively handle urban scenes with diverse object densities and distances. These computational results substantiate MSF-CSCNet’s effectiveness and suitability for real-time semantic perception in large-scale autonomous driving applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MSF-CSCNet: a supercomputing-ready 3D semantic segmentation network for urban point clouds via multi-scale fusion and context-aware channel modeling

  • Yun Bai,
  • Yuxuan Gong,
  • Jinlei Wang,
  • Feng Wei

摘要

Accurate semantic segmentation of LiDAR point clouds demands substantial computational resources due to massive data volumes and complex feature extraction, particularly in autonomous driving scenarios. To mitigate these computational bottlenecks, this paper proposes MSF-CSCNet, an advanced 3D semantic segmentation framework explicitly optimized for parallel and distributed processing. MSF-CSCNet integrates multi-scale voxel feature fusion and enhanced channel context modeling, enabling efficient utilization of parallel computing environments. The multi-scale spatial fusion module (MSF) effectively addresses spatial density variations through efficient cross-distance feature fusion in cylindrical coordinates, thereby enabling parallel execution. Additionally, the channel context and spatial decomposition module (CSC) dynamically reorganizes feature channels and implements dual-path attention, balancing computational loads across multiple GPUs. Furthermore, the Pinwheel3DConv module, employing asymmetric convolution kernels, significantly enhances segmentation granularity while maintaining computational efficiency. Experimental evaluations conducted on a high-performance computing cluster equipped with NVIDIA RTX 3090 GPUs demonstrate near-linear scalability for MSF-CSCNet across multiple GPUs, achieving end-to-end throughput exceeding seven frames per second. Evaluations on the SemanticKITTI dataset further validate MSF-CSCNet’s superior accuracy (65.81% mIoU, 91.55% Acc, and 71.44% Acc_cls), confirming the method’s capability to effectively handle urban scenes with diverse object densities and distances. These computational results substantiate MSF-CSCNet’s effectiveness and suitability for real-time semantic perception in large-scale autonomous driving applications.