Implementing efficient Deep Neural Networks (DNNs) for dense-prediction vision applications on embedded heterogeneous SoCs comes with many challenges, such as latency and energy constraints. To tackle them, we propose a novel and practical multi-objective Hardware-aware Neural Architecture Search (HW-NAS) framework able, for the first time, to handle complex search spaces while considering the hardware manufacturer’s expertise. This HW-NAS flow targeting Nvidia’s Orin SoCs relies on (1) a practical strategy to reduce the total exploration duration, and (2) a compact enhancement of the existing TensorRT deployment flow. On the FasterSeg’s search space, our framework can obtain a latency-power-mIoU Pareto front for multiple power modes in only 66 h (-33 % than the inital flow) using 8 Nvidia A100 GPUs. Compared to default mappings, these results demonstrate that our novel mapping strategy can obtain practical solutions with either 50 % less power consumption or 80 % less latency for the same accuracy performance, or achieve a better accuracy (+6 %) with 30 % less power consumption.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Practical HW-Aware NAS Flow for AI Vision Applications on Embedded Heterogeneous SoCs

  • Agathe Archet,
  • Nicolas Ventroux,
  • Nicolas Gac,
  • François Orieux

摘要

Implementing efficient Deep Neural Networks (DNNs) for dense-prediction vision applications on embedded heterogeneous SoCs comes with many challenges, such as latency and energy constraints. To tackle them, we propose a novel and practical multi-objective Hardware-aware Neural Architecture Search (HW-NAS) framework able, for the first time, to handle complex search spaces while considering the hardware manufacturer’s expertise. This HW-NAS flow targeting Nvidia’s Orin SoCs relies on (1) a practical strategy to reduce the total exploration duration, and (2) a compact enhancement of the existing TensorRT deployment flow. On the FasterSeg’s search space, our framework can obtain a latency-power-mIoU Pareto front for multiple power modes in only 66 h (-33 % than the inital flow) using 8 Nvidia A100 GPUs. Compared to default mappings, these results demonstrate that our novel mapping strategy can obtain practical solutions with either 50 % less power consumption or 80 % less latency for the same accuracy performance, or achieve a better accuracy (+6 %) with 30 % less power consumption.