Efficient Post-training Augmentation for Adaptive Inference in Heterogeneous and Distributed IoT Environments
摘要
Early Exit Neural Networks (EENNs) achieve enhanced efficiency compared to traditional models, but creating them is challenging due to the many additional design choices required. To address this, we propose an automated augmentation flow that converts existing models into EENNs, making all necessary design decisions for deployment on heterogeneous or distributed embedded targets. Our framework is the first to perform all these steps, including EENN architecture construction, subgraph mapping, and decision mechanism configuration. We evaluated our approach on embedded Deep Learning scenarios, achieving significant performance improvements. Our solution reduced latency by 65.95% on a speech command detection problem and mean operations per inference by 78.3% on an ECG classification task. This showcases the potential for EENNs in embedded applications.