Online Optimization for Diverse Edge DNN Inference Serving at Scale
摘要
In this chapter, we study diverse edge DNN inference serving for multi-user and multi-application at scale, which aims to navigate the three-way trade-off between inference accuracy, latency, and resource cost via jointly optimizing the application configuration adaption, DNN model selection, and edge resource provisioning on-the-fly.