Lotus: Loading Cost-Aware Joint Mining Service Caching, Request Routing, and Bandwidth Orchestration in Cooperative MEC Networks
摘要
By deploying deep neural network (DNN) models at the edge server, the computing capability of mobile devices running data mining applications can be greatly expanded having shorter delays than traditional cloud based paradigms. However, there exists a performance gap between theory and practise for current AI service caching and inference request routing solutions. Loading cost is one of the dominating factors that affect the edge-AI system performance, which cannot be neglected. Distinct from existing studies, we jointly optimize caching, computation, and communication in a cooperative multi-access edge computing (MEC) network, having loading cost in mind. The problem is modeled as a mixed-integer nonlinear programming, which is proven to be NP-hard. We then devise an on-line approximation algorithm on the basis of approximate submodular property. Extensive simulation results demonstrate the superiority of proposed algorithm.