Resource-Efficient Model Deployment for Enterprise AI
摘要
While the rise of foundation models, like large language models, has brought immense opportunities for some enterprises to roll out new products and further optimize their operational logistics, the ever-expanding digitalization gap caused by the skewed ownership of computing power is also evident, preventing AI from realizing wider business benefits. This implies that democratizing AI among all enterprises requires addressing the key challenge of making the use of state-of-the-art AI technologies less dependent on the availability of computing resources. In this chapter, with a main focus on deploying advanced yet resource-intensive AI models in less resourceful enterprise environments, we will provide an in-depth analysis of the research and applications within this context. Firstly, we will review and taxonomize the main technical pathways for reducing the resource footprint of a bulky AI model while keeping the utility trade-off minimal. Then, we will associate each category of the solutions with real examples, so as to discuss when and how their advantages can be fully leveraged. Finally, by providing a roadmap that entails high-potential R&D areas, we will conclude this chapter by pointing out future directions for developing technologies for resource-efficient enterprise AI.