DRL Based SFC Orchestration in SDN/NFV Environments Subject to Transient Unavailability
摘要
In this paper, we address how complex and dynamic environments characterized by variable and limited capacities and subject to transient unavailability may pose significant challenges for Deep Reinforcement Learning (DRL) agents. The investigation concerns the context of the Service Function Chaining (SFC) orchestration problem in Software-Defined Networking (SDN) and Network Function Virtualization (NFV)-based environments using the DRL approach, implemented through Deep Q-Network (DQN), and aims to maximize Quality of Experience (QoE) while meeting Quality of Service (QoS) constraints. We show through numerical results how limited capacity in the Physical Substrate Network (PSN) complicates the training process in terms of finding a suitable compromise between performance and convergence. We highlight how replay buffers may mitigate the transient unavailability of PSN nodes and what the limits of such a solution are when the unavailability becomes more prolonged in time or more severe (simultaneous unavailability of more than one node).