Twin-Based Reinforcement Learning for Solving Multi-period Portfolio Optimization Problem
摘要
Portfolio optimization (PO) remains a key area of interest in finance. However, traditional methods face difficulties in creating an efficient and stable investment strategy due to the ever-changing nature of financial markets. These markets are influenced by internal factors that affect investment performance, such as volatility, trends, asset correlations, etc. As a result, a more dynamic and adaptive approach is needed to handle the market’s volatility and uncertainty. This paper introduces a twin-based deep reinforcement learning framework to address these challenges. The framework consists of two agents: one manages the relaxed constraint problem, and the other handles the full constraint problem. The policies from the full constraint agent are periodically merged with those from the relaxed constraint agent, improving overall portfolio performance and risk management. We selected Deep Deterministic Policy Gradient (DDPG) as the core reinforcement learning algorithm. Additionally, we introduced constraints to mitigate market shocks during trading and prevent capital erosion. This study utilizes financial data from various Vietnamese stock classes, covering the period from January 2009 to December 2023. Fundamental stock price data is used to compute key metrics such as price change rates, cash dividends, and technical indicators like moving average convergence divergence (MACD). These comprehensive inputs boost model performance, while the interaction between the two modules further strengthens the framework through policy blending. Empirical experiments demonstrated that DDPG outperformed other well-known algorithms, including PPO, A2C, SAC, and TD3, while our twin-based approach improved overall portfolio performance and risk management.