Ensemble Deep Reinforcement Learning for Financial Trading
摘要
Stocks trading strategy plays an important role in financial investment. However, it is challenging to come up with an optimal profit-making portfolio in a volatile market. In this thesis, we proposed a couple of ensemble methods that use a few deep reinforcement learning (DRL) architectures to train on dynamic markets and learn complex trading strategies to achieve maximum returns on investments. We proposed three ensemble strategies with three different RL Actor-Critic algorithms as constituents: Twin Delayed Deep Deterministic Policy Gradient (TD3), Deep Deterministic Policy Gradient (DDPG), and Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC). These three ensembles are as follows: (i) PPO, TD3, and DDPG (ii) SAC, PPO, and TD3 (iii) DDPG, SAC, and PPO and compared their performance with that of the state-of-the-art ensemble method, performance namely, Advantage Actor Critic (A2C), PPO, and DDPG. The ensemble techniques adapt to various market conditions by utilizing the best aspects of all three algorithms. The effectiveness of these ensembles is demonstrated on 30 Sensex stocks with sufficient liquidity and 30 Dow Jones Industrial Average (DJIA) indexed stocks. The Sharpe ratio and maximum drawdown are employed to evaluate the performance of the ensemble methods.