Exploring Diverse Configurations of Cellular Automata Based S-Boxes Using Reinforcement Learning
摘要
This paper presents an approach for constructing efficient substitution boxes (S-boxes) for cryptographic applications by combining cellular automata (CA) and reinforcement learning (RL). Semi-bent Boolean functions derived from CA rules are used to generate the S-box output array with desirable cryptographic properties like high nonlinearity. The selection of optimal CA rules is formulated as a Markov Decision Process (MDP), where a reinforcement learning agent explores the state space of rule combinations to maximize a reward signal based on nonlinearity and differential uniformity. Various configurations for applying the selected CA rules to generate multi-layered S-boxes are explored. The proposed methodology offers advantages such as reduced memory footprint, exploration of a vast solution space, and inherent parallelism suitable for hardware implementations. Experimental results demonstrate that the generated S-boxes outperform previously proposed CA based S-boxes in terms of cryptographic strength.