Adversarial Robust Reinforcement Learning for Secure and Regulation-Aware Internet of Things Systems
Abstract
Reinforcement learning has emerged as a promising paradigm for adaptive Internet of Things (IoT) security due to its ability to optimize sequential defense decisions in dynamic environments. Yet, the vast majority of reinforcement learning–based IoT security solutions take benign learning conditions and impose regulatory imperatives as static post-processing constraints. You are aware that these assumptions allow adversarial exploitation of the learning loop with respect to state observations, reward feedback, and policy updates while still adhering to the rules. To address this limitation, in this paper, we propose an adversarially robust and regulation-aware reinforcement learning framework that explicitly models the learning process itself as a possible attack surface. By integrating bounded adversarial threat modeling and dynamic, compliance-aware policy optimization, the approach leverages uncertainty-aware reward shaping and constrained robust optimization to maintain policy stability and regulatory compliance. Experimental assessment over benchmark IoT intrusion detection datasets shows that the framework achieves up to 35% more stability in policy, reduces security degradation by more than 30% under adversarial conditions, and dramatically lowers compliance violation rates compared to classical architectures relying solely on reinforcement learning. These results emphasize how the security of the learning process itself is essential to the establishment of trust-based and regulation-compliant IoT security systems.
© 2026 Mohammed Farsi, Elsayed Atlam, published by Cerebration Science Publishing Co., Limited
This work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License.