Figure 1.

Figure 2.

Figure 3.

Figure 4.

Figure 5.

Figure 6.

Regulation Compliance and Robust Performance Trade-off
| Dataset | Method | Compliance Violations (%) ↓ | Avg. Constraint Slack ↓ | Security Performance (%) ↑ |
|---|---|---|---|---|
| NSL-KDD | Static Constraints | 18.6 | 0.041 | 84.3 |
| Proposed | 3.2 | 0.009 | 88.7 | |
| UNSW-NB15 | Static Constraints | 24.1 | 0.053 | 81.6 |
| Proposed | 4.8 | 0.012 | 86.9 | |
| BoT-IoT | Static Constraints | 21.7 | 0.047 | 79.2 |
| Proposed | 5.5 | 0.015 | 83.8 |
Stability in Decentralized and Federated Learning
| Dataset | Method | Policy Divergence (|| θ – θ\* ||) ↓ | Malicious Agents (%) | Global Performance Drop (%) ↓ |
|---|---|---|---|---|
| NSL-KDD | Federated RL | 0.217 | 20 | 26.5 |
| Proposed | 0.082 | 20 | 9.4 | |
| UNSW-NB15 | Federated RL | 0.284 | 25 | 33.8 |
| Proposed | 0.096 | 25 | 11.8 | |
| BoT-IoT | Federated RL | 0.312 | 30 | 37.4 |
| Proposed | 0.124 | 30 | 14.6 |
Policy Stability and Security Degradation under Learning-Loop_
| Dataset | Method | Policy Instability (Var[J]) ↓ | Convergence Steps ↓ | Security Degradation (%) ↓ |
|---|---|---|---|---|
| NSL-KDD | Classical RL | 0.182 | 4,200 | 31.4 |
| Proposed | 0.061 | 2,750 | 11.2 | |
| UNSW-NB15 | Classical RL | 0.247 | 5,100 | 38.9 |
| Proposed | 0.073 | 3,100 | 13.6 | |
| BoT-IoT | Classical RL | 0.301 | 5,800 | 42.7 |
| Proposed | 0.098 | 3,950 | 17.9 |