An Explainable Deep Learning Approach for Distinguishing Cyber Attacks and Sensor Faults in Critical IoT Systems

Abstract
The increasing rate of cyber-attacks in critical Internet of Things (IoT) systems requires robust and interpretable detection systems. In this paper, an interpretable deep learning approach has been introduced to detect various types of cyber-attacks and distinguish them from sensor faults. The proposed model uses Convolutional Neural Networks-Bidirectional Long Short-Term Memory (CNN-BiLSTM) to extract features. The attention mechanism has also been employed to improve the interpretability of the model. In addition, the SHAP approach has been used to interpret the model’s predictions by measuring the contribution of features. The performance of the model has been evaluated using the TON_IoT dataset with an accuracy of 97.3% for multi-class classification and 94.25% for binary classification. The results show that the proposed approach offers high detection accuracy and interpretability of the results.
© 2026 Nabeel I. Zanoon, Abdullah Odeh Al-Zaghameem, Khalid Alkharabsheh, published by Bulgarian Academy of Sciences, Institute of Information and Communication Technologies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.