Skip to main content
Have a personal or library account? Click to login
Towards Automated Identification of Block Cipher Structures Using Machine Learning Cover

Towards Automated Identification of Block Cipher Structures Using Machine Learning

Open Access
|Jul 2026

Abstract

This paper focuses on the classification of block ciphers based on their structure, which is a crucial step in automated cryptanalysis. For this purpose, five machine learning algorithms, including Logistic Regression, Naive Bayes, k-Nearest Neighbour, Random Forest, and XGBoost, are applied to ciphertext data to differentiate between Substitution–Permutation Networks (SPN) and Feistel block cypher structures. A comprehensive experimental framework is developed using 16 block ciphers, including AES, DES, Blowfish, Twofish, and Camelia, etc evaluated across two encryption modes (ECB and CBC) with 5-fold cross validation under four different scenarios. To improve generalization and eliminate dataset bias, ciphertext datasets are generated using fixed key and multiple random keys and diverse plaintext sources, including text, source code, and binary data. A feature extraction pipeline based on byte-level statistics, compression characteristics, and higher-order histogram features is employed, followed by classification using five machine learning and one deep learning models. NIST SP800-22 suite tests were conducted on ciphertext dataset. Results confirmed that despite ciphertexts exhibit near-ideal entropy and pass standard randomness tests; however, small higher-order statistical deviations remain observable. Experimental results demonstrate that these weak statistical signals can be leveraged to achieve classification accuracies exceeding 90% for ciphertext, with performance increasing as ciphertext size increases. The classification accuracy is highest for ECB and decreases for CBC, consistent with increasing diffusion and inter-block dependency. Feature importance and ablation studies further reveal that statistical features provide the dominant discriminative signal. These findings highlight the presence of finite-sample statistical artifacts that can be detected using machine learning, and provide new insights into the intersection of statistical cryptanalysis and machine learning, with implications for automated cryptographic analysis.

DOI: https://doi.org/10.2478/ias-2026-0007 | Journal eISSN: 1554-1029 | Journal ISSN: 1554-1010
Language: English
Page range: 119 - 136
Published on: Jul 8, 2026
In partnership with: Paradigm Publishing Services
Publication frequency: 6 issues per year

© 2026 Uroosa Kiran, Hammad Tanveer Butt, Zunera Jalil, published by Cerebration Science Publishing Co., Limited
This work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 License.