Interpretable Region-Attribute Graph Matching Framework for Structure-Aware and Explainable Content-Based Image Retrieval

Abstract
In recent years, Content-Based Image Retrieval (CBIR) has advanced significantly with Deep Learning (DL), yet most of the existing models rely on global descriptors and attention-based local features, which struggle to capture explicit structural relationships within images. These models often operate as black boxes and provide limited interpretability; they exhibit reduced robustness under viewpoint and geometric variations. To address these limitations, an Interpretable Region-Attribute Graph Network (IR-AGN) is proposed in this research. Specifically, each image is decomposed into semantically meaningful regions, where each node encodes visual and semantic descriptors while edges represent geometric, appearance, and contextual relations. Consequently, a joint node-edge graph matching strategy computes similarity by aligning visual descriptors and relational structures between query and retrieval images. The results of the proposed IR-AGN model exhibit a higher mean Average Precision (mAP) of 95.19% on the Revisited Paris (RParis) dataset, when compared with the existing Multi-Layer Orientation Histogram (MLOH) model.
© 2026 Prabhuraj Metipatil, M. S. Mrutyunjaya, Sanjeev Prakashrao Kaulgud, Sunil Manoli, published by Bulgarian Academy of Sciences, Institute of Information and Communication Technologies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.