Skip to main content
Have a personal or library account? Click to login
A Unified Hybrid Encoder–Decoder Model for Vision–Language Image Captioning Cover

A Unified Hybrid Encoder–Decoder Model for Vision–Language Image Captioning

Open Access
|Jul 2026

References

  1. O. Vinyals, A. Toshev, S. Bengio, D. Erhan, “Show and tell: A neural image caption generator,” In CVPR, doi:10.1109/CVPR.2015.7298935, 2015.
  2. G. Hoxha, F. Melgani, J. Slaghenauffi, “A new CNN-RNN framework for remote sensing image captioning,” In: 2020 Mediterranean and Middle-East Geoscience and Remote Sensing Symposium (M2GARSS), Tunis, Tunisia, p. 1–4, June, 2020, doi:10.1109/M2GARSS47143.2020.9105191.
  3. P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, L. Zhang, “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering”, In CVPR, 2018, https://doi.org/10.48550/arXiv.1707.07998.
  4. S.J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, V. Goel, “Self-critical sequence training for image captioning,” In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  5. K. Xu et al., “Show, attend and tell: Neural image caption generation with visual attention,” Int. Conf. Mach. Learn., pp. 2048–2057, 2015.
  6. M. Cornia, M. Stefanini, L. Baraldi, R. Cucchiara, “Meshed-Memory Transformer for Image Captioning,” In CVPR, 2020, https://doi.org/10.48550/arXiv.1912.08226.
  7. X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, et al., “OSCAR: object-semantics aligned pre-training for vision-language tasks,” In: European Conference on Computer Vision (ECCV), p. 121–137, 2020.
  8. J. Sudhakar, V.V. Iyer, S.T. Sharmila, “Image caption generation using deep neural networks,” In: International Conference for Advancement in Technology (ICONAT), Goa, India. p. 1–3, March, 2022, doi:10.1109/ICONAT53423.2022.9726074.
  9. P. Li, M. Zhang, P. Lin, J. Wan, M. Jiang, “Visual-text reference pre-training model for image captioning,” Comput Intell Neurosci, 2022:9400999, 2022.
  10. N. Li, Z. Chen, S. Liu, “Meta-learning for image captioning,” Proc AAAI Conf Artif Intell, 2019, 33:8626–8633.
  11. W. Kang, W. Hu, “A survey of image caption tasks,” In: 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI), Nanjing, China, p. 71–74, November, 2022, doi:10.1109/CEI57409.2022.9950150.
  12. H. Maru, T. Chandana, D. Naik, “Comparison of image encoder architectures for image captioning,” In: 5th International Conference on Computing Methodologies and Communication (ICCMC), Erode, India, p. 740–744, April, 2021, doi:10.1109/ICCMC51019.2021.9418234.
  13. S. Yıldız, A. Memiş, S. Varlı, “Automatic Turkish image captioning: the impact of deep machine translation,” In: 8th International Conference on Computer Science and Engineering (UBMK), Burdur, Turkiye, p. 414–419, October, 2023, doi:10.1109/UBMK59864.2023.10286693.
  14. P. Tian, H. Mo, L. Jiang, “Improved image captioning via semantic feature update,” In: 40th Chinese Control Conference (CCC), Shanghai, China, p. 7938–7943, July, 2021, doi:10.23919/CCC52363.2021.9549991.
  15. A. P. Singh, M. Manoria and S. Joshi, “A Review on Automatic Image Caption Generation for Various Deep Learning Approaches,” 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, pp. 1–5, doi: 10.1109/ICCCNT56998.2023.10308085, 2023.
  16. Yu Z, Fu K, Jin H, Bai J, Zhang H, Li Y, “Local and global multimodal interaction for image caption,” In: 4th International Conference on Electronic Communication and Artificial Intelligence (ICECAI), Guangzhou, China, p. 164–169, May, 2023, doi:10.1109/ICECAI58670.2023.10176671.
  17. Krasin I, Duerig T, Alldrin N, Ferrari V, Abu-El-Haija S, Kuznetsova A, et al., “Open images: a public dataset for large-scale multi-label and multi-class image classification [dataset on the Internet],”, June, 2017, Available from: https://github.com/openimages.
  18. Kumari A, Chauhan A, Singhal A, “Vision 360: image caption generation using encoder-decoder model,” In: 12th International Conference on Cloud Computing, Data Science & Engineering (Confluence), Noida, India, p. 312–317, January, 2022, doi:10.1109/Confluence52989.2022.9734167.
  19. D. J. B. Saini, S. Kumar, K. Joshi, A. K. Pathak, S. Jain and A. Singh, “A Novel Approach of Image Caption Generator using Deep Learning,” Third International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), Gobichettipalayam, India, 2023, pp. 24–29, 2023, doi: 10.1109/ICUIS60567.2023.00012.
  20. Z. U. Kamangar, G. M. Shaikh, S. Hassan, N. Mughal and U. A. Kamangar, “Image Caption Generation Related to Object Detection and Colour Recognition Using Transformer-Decoder,” 4th International Conference on Computing, Mathematics and Engineering Technologies (iCoMET), Sukkur, Pakistan, pp. 1–5, 2023, doi: 10.1109/iCoMET57998.2023.10099161.
  21. Krishna, R., Zhu, Y., Groth, O. et al. Visual Genome: Connecting Language and Vision Utilizing Crowd sourced Dense Image Annotations. Int J Comput Vision 123, 32–73, 2017, https://doi.org/10.1007/s11263-016-0981-7.
  22. Z. U. Kamangar, G. M. Shaikh, S. Hassan, N. Mughal and U. A. Kamangar, “Image Caption Generation Related to Object Detection and Colour Recognition Using Transformer-Decoder,” 4th International Conference on Computing, Mathematics and Engineering Technologies (iCoMET), Sukkur, Pakistan, 2023, pp. 1–5, 2023, doi: 10.1109/iCoMET57998.2023.10099161.
  23. A. P. Singh, M. Manoria and S. Joshi, “A Review on Automatic Image Caption Generation for Various Deep Learning Approaches,” 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, pp. 1–5, 2023, doi: 10.1109/ICCCNT56998.2023.10308085.
  24. D. J. B. Saini, S. Kumar, K. Joshi, A. K. Pathak, S. Jain and A. Singh, “A Novel Approach of Image Caption Generator using Deep Learning,” 2023 Third International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), Gobichettipalayam, India, pp. 24–29, 2023, doi: 10.1109/ICUIS60567.2023.00012.
  25. K. Fatema, S. Montaha, M. A. H. Rony, S. Azam, M. Z. Hasan, and M. Jonkman, “A robust framework combining image processing and deep learning hybrid model to classify cardiovascular diseases using a limited number of paper-based complex ECG images,” Biomedicines, 10, 2835, 2022, https://doi.org/10.3390/biomedicines10112835.
  26. J. Yang and N. A. Mat Isa, “YOLOv10-MsA: Attention-Augmented Real-Time Insulator Defect Detection from UAV Imagery,” Emerging Science Journal, vol. 9, no. 6, pp. 2899–2914, 2025, doi: 10.28991/ESJ-2025-09-06-02.
  27. Z. Deng, X. Li, R. Yang, “RML-YOLO: An Insulator Defect Detection Method for UAV Aerial Images,” Applied Sciences, 15(11):6117, 2025, https://doi.org/10.3390/app15116117.
  28. R. D. Puriyanto, I.D. Yunandha, H. Maghfiroh, A. Ma’arif, Furizal and I. Suwarno, “Ball Detection System for a Soccer on Wheeled Robot Using the MobileNetV2 SSD Method,” Emerging Science Journal 9, 5, 2782–2796, 2025, doi: https://doi.org/10.28991/ESJ-2025-09-05-028.
  29. X. Lian, D. Wang, “Insulator defect detection algorithm based on improved YOLOv5,” Frontiers in Computing and Intelligent Systems, 3(2), 44–47, 2023. https://doi.org/10.54097/fcis.v3i2.7168.
  30. J. Liu, C. Liu, Y. Wu, Z. Sun, H. Xu, “Insulators’ Identification and Missing Defect Detection in Aerial Images Based on Cascaded YOLO Models,” Comput Intell Neurosci, 2022:7113765, 2022, August, doi: 10.1155/2022/7113765. PMID: 36035858; PMCID: PMC9402330.
  31. X. Sun, S. Qi, “A Study on Defect Detection of YOLOV8 Insulators Based on Improvements,” Academic Journal of Computing & Information Science, Vol. 7, Issue 5: 37–43, 2024, https://doi.org/10.25236/AJCIS.2024.070505.
  32. N. Fahad, M. J. Hossen, and M. S. Sayeed, “Efficient object detection with an optimized YOLOv8x model,” HighTech and Innovation Journal, vol. 6, no. 3, Article 09, Sep. 2025, doi: 10.28991/HIJ-2025-06-03-09.
  33. C. Jia, D. Wang, J. Liu, W. Deng, “Performance Optimization and Application Research of YOLOv8 Model in Object Detection,” Academic Journal of Science and Technology, 10(1), 325–329, 2024, https://doi.org/10.54097/p9w3ax47.
  34. Y. Zhao, N. C. Rodelas, “Enhanced Small Target Detection Methodology via Optimized YOLOv8 Framework,” Academic Journal of Science and Technology, 11(3), 80–84, 2024, https://doi.org/10.54097/7terh631.
  35. B. Yilmaz, U. Kutbay, “YOLOv8-Based Drone Detection: Performance Analysis and Optimization,” Computers, 13(9):234, 2024, https://doi.org/10.3390/computers13090234.
  36. U. Verma, “YOLOV8: An Enhanced Object Detection Model for Distance Estimation,” International Journal of Intelligent Systems and Applications in Engineering, 12(21s), 3852, 2024, https://www.ijisae.org/index.php/IJISAE/article/view/6156.
  37. S. Erniwati, V. Afifah and B. Imran, “Mask Region-based Convolutional Neural Network in Object Detection: A Review,” IJACI: International Journal of Advanced Computing and Informatics, vol. 1, no. 2, pp. 106–117, 2025, doi: 10.71129/ijaci.v1i2.pp106-117.
  38. D. Juhartini, D. Arwidiyarti, and Desmiwati, “Single Shot Multibox Detector (SSD) in Object Detection: A Review,” IJACI: International Journal of Advanced Computing and Informatics, vol. 1, no. 2, pp. 118–127, 2025, doi: 10.71129/ijaci.v1i2.pp118-127.
  39. R. Sastra, D. Hariyanto Dicky, and B. Apriyansyah, “Fast Region-based Convolutional Neural Network in Object Detection: A Review,” IJACI: International Journal of Advanced Computing and Informatics, vol. 2, no. 1, pp. 34–40, 2026, doi: 10.71129/ijaci.v2i1.pp34-40.
  40. W. Ilham and A. Ahmad, “A comprehensive review of ConvNeXt architecture in image classification: Performance, applications, and prospects,” IJACI: International Journal of Advanced Computing and Informatics, vol. 2, no. 2, pp. 108–114, 2026, doi: 10.71129/ijaci.v2i2.pp108-114.
  41. R. D. Ramadhan and Y. A. Adam, “Efficient tomato leaf disease classification using adapter based fine tuning of Vision Transformers,” IJACI: International Journal of Advanced Computing and Informatics, vol. 2, no. 2, pp. 127–138, 2026, doi: 10.71129/ijaci.v2i2.pp127-138.
  42. V. Eriyandi and A. N. Ahmad, “A comparative machine learning and deep learning models with ElasticNet regularization for predicting student outcomes: LGBM, CatBoost, ANN, DNN, and WDNN,” IJACI: International Journal of Advanced Computing and Informatics, vol. 2, no. 2, pp. 139–150, 2026, doi: 10.71129/ijaci.v2i2.pp139-150.
  43. N. S. Rahmawati and C. I. Amalia, “Lightweight ensemble models for static malware detection: Addressing deep learning trade-offs with the Kaggle PE dataset,” IJACI: International Journal of Advanced Computing and Informatics, vol. 2, no. 1, pp. 57–66, 2026, doi: 10.71129/ijaci.v2i1.pp57-66.
  44. K. He, G. Gkioxari, P. Dollár and R. Girshick, “Mask R-CNN,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 2, pp. 386–397, 2020, doi: 10.1109/TPAMI.2018.2844179.
  45. Z.-Q. Zhao, P. Zheng, S.-T. Xu and X. Wu, “Object detection with deep learning: A review,” IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 11, pp. 3212–3232, 2019, doi: 10.1109/TNNLS.2018.2876865.
  46. T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan and S. Belongie, “Feature Pyramid Networks for object detection,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 936–944, doi: 10.1109/CVPR.2017.106.
  47. H. Li, X. Qi and G.-J. Qi, “A guide to convolutional neural networks for visual recognition,” Foundations and Trends® in Computer Graphics and Vision, vol. 15, no. 3, pp. 207–371, 2021, doi: 10.1561/0600000117.
  48. X. Wang, D. Zhang, G. Wang, and Z. Cui, “Performance improvement techniques for Mask R-CNN in visual detection tasks,” Journal of Visual Communication and Image Representation, vol. 81, p. 103426, 2021, doi: 10.1016/j.jvcir.2021.103426.
Language: English
Submitted on: Dec 11, 2025
Published on: Jul 15, 2026
Published by: International Journal on Smart Sensing and Intelligent Systems
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year

© 2026 Moloy Dhar, Mrinmoy Sen, Bidesh Chakraborty, Suparna Biswas, published by International Journal on Smart Sensing and Intelligent Systems
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.