A Unified Hybrid Encoder–Decoder Model for Vision–Language Image Captioning
Abstract
Image captioning is a combined activity that generates relevant text descriptions of images using natural language processing and computer vision. The outcomes are heavily influenced on the supplied image’s quality and clarity. This research presents a hybrid system that combines ResNet50 and MobileNet into a single encoder in order to enhance captioning performance. The Integrated Model of Vision and Language Processing helps create captions that are more precise, logical, and context-aware by learning both language patterns and visual characteristics simultaneously. After testing the suggested model on the Flickr8k and Flickr30k datasets, it achieved a maximum accuracy of 98%.
© 2026 Moloy Dhar, Mrinmoy Sen, Bidesh Chakraborty, Suparna Biswas, published by International Journal on Smart Sensing and Intelligent Systems
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.