Skip to main content
Have a personal or library account? Click to login
A Unified Hybrid Encoder–Decoder Model for Vision–Language Image Captioning Cover

A Unified Hybrid Encoder–Decoder Model for Vision–Language Image Captioning

Open Access
|Jul 2026

Abstract

Image captioning is a combined activity that generates relevant text descriptions of images using natural language processing and computer vision. The outcomes are heavily influenced on the supplied image’s quality and clarity. This research presents a hybrid system that combines ResNet50 and MobileNet into a single encoder in order to enhance captioning performance. The Integrated Model of Vision and Language Processing helps create captions that are more precise, logical, and context-aware by learning both language patterns and visual characteristics simultaneously. After testing the suggested model on the Flickr8k and Flickr30k datasets, it achieved a maximum accuracy of 98%.

Language: English
Submitted on: Dec 11, 2025
Published on: Jul 15, 2026
Published by: International Journal on Smart Sensing and Intelligent Systems
In partnership with: Paradigm Publishing Services
Publication frequency: 1 issue per year

© 2026 Moloy Dhar, Mrinmoy Sen, Bidesh Chakraborty, Suparna Biswas, published by International Journal on Smart Sensing and Intelligent Systems
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.