International Journal of Innovative Research in Computer and Communication Engineering
ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines
| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |
| TITLE | A Comprehensive Survey on Image Caption Generation Using Deep Learning: Evolution, Comparative Analysis, and Emerging Vision-Language Models |
|---|---|
| ABSTRACT | Image caption generation is a multimodal artificial intelligence task that automatically converts visual content into natural-language descriptions. The problem combines computer vision for visual understanding with natural language processing for sentence generation, making it substantially more complex than conventional image classification. Deep learning has transformed this field from early convolutional neural network (CNN) and recurrent neural network (RNN) encoder-decoder models toward attention-based systems, object-aware models, graph reasoning, reinforcement learning, Transformer architectures, vision-language pretraining, and multimodal large language models (MLLMs). This survey reviews the evolution of these approaches and organizes representative studies according to their central modelling strategy. Particular attention is given to visual attention, object-level representations, scene and relationship modelling, Transformer-based cross-modal interaction, and the recent transition toward pretrained vision-language systems. The survey also compares benchmark datasets such as Flickr8k, Flickr30k, MS COCO, Conceptual Captions, and nocaps, and discusses evaluation metrics including BLEU, METEOR, ROUGE-L, CIDEr, SPICE, CLIPScore, BERTScore, and CHAIR. In addition, major limitations are examined, including object hallucination, weak visual grounding, dataset bias, long-tail recognition, computational cost, and the mismatch between lexical similarity metrics and factual correctness. The review concludes with future directions in grounded captioning, controllable generation, multilingual captioning, parameter-efficient adaptation, knowledge-grounded generation, efficient deployment, and trustworthy multimodal AI. |
| AUTHOR | UMESH N WAGH, PROF. DR. B.D.PHULPAGAR Department of Computer Engineering, P. E. S. Modern College of Engineering, Pune, India |
| VOLUME | 187 |
| DOI | DOI: 10.15680/IJIRCCE.2026.1408022 |
| pdf/22_A Comprehensive Survey on Image Caption Generation Using Deep Learning Evolution, Comparative Analysis, and Emerging Vision-Language Models.pdf | |
| KEYWORDS | |
| References | [1] M. A. Al-Malla, O. Hamdoun, and N. Ghneim, “A comprehensive survey on deep learning approaches for image captioning: A systematic review,” Journal of Big Data, vol. 13, art. 48, 2026. Springer Nature. doi: 10.1186/s40537-026-01377-w. https://link.springer.com/article/10.1186/s40537-026-01377-w [2] A. Sharma et al., “Pixels to prose: A comprehensive survey of image captioning techniques with deep learning and generative artificial intelligence,” Neurocomputing, vol. 667, art. 132385, 2026. Elsevier. doi: 10.1016/j.neucom.2025.132385. https://doi.org/10.1016/j.neucom.2025.132385 [3] H. D. Abdulgalil and O. A. Basir, “Next-generation image captioning: A survey of methodologies and emerging challenges from transformers to Multimodal Large Language Models,” Natural Language Processing Journal, vol. 12, art. 100159, 2025. Elsevier. doi: 10.1016/j.nlp.2025.100159. https://doi.org/10.1016/j.nlp.2025.100159 [4] “Deep image captioning: A review of methods, trends and future challenges,” Neurocomputing, vol. 546, art. 126287, 2023. Elsevier. doi: 10.1016/j.neucom.2023.126287. https://doi.org/10.1016/j.neucom.2023.126287 [5] K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proc. ICML, vol. 37, pp. 2048-2057, 2015. https://proceedings.mlr.press/v37/xuc15.html [6] P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proc. IEEE/CVF CVPR, pp. 6077-6086, 2018. doi: 10.1109/CVPR.2018.00636. https://openaccess.thecvf.com/content_cvpr_2018/html/Anderson_Bottom-Up_and_Top-Down_CVPR_2018_paper.html [7] L. Huang, W. Wang, J. Chen, and X.-Y. Wei, “Attention on attention for image captioning,” in Proc. IEEE/CVF ICCV, pp. 4634-4643, 2019. https://openaccess.thecvf.com/content_ICCV_2019/html/Huang_Attention_on_Attention_for_Image_Captioning_ICCV_2019_paper.html [8] S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proc. IEEE/CVF CVPR, pp. 7008-7024, 2017. https://openaccess.thecvf.com/content_cvpr_2017/html/Rennie_Self-Critical_Sequence_Training_CVPR_2017_paper.html [9] Q. Sun, J. Zhang, Z. Fang, and Y. Gao, “Self-enhanced attention for image captioning,” Neural Processing Letters, vol. 56, art. 131, 2024. Springer Nature. doi: 10.1007/s11063-024-11527-x. https://link.springer.com/article/10.1007/s11063-024-11527-x [10] Y. Shi, J. Xia, M. C. Zhou, and Z. Cao, “A dual-feature-based adaptive shared transformer network for image captioning,” IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1-13, 2024. doi: 10.1109/TIM.2024.3353830. https://ieeexplore.ieee.org/document/10400415/ [11] S. Cao, G. An, Z. Zheng, and Z. Wang, “Vision-enhanced and consensus-aware transformer for image captioning,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 10, pp. 7005-7018, 2022. doi: 10.1109/TCSVT.2022.3178844. https://ieeexplore.ieee.org/document/9784827/ [12] J. Zhang, Y. Xie, W. Ding, and Z. Wang, “Cross on cross attention: Deep fusion transformer for image captioning,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4257-4268, 2023. doi: 10.1109/TCSVT.2023.3243725. https://ieeexplore.ieee.org/document/10041164/ [13] X. Li et al., “OSCAR: Object-semantics aligned pre-training for vision-language tasks,” in Proc. ECCV, 2020. https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/7133_ECCV_2020_paper.php [14] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proc. ICML, vol. 202, pp. 19730-19742, 2023. https://proceedings.mlr.press/v202/li23q.html [15] H. Agrawal et al., “nocaps: Novel object captioning at scale,” in Proc. IEEE/CVF ICCV, pp. 8948-8957, 2019. doi: 10.1109/ICCV.2019.00904. https://openaccess.thecvf.com/content_ICCV_2019/html/Agrawal_nocaps_novel_object_captioning_at_scale_ICCV_2019_paper.html [16] J. Wang, W. Wang, L. Wang, Z. Wang, D. Dagan Feng, and T. Tan, “Learning visual relationship and context-aware attention for image captioning,” Pattern Recognition, vol. 98, art. 107075, 2020. Elsevier. doi: 10.1016/j.patcog.2019.107075. https://doi.org/10.1016/j.patcog.2019.107075 [17] J. Zhang, K. Li, Z. Wang, X. Zhao, and Z. Wang, “Visual enhanced gLSTM for image captioning,” Expert Systems with Applications, vol. 184, art. 115462, 2021. Elsevier. doi: 10.1016/j.eswa.2021.115462. https://doi.org/10.1016/j.eswa.2021.115462 [18] Y.-L. Chang, H.-S. Ma, S.-C. Li, B. P. Jaysawal, et al., “GDVT: geometrically aware dual transformer encoding visual and textual features for image captioning,” International Journal of Data Science and Analytics, vol. 20, pp. 6507-6525, 2025. Springer Nature. doi: 10.1007/s41060-025-00836-6. https://link.springer.com/article/10.1007/s41060-025-00836-6 [19] L. Chen and K. Li, “Dual-adaptive interactive transformer with textual and visual context for image captioning,” Expert Systems with Applications, vol. 243, art. 122955, 2024. Elsevier. doi: 10.1016/j.eswa.2023.122955. https://doi.org/10.1016/j.eswa.2023.122955 [20] J. Hu, Z. Li, Q. Su, Z. Tang, and H. Ma, “Exploring refined dual visual features cross-combination for image captioning,” Neural Networks, vol. 180, art. 106710, 2024. Elsevier. doi: 10.1016/j.neunet.2024.106710. https://doi.org/10.1016/j.neunet.2024.106710 [21] “Divergent-convergent attention for image captioning,” Pattern Recognition, vol. 115, art. 107928, 2021. Elsevier. doi: 10.1016/j.patcog.2021.107928. https://doi.org/10.1016/j.patcog.2021.107928 [22] “Image captioning: Semantic selection unit with stacked residual attention,” Image and Vision Computing, vol. 144, art. 104965, 2024. Elsevier. doi: 10.1016/j.imavis.2024.104965. https://doi.org/10.1016/j.imavis.2024.104965 [23] M. A. Al-Malla, A. Jafar, and N. Ghneim, “Image captioning model using attention and object features to mimic human image understanding,” Journal of Big Data, vol. 9, art. 20, 2022. Springer Nature. doi: 10.1186/s40537-022-00571-w. https://link.springer.com/article/10.1186/s40537-022-00571-w |