A Novel Image Captioning Technique Using Deep Learning Methodology
Research Article  ·  Published: 01 August 2025
Issue cover
ICCK Transactions on Machine Intelligence
Volume 1, Issue 2, 2025: 52-68
Research Article Free to Read

A Novel Image Captioning Technique Using Deep Learning Methodology

1 Department of the AIML-CSE Apex Institute of Technology, Chandigarh University, Mohali, India
* Corresponding Author: Jaswinder Singh, [email protected]
Volume 1, Issue 2
You have access to this article · Limited-Time Free Access

Article Information

Abstract

The capacity of AI systems to generate captions for images autonomously represents a significant advancement in artificial intelligence and language understanding. This paper presents an advanced image captioning system that employs deep learning techniques, particularly convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to produce contextually appropriate and meaningful descriptions of visual content. The proposed method extracts features using the DenseNet201 model, enabling a more comprehensive and hierarchical understanding of image components. These extracted features are then fed into a long short-term memory (LSTM) network, a specialized RNN variant designed to capture sequential dependencies in language, yielding coherent and fluent captions. The model is trained and evaluated on the well-known Flickr8k dataset, achieving competitive performance as measured by BLEU score metrics and demonstrating its ability to generate human-like descriptions. This integration of CNNs and RNNs highlights the effectiveness of combining computer vision and natural language processing for automated caption generation. The approach has potential applications across various domains, including assistive technologies for the visually impaired, automated content creation for digital media, enhanced indexing and retrieval of multimedia assets, and improved human-computer interaction. Furthermore, advancements in attention mechanisms and transformer-based models present opportunities to further enhance the accuracy and contextual relevance of image captioning systems. The study underscores the broader implications of machine-generated captions for improving accessibility, boosting searchability in large-scale databases, and enabling seamless AI-human collaboration in content interpretation and storytelling.

Graphical Abstract

A Novel Image Captioning Technique Using Deep Learning Methodology

Keywords

convolutional neural networks (CNN) recurrent neural networks (RNN) deep learning image captioning LSTM DenseNet201 attention mechanism BLEU scor natural language processing (NLP) multimodal learning content retrieval

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Aneja, J., Deshpande, A., & Schwing, A. G. (2018). Convolutional image captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5561–5570.
    [CrossRef] [Google Scholar]
  2. Bai, S., & An, S. (2018). A survey on automatic image caption generation. Neurocomputing, 311, 291–304.
    [CrossRef] [Google Scholar]
  3. Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2017). Bottom-up and top-down attention for image captioning and vqa. arXiv preprint arXiv:1707.07998, 2(4), 8.
    [Google Scholar]
  4. Chen, X., & Zitnick, C. L. (2015). Mind’s eye: A recurrent visual representation for image caption generation. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2422–2431.
    [CrossRef] [Google Scholar]
  5. Ghandi, T., Pourreza, H., & Mahyar, H. (2023). Deep learning approaches on image captioning: A review. ACM Computing Surveys, 56(3), 1–39.
    [CrossRef] [Google Scholar]
  6. Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3128–3137.
    [CrossRef] [Google Scholar]
  7. Hossain, M. Z., Sohel, F., Shiratuddin, M. F., & Laga, H. (2019). A comprehensive survey of deep learning for image captioning. ACM Computing Surveys (CsUR), 51(6), 1-36.
    [CrossRef] [Google Scholar]
  8. Rastogi, R., Rawat, V., & Kaushal, S. (2024). Demonstration and analysing the performance of image caption generator: Efforts for visually impaired candidates for Smart Cities 5.0. International Journal of Advanced Mechatronic Systems, 11(3), 161–178.
    [CrossRef] [Google Scholar]
  9. Jamil, A., Mahmood, K., Villar, M. G., Prola, T., Diez, I. D. L. T., Samad, M. A., & Ashraf, I. (2024). Deep learning approaches for image captioning: Opportunities, challenges and future potential. IEEE Access, 12, 12345–12367.
    [CrossRef] [Google Scholar]
  10. Vo-Ho, V. K., Luong, Q. A., Nguyen, D. T., Tran, M. K., & Tran, M. T. (2019). A smart system for text-lifelog generation from wearable cameras in smart environment using concept-augmented image captioning with modified beam search strategy. Applied Sciences, 9(9), 1886.
    [CrossRef] [Google Scholar]
  11. Cornia, M., Baraldi, L., & Cucchiara, R. (2019). Show, control and tell: A framework for generating controllable and grounded captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8307–8316.
    [CrossRef] [Google Scholar]
  12. Chen, L., Jiang, Z., Xiao, J., & Liu, W. (2021). Human-like controllable image captioning with verb-specific semantic roles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16846–16856.
    [CrossRef] [Google Scholar]
  13. Yang, X., Yang, Y., Ma, S., Li, Z., Dong, W., & Woz´niak, M. (2024). SAMT-generator: A second-attention for image captioning based on multi-stage transformer network. Neurocomputing, 593, 127823.
    [CrossRef] [Google Scholar]
  14. Yang, X., Tang, K., Zhang, H., & Cai, J. (2019). Auto-encoding scene graphs for image captioning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 10685-10694).
    [CrossRef] [Google Scholar]
  15. Wang, Q., Deng, H., Wu, X., Yang, Z., Liu, Y., Wang, Y., & Hao, G. (2023). LCM-Captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text. Neural Networks, 162, 318–329.
    [CrossRef] [Google Scholar]
  16. Nag, I. (2021). Systematic literature review of deep visual and audio captioning (Master’s thesis, Tampere University). Tampere University Institutional Repository. Available at: https://trepo.tuni.fi/handle/10024/133909
    [Google Scholar]
  17. Li, X., Xu, C., Wang, X., Lan, W., Jia, Z., Yang, G., & Xu, J. (2019). COCO-CN for cross-lingual image tagging, captioning, and retrieval. IEEE Transactions on Multimedia, 21(9), 2347–2360.
    [CrossRef] [Google Scholar]
  18. Khubchandani, V. (2024). Image caption generator using DenseNet201 and ResNet50. International Journal of Future Computer and Communication, 13(3), 55–59.
    [CrossRef] [Google Scholar]
  19. Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2016). Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 652–663.
    [CrossRef] [Google Scholar]
  20. Jaiswal, S., Pallthadka, H., Chinchewadi, P., & Jaiswal, T. (2023). An extensive analysis of image captioning models, evaluation measures, and datasets. International Journal of Multidisciplinary Science Research Review, 1(1), 21–37.
    [Google Scholar]

Cited By (11)

  1. Gaurav Dhiman, Kiran Deep Singh, Prabh Deep Singh, Norah Saleh Alghamdi, Ghadah Shukri Albakri. A novel approach to reliable and flexible distributed computing with virtualization in smart healthcare applications. Scientific Reports, 2026 , 16 (1).
    [CrossRef]
  2. Yiding Zhang, Zonghuan Han, Jian Chen, Chang Xu, Yuanze Qin, Bo Wang, Xiaoshuan Zhang, Lingxian Zhang. CropGPT: A large multimodal model for precise and explainable diagnosis of crop pests and diseases. Journal of Industrial Information Integration, 2026 , 51 .
    [CrossRef]
  3. Bin Sun. Lightweight GIS-based large-scale urban fire spread simulation method. Journal of Industrial Information Integration, 2026 , 50 .
    [CrossRef]
  4. İlknur Dönmez, Faruk Bulut. Multidimensional Diversity In Video Recommender Systems: A Holistic Framework of Literature Gaps and Future Directions. Intelligent Systems with Applications, 2026 .
    [CrossRef]
  5. Yingying Jiao. Research on an interactive film and entertainment style generation system based on multimodal input and AI models. Entertainment Computing, 2026 , 57 .
    [CrossRef]
  6. Shubhani Aggarwal, Arzoo Miglani, Norah Saleh Alghamdi, Gaurav Dhiman. Resilient and decentralized demand-side management in smart grids using blockchain. Scientific Reports, 2026 , 16 (1).
    [CrossRef]
  7. Ashok Singh Bhandari, Nitin Uniyal, Sukhveer Singh, Abhishek Sharma, Norah Saleh Alghamdi, Gaurav Dhiman. Heterogeneous Component Mixing With Cold Standby for Optimising Reliability and Redundancy. Expert Systems, 2026 , 43 (3).
    [CrossRef]
  8. Beihua Yang, Peng Song, Yunpeng Zeng. COALN-MvC: A continuous optimized anchor learning network for multi-view clustering. Knowledge-Based Systems, 2026 , 345 .
    [CrossRef]
  9. Nidhi B. Shah, Amit P. Ganatra. Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning. Journal of Innovative Image Processing, 2026 , 8 (3).
    [CrossRef]
  10. Niva Tripathy, Sampa Sahoo, Norah Saleh Alghamdi, Wattana Viriyasitavat, Gaurav Dhiman. Energy and makespan optimised task mapping in fog enabled IoT application: a hybrid approach. Scientific Reports, 2026 , 16 (1).
    [CrossRef]
  11. Praveen Kumar Tripathi, Shambhu Bharadwaj. . 2025 14th International Conference on System Modeling & Advancement in Research Trends (SMART), 2025 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Khan, A., & Singh, J. (2025). A Novel Image Captioning Technique Using Deep Learning Methodology. ICCK Transactions on Machine Intelligence, 1(2), 52–68. https://doi.org/10.62762/TMI.2025.886122
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Khan, Abdullah
AU  - Singh, Jaswinder
PY  - 2025
DA  - 2025/08/01
TI  - A Novel Image Captioning Technique Using Deep Learning Methodology
JO  - ICCK Transactions on Machine Intelligence
T2  - ICCK Transactions on Machine Intelligence
JF  - ICCK Transactions on Machine Intelligence
VL  - 1
IS  - 2
SP  - 52
EP  - 68
DO  - 10.62762/TMI.2025.886122
UR  - https://www.icck.org/article/abs/TMI.2025.886122
KW  - convolutional neural networks (CNN)
KW  - recurrent neural networks (RNN)
KW  - deep learning
KW  - image captioning
KW  - LSTM
KW  - DenseNet201
KW  - attention mechanism
KW  - BLEU scor
KW  - natural language processing (NLP)
KW  - multimodal learning
KW  - content retrieval
AB  - The capacity of AI systems to generate captions for images autonomously represents a significant advancement in artificial intelligence and language understanding. This paper presents an advanced image captioning system that employs deep learning techniques, particularly convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to produce contextually appropriate and meaningful descriptions of visual content. The proposed method extracts features using the DenseNet201 model, enabling a more comprehensive and hierarchical understanding of image components. These extracted features are then fed into a long short-term memory (LSTM) network, a specialized RNN variant designed to capture sequential dependencies in language, yielding coherent and fluent captions. The model is trained and evaluated on the well-known Flickr8k dataset, achieving competitive performance as measured by BLEU score metrics and demonstrating its ability to generate human-like descriptions. This integration of CNNs and RNNs highlights the effectiveness of combining computer vision and natural language processing for automated caption generation. The approach has potential applications across various domains, including assistive technologies for the visually impaired, automated content creation for digital media, enhanced indexing and retrieval of multimedia assets, and improved human-computer interaction. Furthermore, advancements in attention mechanisms and transformer-based models present opportunities to further enhance the accuracy and contextual relevance of image captioning systems. The study underscores the broader implications of machine-generated captions for improving accessibility, boosting searchability in large-scale databases, and enabling seamless AI-human collaboration in content interpretation and storytelling.
SN  - 3068-7403
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Khan2025A,
  author = {Abdullah Khan and Jaswinder Singh},
  title = {A Novel Image Captioning Technique Using Deep Learning Methodology},
  journal = {ICCK Transactions on Machine Intelligence},
  year = {2025},
  volume = {1},
  number = {2},
  pages = {52-68},
  doi = {10.62762/TMI.2025.886122},
  url = {https://www.icck.org/article/abs/TMI.2025.886122},
  abstract = {The capacity of AI systems to generate captions for images autonomously represents a significant advancement in artificial intelligence and language understanding. This paper presents an advanced image captioning system that employs deep learning techniques, particularly convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to produce contextually appropriate and meaningful descriptions of visual content. The proposed method extracts features using the DenseNet201 model, enabling a more comprehensive and hierarchical understanding of image components. These extracted features are then fed into a long short-term memory (LSTM) network, a specialized RNN variant designed to capture sequential dependencies in language, yielding coherent and fluent captions. The model is trained and evaluated on the well-known Flickr8k dataset, achieving competitive performance as measured by BLEU score metrics and demonstrating its ability to generate human-like descriptions. This integration of CNNs and RNNs highlights the effectiveness of combining computer vision and natural language processing for automated caption generation. The approach has potential applications across various domains, including assistive technologies for the visually impaired, automated content creation for digital media, enhanced indexing and retrieval of multimedia assets, and improved human-computer interaction. Furthermore, advancements in attention mechanisms and transformer-based models present opportunities to further enhance the accuracy and contextual relevance of image captioning systems. The study underscores the broader implications of machine-generated captions for improving accessibility, boosting searchability in large-scale databases, and enabling seamless AI-human collaboration in content interpretation and storytelling.},
  keywords = {convolutional neural networks (CNN), recurrent neural networks (RNN), deep learning, image captioning, LSTM, DenseNet201, attention mechanism, BLEU scor, natural language processing (NLP), multimodal learning, content retrieval},
  issn = {3068-7403},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
11
Scopus
10
Views
5517
PDF Downloads
541

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Machine Intelligence
ICCK Transactions on Machine Intelligence
ISSN: 3068-7403 (Online)
Portico
Preserved at
Portico