Text Localization and Recognition of Chinese Characters in Natural Scenes Based on Improved Faster Region-Based Convolutional Neural Network
Article Information
Abstract
To solve the problems faced by Chinese character recognition, such as complex shapes and diverse structures, this paper adopts the Visual Geometry Group 16 (VGG-16) model for feature extraction and introduces a two-layer bidirectional Long Short-Term Memory (LSTM) network. It improves the Faster Region-based Convolutional Neural Network (Faster R-CNN) by using a Region Proposal Network (RPN) to extract candidate boxes and adjust the positions of candidate regions. The improved model, namely Faster BLSTM-CNN, is tested through three types of experiments: validation of feature extraction effectiveness, comparative analysis of the algorithm before and after improvement, and comparison with traditional recognition algorithms. Finally, an experimental comparison combining text recognition and localization is conducted. On the ReCTS and MSRA-TD500 datasets, the proposed Faster BLSTM-CNN achieves excellent performance in Chinese character localization and recognition: in natural scenes (ReCTS dataset), its total image identification rate (TIIR) reaches 81.54%, recognition precision rate is 88.14%, and inference speed is 86 ms per image. Compared with the end-to-end benchmark algorithms DETR+CRNN-RMC and CTPN+CRNN-RMC, it improves the recognition precision rate by up to 7.54%, the TIIR by up to 10.04%, and the inference speed by up to 21%. These gains are mainly attributed to the two-layer bidirectional LSTM for contextual feature modeling and the element-wise addition fusion strategy that avoids dimension explosion.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Kantipudi, M. P., Kumar, S., & Kumar Jha, A. (2021). Scene text recognition based on bidirectional LSTM and deep neural network. Computational Intelligence and Neuroscience, 2021(1), 2676780.
[CrossRef] [Google Scholar] - Ye, Q., & Doermann, D. (2014). Text detection and recognition in imagery: A survey. IEEE transactions on pattern analysis and machine intelligence, 37(7), 1480-1500.
[CrossRef] [Google Scholar] - Raisi, Z., Naiel, M. A., Fieguth, P., Wardell, S., & Zelek, J. (2021). 2D positional embedding-based transformer for scene text recognition. Journal of Computational Vision and Imaging Systems, 6(1), 1-4.
[CrossRef] [Google Scholar] - Yao, C., Bai, X., & Liu, W. (2014). A unified framework for multioriented text detection and recognition. IEEE Transactions on Image Processing, 23(11), 4737-4749.
[CrossRef] [Google Scholar] - Tong, G., Li, Y., Gao, H., Chen, H., Wang, H., & Yang, X. (2020). MA-CRNN: a multi-scale attention CRNN for Chinese text line recognition in natural scenes. International Journal on Document Analysis and Recognition (IJDAR), 23(2), 103-114.
[CrossRef] [Google Scholar] - Ren, S., He, K., Girshick, R., & Sun, J. (2016). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6), 1137-1149.
[CrossRef] [Google Scholar] - Long, S., He, X., & Yao, C. (2021). Scene text detection and recognition: The deep learning era. International Journal of Computer Vision, 129(1), 161-184.
[CrossRef] [Google Scholar] - Liao, M., Wan, Z., Yao, C., Chen, K., & Bai, X. (2020, April). Real-time scene text detection with differentiable binarization. In Proceedings of the AAAI conference on artificial intelligence (Vol. 34, No. 07, pp. 11474-11481).
[CrossRef] [Google Scholar] - Sri, M. S., Naik, B. R., & Sankar, K. J. (2021). Object detection based on Faster R-CNN. International Journal of Engineering and Advanced Technology, 10(3), 72-76.
[CrossRef] [Google Scholar] - LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436-444.
[CrossRef] [Google Scholar] - Girshick, R. (2015). Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 1440-1448).
[CrossRef] [Google Scholar] - Schuster, M., & Paliwal, K. K. (1997). Bidirectional recurrent neural networks. IEEE transactions on Signal Processing, 45(11), 2673-2681.
[CrossRef] [Google Scholar] - Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., & Ren, D. (2020). Distance-IoU loss: Faster and better learning for bounding box regression. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 34, No. 07, pp. 12993-13000).
[CrossRef] [Google Scholar] - Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
[CrossRef] [Google Scholar] - Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735-1780.
[CrossRef] [Google Scholar] - Gers, F. A., Schmidhuber, J., & Cummins, F. (2000). Learning to forget: Continual prediction with LSTM. Neural computation, 12(10), 2451-2471.
[CrossRef] [Google Scholar] - Yu, Y., Si, X., Hu, C., & Zhang, J. (2019). A review of recurrent neural networks: LSTM cells and network architectures. Neural computation, 31(7), 1235-1270.
[CrossRef] [Google Scholar] - Shi, B., Bai, X., & Belongie, S. (2017, July). Detecting oriented text in natural images by linking segments. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 3482-3490). IEEE.
[CrossRef] [Google Scholar] - Gupta, N., & Jalal, A. S. (2022). Traditional to transfer learning progression on scene text detection and recognition: a survey. Artificial Intelligence Review, 55(4), 3457-3502.
[CrossRef] [Google Scholar] - Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLO (Version 8.0.0) [Software]. Zenodo.
[CrossRef] [Google Scholar] - Tian, Z., Huang, W., He, T., He, P., & Qiao, Y. (2016, September). Detecting text in natural image with connectionist text proposal network. In European conference on computer vision (pp. 56-72). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Shi, B., Bai, X., & Yao, C. (2016). An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE transactions on pattern analysis and machine intelligence, 39(11), 2298-2304.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Li, Yuejie AU - Li, Shijun AU - Luo, Zhendong PY - 2026 DA - 2026/09/27 TI - Text Localization and Recognition of Chinese Characters in Natural Scenes Based on Improved Faster Region-Based Convolutional Neural Network JO - ICCK Transactions on Emerging Topics in Artificial Intelligence T2 - ICCK Transactions on Emerging Topics in Artificial Intelligence JF - ICCK Transactions on Emerging Topics in Artificial Intelligence VL - 3 IS - 3 SP - 202 EP - 217 DO - 10.62762/TETAI.2026.620822 UR - https://www.icck.org/article/abs/TETAI.2026.620822 KW - text localization and recognition KW - deep learning KW - faster region-based convolutional neural network KW - long short-term memory AB - To solve the problems faced by Chinese character recognition, such as complex shapes and diverse structures, this paper adopts the Visual Geometry Group 16 (VGG-16) model for feature extraction and introduces a two-layer bidirectional Long Short-Term Memory (LSTM) network. It improves the Faster Region-based Convolutional Neural Network (Faster R-CNN) by using a Region Proposal Network (RPN) to extract candidate boxes and adjust the positions of candidate regions. The improved model, namely Faster BLSTM-CNN, is tested through three types of experiments: validation of feature extraction effectiveness, comparative analysis of the algorithm before and after improvement, and comparison with traditional recognition algorithms. Finally, an experimental comparison combining text recognition and localization is conducted. On the ReCTS and MSRA-TD500 datasets, the proposed Faster BLSTM-CNN achieves excellent performance in Chinese character localization and recognition: in natural scenes (ReCTS dataset), its total image identification rate (TIIR) reaches 81.54%, recognition precision rate is 88.14%, and inference speed is 86 ms per image. Compared with the end-to-end benchmark algorithms DETR+CRNN-RMC and CTPN+CRNN-RMC, it improves the recognition precision rate by up to 7.54%, the TIIR by up to 10.04%, and the inference speed by up to 21%. These gains are mainly attributed to the two-layer bidirectional LSTM for contextual feature modeling and the element-wise addition fusion strategy that avoids dimension explosion. SN - 3068-6652 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Li2026Text,
author = {Yuejie Li and Shijun Li and Zhendong Luo},
title = {Text Localization and Recognition of Chinese Characters in Natural Scenes Based on Improved Faster Region-Based Convolutional Neural Network},
journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
year = {2026},
volume = {3},
number = {3},
pages = {202-217},
doi = {10.62762/TETAI.2026.620822},
url = {https://www.icck.org/article/abs/TETAI.2026.620822},
abstract = {To solve the problems faced by Chinese character recognition, such as complex shapes and diverse structures, this paper adopts the Visual Geometry Group 16 (VGG-16) model for feature extraction and introduces a two-layer bidirectional Long Short-Term Memory (LSTM) network. It improves the Faster Region-based Convolutional Neural Network (Faster R-CNN) by using a Region Proposal Network (RPN) to extract candidate boxes and adjust the positions of candidate regions. The improved model, namely Faster BLSTM-CNN, is tested through three types of experiments: validation of feature extraction effectiveness, comparative analysis of the algorithm before and after improvement, and comparison with traditional recognition algorithms. Finally, an experimental comparison combining text recognition and localization is conducted. On the ReCTS and MSRA-TD500 datasets, the proposed Faster BLSTM-CNN achieves excellent performance in Chinese character localization and recognition: in natural scenes (ReCTS dataset), its total image identification rate (TIIR) reaches 81.54\%, recognition precision rate is 88.14\%, and inference speed is 86 ms per image. Compared with the end-to-end benchmark algorithms DETR+CRNN-RMC and CTPN+CRNN-RMC, it improves the recognition precision rate by up to 7.54\%, the TIIR by up to 10.04\%, and the inference speed by up to 21\%. These gains are mainly attributed to the two-layer bidirectional LSTM for contextual feature modeling and the element-wise addition fusion strategy that avoids dimension explosion.},
keywords = {text localization and recognition, deep learning, faster region-based convolutional neural network, long short-term memory},
issn = {3068-6652},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico