Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network
Research Article  ·  Published: 31 July 2026
Issue cover
ICCK Transactions on Advanced Computing and Systems
Volume 2, Issue 3, 2026: 225-254
Research Article Open Access

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

1 School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing 400065, China
2 Institute of Computer Sciences and IT, The University of Agriculture, Peshawar 25000, Pakistan
* Corresponding Author: Altaf Hussain, [email protected]
Volume 2, Issue 3

Article Information

Abstract

Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5% accuracy on KTH, 93.0% accuracy on UCF-50, and 96.0% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.

Graphical Abstract

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

Keywords

human action recognition CNN AlexNet model feature evaluation LSTM real-time video processing

Data Availability Statement

The KTH, Hollywood-2, and UCF-50 datasets analyzed in this study are publicly available from their original providers. The code and trained models are available from the corresponding author upon reasonable request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

AI Use Statement

The authors declare that no generative AI was used in the preparation of this manuscript.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Aggarwal, J. K., & Ryoo, M. S. (2011). Human activity analysis: A review. Acm Computing Surveys (Csur), 43(3), 1-43.
    [CrossRef] [Google Scholar]
  2. Pareek, P., & Thakkar, A. (2021). A survey on video-based human action recognition: recent updates, datasets, challenges, and applications. Artificial Intelligence Review, 54(3), 2259-2322.
    [CrossRef] [Google Scholar]
  3. Ali, S., & Shah, M. (2008). Human action recognition in videos using kinematic features and multiple instance learning. IEEE transactions on pattern analysis and machine intelligence, 32(2), 288-303.
    [CrossRef] [Google Scholar]
  4. Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van Gool, L. (2016, September). Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision (pp. 20-36). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  5. Blank, M., Gorelick, L., Shechtman, E., Irani, M., & Basri, R. (2005, October). Actions as space-time shapes. In Tenth IEEE International Conference on Computer Vision (ICCV'05) Volume 1 (Vol. 2, pp. 1395-1402). IEEE.
    [CrossRef] [Google Scholar]
  6. Ahad, M. A. R., Tan, J., Kim, H., & Ishikawa, S. (2010, June). Action recognition by employing combined directional motion history and energy images. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops (pp. 73-78). IEEE.
    [CrossRef] [Google Scholar]
  7. Calderara, S., Cucchiara, R., & Prati, A. (2008, September). Action signature: A novel holistic representation for action recognition. In 2008 IEEE Fifth International Conference on Advanced Video and Signal Based Surveillance (pp. 121-128). IEEE.
    [CrossRef] [Google Scholar]
  8. Herath, S., Harandi, M., & Porikli, F. (2017). Going deeper into action recognition: A survey. Image and vision computing, 60, 4-21.
    [CrossRef] [Google Scholar]
  9. Weinland, D., Ronfard, R., & Boyer, E. (2011). A survey of vision-based methods for action representation, segmentation and recognition. Computer vision and image understanding, 115(2), 224-241.
    [CrossRef] [Google Scholar]
  10. Chaaraoui, A. A., Climent-Pérez, P., & Flórez-Revuelta, F. (2013). Silhouette-based human action recognition using sequences of key poses. Pattern Recognition Letters, 34(15), 1799-1807.
    [CrossRef] [Google Scholar]
  11. Olatunji, I. E. (2018, August). Human activity recognition for mobile robot. In Journal of Physics: Conference Series (Vol. 1069, No. 1, p. 012148). IOP Publishing.
    [CrossRef] [Google Scholar]
  12. Franco, A., Magnani, A., & Maio, D. (2020). A multimodal approach for human activity recognition based on skeleton and RGB data. Pattern Recognition Letters, 131, 293-299.
    [CrossRef] [Google Scholar]
  13. Davis, J. W. (2001, July). Hierarchical motion history images for recognizing human motion. In Proceedings IEEE Workshop on Detection and Recognition of Events in Video (pp. 39-46). IEEE.
    [CrossRef] [Google Scholar]
  14. Davis, J. W., & Bobick, A. F. (1997, June). The representation and recognition of human movement using temporal templates. In Proceedings of IEEE computer society conference on computer vision and pattern recognition (pp. 928-934). IEEE.
    [CrossRef] [Google Scholar]
  15. Papadopoulos, G. T., Axenopoulos, A., & Daras, P. (2014, January). Real-time skeleton-tracking-based human action recognition using kinect data. In International conference on multimedia modeling (pp. 473-483). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  16. Ji, S., Xu, W., Yang, M., & Yu, K. (2013). 3D convolutional neural networks for human action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(01), 221-231.
    [CrossRef] [Google Scholar]
  17. Dollár, P., Rabaud, V., Cottrell, G., & Belongie, S. (2005, October). Behavior recognition via sparse spatio-temporal features. In 2005 IEEE international workshop on visual surveillance and performance evaluation of tracking and surveillance (pp. 65-72). IEEE.
    [CrossRef] [Google Scholar]
  18. Vemulapalli, R., Arrate, F., & Chellappa, R. (2014, June). Human Action Recognition by Representing 3D Skeletons as Points in a Lie Group. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 588-595). IEEE.
    [CrossRef] [Google Scholar]
  19. Cippitelli, E., Gasparrini, S., Gambi, E., & Spinsante, S. (2016). A human activity recognition system using skeleton data from RGBD sensors. Computational intelligence and neuroscience, 2016(1), 4351435.
    [CrossRef] [Google Scholar]
  20. Marszalek, M., Laptev, I., & Schmid, C. (2009). Actions in context. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2929-2936). IEEE.
    [CrossRef] [Google Scholar]
  21. He, Y., Shirakabe, S., Satoh, Y., & Kataoka, H. (2016, October). Human action recognition without human. In European Conference on Computer Vision (pp. 11-17). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  22. Jain, M., Jégou, H., & Bouthemy, P. (2016). Improved motion description for action classification. Frontiers in ICT, 2, 28.
    [CrossRef] [Google Scholar]
  23. Cai, J. X., Feng, G. C., & Tang, X. (2013, September). Human action recognition using oriented holistic feature. In 2013 IEEE International Conference on Image Processing (pp. 2420-2424). IEEE.
    [CrossRef] [Google Scholar]
  24. Wang, H., & Schmid, C. (2013). Action recognition with improved trajectories. In Proceedings of the IEEE international conference on computer vision (pp. 3551-3558).
    [CrossRef] [Google Scholar]
  25. Junejo, I. N., Dexter, E., Laptev, I., & P\'{erez, P. (2011). View-independent action recognition from temporal self-similarities. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(1), 172-185.
    [CrossRef] [Google Scholar]
  26. Niebles, J. C., Wang, H., & Fei-Fei, L. (2008). Unsupervised learning of human action categories using spatial-temporal words. International journal of computer vision, 79(3), 299-318.
    [CrossRef] [Google Scholar]
  27. Ke, Y., Sukthankar, R., & Hebert, M. (2005, October). Efficient visual event detection using volumetric features. In Tenth IEEE International Conference on Computer Vision (ICCV'05) Volume 1 (Vol. 1, pp. 166-173). IEEE.
    [CrossRef] [Google Scholar]
  28. Kovashka, A., & Grauman, K. (2010, June). Learning a hierarchy of discriminative space-time neighborhood features for human action recognition. In 2010 IEEE computer society conference on computer vision and pattern recognition (pp. 2046-2053). IEEE.
    [CrossRef] [Google Scholar]
  29. Shi, L., Zhang, Y., Cheng, J., & Lu, H. (2020). Skeleton-based action recognition with multi-stream adaptive graph convolutional networks. IEEE Transactions on Image Processing, 29, 9532-9545.
    [CrossRef] [Google Scholar]
  30. Poppe, R. (2010). A survey on vision-based human action recognition. Image and vision computing, 28(6), 976-990.
    [CrossRef] [Google Scholar]
  31. Xu, Z., Zhang, J., Zhang, P., & Ding, P. (2024). Skeleton-weighted and multi-scale temporal-driven network for video action recognition. Journal of Electronic Imaging, 33(6), 063056-063056.
    [CrossRef] [Google Scholar]
  32. Li, H., Huang, J., Zhou, M., Shi, Q., & Fei, Q. (2022). Self-attention pooling-based long-term temporal network for action recognition. IEEE Transactions on Cognitive and Developmental Systems, 15(1), 65-77.
    [CrossRef] [Google Scholar]
  33. Chen, J., & Ho, C. M. (2022, January). MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 786-797). IEEE.
    [CrossRef] [Google Scholar]
  34. Noor, N., & Park, I. K. (2023, October). A Lightweight Skeleton-Based 3D-CNN for Real-Time Fall Detection and Action Recognition. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (pp. 2171-2180). IEEE.
    [CrossRef] [Google Scholar]
  35. Hong, Z., Li, Z., Zhong, S., Lyu, W., Wang, H., Ding, Y., ... & Zhang, D. (2024). Crosshar: Generalizing cross-dataset human activity recognition via hierarchical self-supervised pretraining. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 8(2), 1-26.
    [CrossRef] [Google Scholar]
  36. Ullah, A., Ahmad, J., Muhammad, K., Sajjad, M., & Baik, S. W. (2017). Action recognition in video sequences using deep bi-directional LSTM with CNN features. IEEE access, 6, 1155-1166.
    [CrossRef] [Google Scholar]
  37. Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems, 27.
    [Google Scholar]
  38. Kilic, U., Karadag, O. O., & Ozyer, G. T. (2025). AGMS-GCN: Attention-guided multi-scale graph convolutional networks for skeleton-based action recognition. Knowledge-Based Systems, 311, 113045.
    [CrossRef] [Google Scholar]
  39. Shabaninia, E., Nezamabadi-pour, H., & Shafizadegan, F. (2024). Multimodal action recognition: a comprehensive survey on temporal modeling. Multimedia Tools and Applications, 83(20), 59439-59489.
    [CrossRef] [Google Scholar]
  40. Alsarhan, T., Ali, S. S., Ganapathi, I. I., Ali, A., & Werghi, N. (2024). PH-GCN: Boosting human action recognition through multi-level granularity with pair-wise hyper GCN. IEEE Access.
    [CrossRef] [Google Scholar]
  41. Xia, L., Chen, C. C., & Aggarwal, J. K. (2012, June). View invariant human action recognition using histograms of 3d joints. In 2012 IEEE computer society conference on computer vision and pattern recognition workshops (pp. 20-27). IEEE.
    [CrossRef] [Google Scholar]
  42. Geng, T., Zheng, F., Hou, X., Lu, K., Qi, G. J., & Shao, L. (2022). Spatial-temporal pyramid graph reasoning for action recognition. IEEE Transactions on Image Processing, 31, 5484-5497.
    [CrossRef] [Google Scholar]
  43. Dong, Z., Xie, M., & Li, X. (2023). Multi-scale receptive fields convolutional network for action recognition. Applied Sciences, 13(6), 3403.
    [CrossRef] [Google Scholar]
  44. Li, J., & Yang, Z. (2024, September). Human action recognition using CNN and ViT. In 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS) (pp. 107-111). IEEE.
    [CrossRef] [Google Scholar]
  45. Wu, H., Ma, X., & Li, Y. (2023). Multi-level channel attention excitation network for human action recognition in videos. Signal Processing: Image Communication, 114, 116940.
    [CrossRef] [Google Scholar]
  46. Wu, X., Zhu, J., & Yang, L. (2024). Faster-slow network fused with enhanced fine-grained features for action recognition. Journal of Visual Communication and Image Representation, 105, 104328.
    [CrossRef] [Google Scholar]
  47. Tejonidhi, M. R., Aravinda, C. V., Kumar, S. A., Madhu, C. K., & Vinod, A. M. (2025). Optimizing group activity recognition with actor relation graphs and GCN-LSTM architectures. IEEE Access.
    [CrossRef] [Google Scholar]
  48. Liu, B., Zheng, T., Zheng, P., Liu, D., Qu, X., Gao, J., ... & Wang, X. (2023, October). Lite-mkd: A multi-modal knowledge distillation framework for lightweight few-shot action recognition. In Proceedings of the 31st ACM International Conference on Multimedia (pp. 7283-7294).
    [CrossRef] [Google Scholar]
  49. Xu, M., Xiong, Y., Chen, H., Li, X., Xia, W., Tu, Z., & Soatto, S. (2021). Long short-term transformer for online action detection. Advances in Neural Information Processing Systems, 34, 1086-1099.
    [Google Scholar]
  50. Guo, G., & Lai, A. (2014). A survey on still image based human action recognition. Pattern Recognition, 47(10), 3343-3361.
    [CrossRef] [Google Scholar]
  51. Reddy, K. K., & Shah, M. (2013). Recognizing 50 human action categories of web videos. Machine vision and applications, 24(5), 971-981.
    [CrossRef] [Google Scholar]
  52. Hussain, A. (2025). Detection and Recognition of Real-Time Violence and Human Actions Recognition in Surveillance using Lightweight MobileNet Model. ICCK Journal of Image Analysis and Processing, 1(3), 125-146.
    [CrossRef] [Google Scholar]
  53. Ji, X., & Liu, H. (2009). Advances in view-invariant human motion analysis: A review. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 40(1), 13-24.
    [CrossRef] [Google Scholar]
  54. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
    [CrossRef] [Google Scholar]
  55. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.
    [CrossRef] [Google Scholar]

Cite This Article

APA Style
Khan, H. U., & Hussain, A. (2026). Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network. ICCK Transactions on Advanced Computing and Systems, 2(3), 225-254. https://doi.org/10.62762/TACS.2025.499786
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Khan, Hidayat Ullah
AU  - Hussain, Altaf
PY  - 2026
DA  - 2026/07/31
TI  - Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network
JO  - ICCK Transactions on Advanced Computing and Systems
T2  - ICCK Transactions on Advanced Computing and Systems
JF  - ICCK Transactions on Advanced Computing and Systems
VL  - 2
IS  - 3
SP  - 225
EP  - 254
DO  - 10.62762/TACS.2025.499786
UR  - https://www.icck.org/article/abs/TACS.2025.499786
KW  - human action recognition
KW  - CNN
KW  - AlexNet model
KW  - feature evaluation
KW  - LSTM
KW  - real-time video processing
AB  - Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5% accuracy on KTH, 93.0% accuracy on UCF-50, and 96.0% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.
SN  - 3068-7969
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Khan2026Deep,
  author = {Hidayat Ullah Khan and Altaf Hussain},
  title = {Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network},
  journal = {ICCK Transactions on Advanced Computing and Systems},
  year = {2026},
  volume = {2},
  number = {3},
  pages = {225-254},
  doi = {10.62762/TACS.2025.499786},
  url = {https://www.icck.org/article/abs/TACS.2025.499786},
  abstract = {Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5\% accuracy on KTH, 93.0\% accuracy on UCF-50, and 96.0\% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.},
  keywords = {human action recognition, CNN, AlexNet model, feature evaluation, LSTM, real-time video processing},
  issn = {3068-7969},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
0
Scopus
0
Views
21
PDF Downloads
3

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Advanced Computing and Systems
ICCK Transactions on Advanced Computing and Systems
ISSN: 3068-7969 (Online)
Portico
Preserved at
Portico