Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network
Article Information
Abstract
Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5% accuracy on KTH, 93.0% accuracy on UCF-50, and 96.0% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Aggarwal, J. K., & Ryoo, M. S. (2011). Human activity analysis: A review. Acm Computing Surveys (Csur), 43(3), 1-43.
[CrossRef] [Google Scholar] - Pareek, P., & Thakkar, A. (2021). A survey on video-based human action recognition: recent updates, datasets, challenges, and applications. Artificial Intelligence Review, 54(3), 2259-2322.
[CrossRef] [Google Scholar] - Ali, S., & Shah, M. (2008). Human action recognition in videos using kinematic features and multiple instance learning. IEEE transactions on pattern analysis and machine intelligence, 32(2), 288-303.
[CrossRef] [Google Scholar] - Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van Gool, L. (2016, September). Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision (pp. 20-36). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Blank, M., Gorelick, L., Shechtman, E., Irani, M., & Basri, R. (2005, October). Actions as space-time shapes. In Tenth IEEE International Conference on Computer Vision (ICCV'05) Volume 1 (Vol. 2, pp. 1395-1402). IEEE.
[CrossRef] [Google Scholar] - Ahad, M. A. R., Tan, J., Kim, H., & Ishikawa, S. (2010, June). Action recognition by employing combined directional motion history and energy images. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops (pp. 73-78). IEEE.
[CrossRef] [Google Scholar] - Calderara, S., Cucchiara, R., & Prati, A. (2008, September). Action signature: A novel holistic representation for action recognition. In 2008 IEEE Fifth International Conference on Advanced Video and Signal Based Surveillance (pp. 121-128). IEEE.
[CrossRef] [Google Scholar] - Herath, S., Harandi, M., & Porikli, F. (2017). Going deeper into action recognition: A survey. Image and vision computing, 60, 4-21.
[CrossRef] [Google Scholar] - Weinland, D., Ronfard, R., & Boyer, E. (2011). A survey of vision-based methods for action representation, segmentation and recognition. Computer vision and image understanding, 115(2), 224-241.
[CrossRef] [Google Scholar] - Chaaraoui, A. A., Climent-Pérez, P., & Flórez-Revuelta, F. (2013). Silhouette-based human action recognition using sequences of key poses. Pattern Recognition Letters, 34(15), 1799-1807.
[CrossRef] [Google Scholar] - Olatunji, I. E. (2018, August). Human activity recognition for mobile robot. In Journal of Physics: Conference Series (Vol. 1069, No. 1, p. 012148). IOP Publishing.
[CrossRef] [Google Scholar] - Franco, A., Magnani, A., & Maio, D. (2020). A multimodal approach for human activity recognition based on skeleton and RGB data. Pattern Recognition Letters, 131, 293-299.
[CrossRef] [Google Scholar] - Davis, J. W. (2001, July). Hierarchical motion history images for recognizing human motion. In Proceedings IEEE Workshop on Detection and Recognition of Events in Video (pp. 39-46). IEEE.
[CrossRef] [Google Scholar] - Davis, J. W., & Bobick, A. F. (1997, June). The representation and recognition of human movement using temporal templates. In Proceedings of IEEE computer society conference on computer vision and pattern recognition (pp. 928-934). IEEE.
[CrossRef] [Google Scholar] - Papadopoulos, G. T., Axenopoulos, A., & Daras, P. (2014, January). Real-time skeleton-tracking-based human action recognition using kinect data. In International conference on multimedia modeling (pp. 473-483). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Ji, S., Xu, W., Yang, M., & Yu, K. (2013). 3D convolutional neural networks for human action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(01), 221-231.
[CrossRef] [Google Scholar] - Dollár, P., Rabaud, V., Cottrell, G., & Belongie, S. (2005, October). Behavior recognition via sparse spatio-temporal features. In 2005 IEEE international workshop on visual surveillance and performance evaluation of tracking and surveillance (pp. 65-72). IEEE.
[CrossRef] [Google Scholar] - Vemulapalli, R., Arrate, F., & Chellappa, R. (2014, June). Human Action Recognition by Representing 3D Skeletons as Points in a Lie Group. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 588-595). IEEE.
[CrossRef] [Google Scholar] - Cippitelli, E., Gasparrini, S., Gambi, E., & Spinsante, S. (2016). A human activity recognition system using skeleton data from RGBD sensors. Computational intelligence and neuroscience, 2016(1), 4351435.
[CrossRef] [Google Scholar] - Marszalek, M., Laptev, I., & Schmid, C. (2009). Actions in context. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2929-2936). IEEE.
[CrossRef] [Google Scholar] - He, Y., Shirakabe, S., Satoh, Y., & Kataoka, H. (2016, October). Human action recognition without human. In European Conference on Computer Vision (pp. 11-17). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Jain, M., Jégou, H., & Bouthemy, P. (2016). Improved motion description for action classification. Frontiers in ICT, 2, 28.
[CrossRef] [Google Scholar] - Cai, J. X., Feng, G. C., & Tang, X. (2013, September). Human action recognition using oriented holistic feature. In 2013 IEEE International Conference on Image Processing (pp. 2420-2424). IEEE.
[CrossRef] [Google Scholar] - Wang, H., & Schmid, C. (2013). Action recognition with improved trajectories. In Proceedings of the IEEE international conference on computer vision (pp. 3551-3558).
[CrossRef] [Google Scholar] - Junejo, I. N., Dexter, E., Laptev, I., & P\'{erez, P. (2011). View-independent action recognition from temporal self-similarities. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(1), 172-185.
[CrossRef] [Google Scholar] - Niebles, J. C., Wang, H., & Fei-Fei, L. (2008). Unsupervised learning of human action categories using spatial-temporal words. International journal of computer vision, 79(3), 299-318.
[CrossRef] [Google Scholar] - Ke, Y., Sukthankar, R., & Hebert, M. (2005, October). Efficient visual event detection using volumetric features. In Tenth IEEE International Conference on Computer Vision (ICCV'05) Volume 1 (Vol. 1, pp. 166-173). IEEE.
[CrossRef] [Google Scholar] - Kovashka, A., & Grauman, K. (2010, June). Learning a hierarchy of discriminative space-time neighborhood features for human action recognition. In 2010 IEEE computer society conference on computer vision and pattern recognition (pp. 2046-2053). IEEE.
[CrossRef] [Google Scholar] - Shi, L., Zhang, Y., Cheng, J., & Lu, H. (2020). Skeleton-based action recognition with multi-stream adaptive graph convolutional networks. IEEE Transactions on Image Processing, 29, 9532-9545.
[CrossRef] [Google Scholar] - Poppe, R. (2010). A survey on vision-based human action recognition. Image and vision computing, 28(6), 976-990.
[CrossRef] [Google Scholar] - Xu, Z., Zhang, J., Zhang, P., & Ding, P. (2024). Skeleton-weighted and multi-scale temporal-driven network for video action recognition. Journal of Electronic Imaging, 33(6), 063056-063056.
[CrossRef] [Google Scholar] - Li, H., Huang, J., Zhou, M., Shi, Q., & Fei, Q. (2022). Self-attention pooling-based long-term temporal network for action recognition. IEEE Transactions on Cognitive and Developmental Systems, 15(1), 65-77.
[CrossRef] [Google Scholar] - Chen, J., & Ho, C. M. (2022, January). MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 786-797). IEEE.
[CrossRef] [Google Scholar] - Noor, N., & Park, I. K. (2023, October). A Lightweight Skeleton-Based 3D-CNN for Real-Time Fall Detection and Action Recognition. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (pp. 2171-2180). IEEE.
[CrossRef] [Google Scholar] - Hong, Z., Li, Z., Zhong, S., Lyu, W., Wang, H., Ding, Y., ... & Zhang, D. (2024). Crosshar: Generalizing cross-dataset human activity recognition via hierarchical self-supervised pretraining. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 8(2), 1-26.
[CrossRef] [Google Scholar] - Ullah, A., Ahmad, J., Muhammad, K., Sajjad, M., & Baik, S. W. (2017). Action recognition in video sequences using deep bi-directional LSTM with CNN features. IEEE access, 6, 1155-1166.
[CrossRef] [Google Scholar] - Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems, 27.
[Google Scholar] - Kilic, U., Karadag, O. O., & Ozyer, G. T. (2025). AGMS-GCN: Attention-guided multi-scale graph convolutional networks for skeleton-based action recognition. Knowledge-Based Systems, 311, 113045.
[CrossRef] [Google Scholar] - Shabaninia, E., Nezamabadi-pour, H., & Shafizadegan, F. (2024). Multimodal action recognition: a comprehensive survey on temporal modeling. Multimedia Tools and Applications, 83(20), 59439-59489.
[CrossRef] [Google Scholar] - Alsarhan, T., Ali, S. S., Ganapathi, I. I., Ali, A., & Werghi, N. (2024). PH-GCN: Boosting human action recognition through multi-level granularity with pair-wise hyper GCN. IEEE Access.
[CrossRef] [Google Scholar] - Xia, L., Chen, C. C., & Aggarwal, J. K. (2012, June). View invariant human action recognition using histograms of 3d joints. In 2012 IEEE computer society conference on computer vision and pattern recognition workshops (pp. 20-27). IEEE.
[CrossRef] [Google Scholar] - Geng, T., Zheng, F., Hou, X., Lu, K., Qi, G. J., & Shao, L. (2022). Spatial-temporal pyramid graph reasoning for action recognition. IEEE Transactions on Image Processing, 31, 5484-5497.
[CrossRef] [Google Scholar] - Dong, Z., Xie, M., & Li, X. (2023). Multi-scale receptive fields convolutional network for action recognition. Applied Sciences, 13(6), 3403.
[CrossRef] [Google Scholar] - Li, J., & Yang, Z. (2024, September). Human action recognition using CNN and ViT. In 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS) (pp. 107-111). IEEE.
[CrossRef] [Google Scholar] - Wu, H., Ma, X., & Li, Y. (2023). Multi-level channel attention excitation network for human action recognition in videos. Signal Processing: Image Communication, 114, 116940.
[CrossRef] [Google Scholar] - Wu, X., Zhu, J., & Yang, L. (2024). Faster-slow network fused with enhanced fine-grained features for action recognition. Journal of Visual Communication and Image Representation, 105, 104328.
[CrossRef] [Google Scholar] - Tejonidhi, M. R., Aravinda, C. V., Kumar, S. A., Madhu, C. K., & Vinod, A. M. (2025). Optimizing group activity recognition with actor relation graphs and GCN-LSTM architectures. IEEE Access.
[CrossRef] [Google Scholar] - Liu, B., Zheng, T., Zheng, P., Liu, D., Qu, X., Gao, J., ... & Wang, X. (2023, October). Lite-mkd: A multi-modal knowledge distillation framework for lightweight few-shot action recognition. In Proceedings of the 31st ACM International Conference on Multimedia (pp. 7283-7294).
[CrossRef] [Google Scholar] - Xu, M., Xiong, Y., Chen, H., Li, X., Xia, W., Tu, Z., & Soatto, S. (2021). Long short-term transformer for online action detection. Advances in Neural Information Processing Systems, 34, 1086-1099.
[Google Scholar] - Guo, G., & Lai, A. (2014). A survey on still image based human action recognition. Pattern Recognition, 47(10), 3343-3361.
[CrossRef] [Google Scholar] - Reddy, K. K., & Shah, M. (2013). Recognizing 50 human action categories of web videos. Machine vision and applications, 24(5), 971-981.
[CrossRef] [Google Scholar] - Hussain, A. (2025). Detection and Recognition of Real-Time Violence and Human Actions Recognition in Surveillance using Lightweight MobileNet Model. ICCK Journal of Image Analysis and Processing, 1(3), 125-146.
[CrossRef] [Google Scholar] - Ji, X., & Liu, H. (2009). Advances in view-invariant human motion analysis: A review. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 40(1), 13-24.
[CrossRef] [Google Scholar] - Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
[CrossRef] [Google Scholar] - Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Khan, Hidayat Ullah AU - Hussain, Altaf PY - 2026 DA - 2026/07/31 TI - Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network JO - ICCK Transactions on Advanced Computing and Systems T2 - ICCK Transactions on Advanced Computing and Systems JF - ICCK Transactions on Advanced Computing and Systems VL - 2 IS - 3 SP - 225 EP - 254 DO - 10.62762/TACS.2025.499786 UR - https://www.icck.org/article/abs/TACS.2025.499786 KW - human action recognition KW - CNN KW - AlexNet model KW - feature evaluation KW - LSTM KW - real-time video processing AB - Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5% accuracy on KTH, 93.0% accuracy on UCF-50, and 96.0% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment. SN - 3068-7969 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2026Deep,
author = {Hidayat Ullah Khan and Altaf Hussain},
title = {Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network},
journal = {ICCK Transactions on Advanced Computing and Systems},
year = {2026},
volume = {2},
number = {3},
pages = {225-254},
doi = {10.62762/TACS.2025.499786},
url = {https://www.icck.org/article/abs/TACS.2025.499786},
abstract = {Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5\% accuracy on KTH, 93.0\% accuracy on UCF-50, and 96.0\% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.},
keywords = {human action recognition, CNN, AlexNet model, feature evaluation, LSTM, real-time video processing},
issn = {3068-7969},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico