Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks
Research Article  ·  Published: 25 October 2024
Issue cover
ICCK Transactions on Sensing, Communication, and Control
Volume 1, Issue 1, 2024: 60-71
Research Article Free to Read

Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks

1 Beijing iQIYI Technology Co., Ltd., China
2 Department of Information Engineering, University of Padua, Italy
* Corresponding Author: Shenglun Yi, [email protected]
Volume 1, Issue 1
You have access to this article · Limited-Time Free Access

Abstract

The efficient extraction and fusion of video features to accurately identify complex and similar actions has consistently remained a significant research endeavor in the field of video action recognition. While adept at feature extraction, prevailing methodologies for video action recognition frequently exhibit suboptimal performance in the context of complex scenes and similar actions. This shortcoming arises primarily from their reliance on uni-dimensional feature extraction, thereby overlooking the interrelations among features and the significance of multi-dimensional fusion. To address this issue, this paper introduces an innovative framework predicated upon a soft correlation strategy aimed at augmenting the representational capacity of features by implementing multi-level, multi-dimensional feature aggregation and concatenating the temporal features produced by the network. Our end-to-end multi-feature encoding soft correlation concatenation aggregation layer, situated at the temporal feature output terminal of the Video Action Recognition network, proficiently aggregates and integrates the output temporal features. This approach culminates in producing a composite feature that cohesively unifies multi-dimensional information, markedly enhancing the network's competency in differentiating analogous video actions. Empirical findings demonstrate that the approach delineated in this paper bolsters the efficacy of video action recognition networks, achieving a more comprehensive characterization of spatiotemporal features, and yielding superior accuracy and robustness.

Graphical Abstract

Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks

Keywords

video action recognition soft correlation spatio-temporal feature extraction concatenation aggregation structure bidirectional LSTM

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

Fafa Wang is affiliated with the Beijing iQIYI Technology Co., Ltd., China. The authors declare that this affiliation had no influence on the study design, data collection, analysis, interpretation, or the decision to publish. Shenglun Yi served as an Associate Editor of the ICCK Transactions on Sensing, Communication, and Control at the time of manuscript submission. To ensure the integrity of the peer-review process, Shenglun Yi was not involved in the editorial handling, peer review, or decision-making process for this manuscript, which was handled independently by another editor. No other competing interests are declared.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Kong, Y., & Fu, Y. (2022). Human action recognition and prediction: A survey. International Journal of Computer Vision, 130(5), 1366-1401.
    [CrossRef] [Google Scholar]
  2. Pareek, P., & Thakkar, A. (2021). A survey on video-based Human Action Recognition: recent updates, datasets, challenges, and applications. Artificial Intelligence Review, 54(3), 2259-2322..
    [CrossRef] [Google Scholar]
  3. Georgiou, T., Liu, Y., Chen, W., & Lew, M. (2020). A survey of traditional and deep learning-based feature descriptors for high dimensional data in computer vision. International Journal of Multimedia Information Retrieval, 9, 135-170.
    [CrossRef] [Google Scholar]
  4. Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van Gool, L. (2016). Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision (pp. 20-36). Springer, Cham.
    [CrossRef] [Google Scholar]
  5. Yang, Z., An, G., Zhang, R., Zheng, Z., & Ruan, Q. (2023). SRI3D: Two‐stream inflated 3D ConvNet based on sparse regularization for action recognition. IET Image Processing, 17(5), 1438-1448.
    [CrossRef] [Google Scholar]
  6. Tran, D., Bourdev, L., Fergus, R., Torresani, L., & Paluri, M. (2015). Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision (pp. 4489-4497).
    [CrossRef] [Google Scholar]
  7. Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6299-6308).
    [CrossRef] [Google Scholar]
  8. Feichtenhofer, C., Fan, H., Malik, J., & He, K. (2019). Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 6202-6211).
    [CrossRef] [Google Scholar]
  9. Ji, S., Xu, W., Yang, M., & Yu, K. (2012). 3D convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence, 35(1), 221-231.
    [CrossRef] [Google Scholar]
  10. Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., & Fei-Fei, L. (2014). Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1725-1732).
    [CrossRef] [Google Scholar]
  11. Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems, 27.
    [Google Scholar]
  12. Lin, J., Gan, C., & Han, S. (2019). Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 7083-7093).
    [CrossRef] [Google Scholar]
  13. Yao, G., Lei, T., & Zhong, J. (2019). A review of convolutional-neural-network-based action recognition. Pattern Recognition Letters, 118, 14-22.
    [CrossRef] [Google Scholar]
  14. Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., & Darrell, T. (2015). Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2625-2634).
    [CrossRef] [Google Scholar]
  15. Girdhar, R., Ramanan, D., Gupta, A., Sivic, J., & Russell, B. (2017). Actionvlad: Learning spatio-temporal aggregation for action classification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 971-980).
    [CrossRef] [Google Scholar]
  16. Bertasius, G., Wang, H., & Torresani, L. (2021, July). Is space-time attention all you need for video understanding?. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 813-824). PMLR. https://proceedings.mlr.press/v139/bertasius21a.html
    [Google Scholar]
  17. Wu, C. Y., Li, Y., Mangalam, K., Fan, H., Xiong, B., Malik, J., & Feichtenhofer, C. (2022). Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition (pp. 13587-13597).
    [CrossRef] [Google Scholar]
  18. Qian, R., Meng, T., Gong, B., Yang, M. H., Wang, H., Belongie, S., & Cui, Y. (2021). Spatiotemporal contrastive video representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 6964-6974).
    [CrossRef] [Google Scholar]
  19. Liu, Z., Wang, L., Wu, W., Qian, C., & Lu, T. (2021). Tam: Temporal adaptive module for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 13708-13718).
    [CrossRef] [Google Scholar]

Cited By (9)

  1. Sangjin Park, Hyeonhak Kim, Richard Murray. International Stock Forecasting Using Ensemble Deep Graph Models and Complex Network Analysis. International Journal of Intelligent Systems, 2026 , 2026 (1).
    [CrossRef]
  2. Qiongmin Gao, Jian Yin. Spatiotemporal feature extraction of regional building energy consumption combining residual network and convolutional attention mechanism. Journal of Measurements in Engineering, 2026 .
    [CrossRef]
  3. Jianfeng Han, Xiongwei Gao, Lili Song, Jiandong Fang, Yongzhao Tao, Haixin Deng, Jie Yao. A New Algorithm for Visual Navigation in Unmanned Aerial Vehicle Water Surface Inspection. Sensors, 2025 , 25 (8).
    [CrossRef]
  4. Zhiqun Lin, Kexin Feng, Diaa Ahmed Mohamed Ahmedien. Improved generative adversarial networks model for movie dance generation. PLOS One, 2025 , 20 (5).
    [CrossRef]
  5. Long Zhao, Jinhui Su, Yusheng Zhong, Weiwei Xie, Jinya Su, Xisong Chen, Congyan Chen, Shihua Li. BeltLineNet: A Shape-Prior-Guided Lightweight Network for Real-Time Deviation Detection in Circular Pipe Conveyors. IEEE Sensors Journal, 2025 , 25 (11).
    [CrossRef]
  6. Biao Xiong, Jianshe Li, Mei Li, Zhuan Jin, Xin Luo, Tiantian Li, Zilin Yao. Medium-long-term prediction of NPP based on an improved spatio-temporal attention mechanism for convolutional long short-term memory network. International Journal of Remote Sensing, 2025 , 46 (14).
    [CrossRef]
  7. Xiaorui Guo, Jiayue Zhang. A Hybrid Sensing-Prediction Framework for Soft Robot Deformation Perception: Coupling Field Simulation With Gated Recurrent Unit. IEEE Sensors Journal, 2025 , 25 (19).
    [CrossRef]
  8. Samuel Appleby, Giacomo Bergami, Gary Ushaw. From Camera Image to Active Target Tracking: Modelling, Encoding and Metrical Analysis for Unmanned Underwater Vehicles. AI, 2025 , 6 (4).
    [CrossRef]
  9. Tae-Moon Seo, Dong-Joong Kang. Disentangled reflectance-ambient feature learning for day-night vehicle re-identification. Applied Soft Computing, 2025 , 181 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Wang, F., & Yi, S. (2024). Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks. ICCK Transactions on Sensing, Communication, and Control, 1(1), 60-71. https://doi.org/10.62762/TSCC.2024.212751
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Wang, Fafa
AU  - Yi, Shenglun
PY  - 2024
DA  - 2024/10/25
TI  - Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks
JO  - ICCK Transactions on Sensing, Communication, and Control
T2  - ICCK Transactions on Sensing, Communication, and Control
JF  - ICCK Transactions on Sensing, Communication, and Control
VL  - 1
IS  - 1
SP  - 60
EP  - 71
DO  - 10.62762/TSCC.2024.212751
UR  - https://www.icck.org/article/abs/TSCC.2024.212751
KW  - video action recognition
KW  - soft correlation
KW  - spatio-temporal feature extraction
KW  - concatenation aggregation structure
KW  - bidirectional LSTM
AB  - The efficient extraction and fusion of video features to accurately identify complex and similar actions has consistently remained a significant research endeavor in the field of video action recognition. While adept at feature extraction, prevailing methodologies for video action recognition frequently exhibit suboptimal performance in the context of complex scenes and similar actions. This shortcoming arises primarily from their reliance on uni-dimensional feature extraction, thereby overlooking the interrelations among features and the significance of multi-dimensional fusion. To address this issue, this paper introduces an innovative framework predicated upon a soft correlation strategy aimed at augmenting the representational capacity of features by implementing multi-level, multi-dimensional feature aggregation and concatenating the temporal features produced by the network. Our end-to-end multi-feature encoding soft correlation concatenation aggregation layer, situated at the temporal feature output terminal of the Video Action Recognition network, proficiently aggregates and integrates the output temporal features. This approach culminates in producing a composite feature that cohesively unifies multi-dimensional information, markedly enhancing the network's competency in differentiating analogous video actions. Empirical findings demonstrate that the approach delineated in this paper bolsters the efficacy of video action recognition networks, achieving a more comprehensive characterization of spatiotemporal features, and yielding superior accuracy and robustness.
SN  - 3068-9287
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Wang2024Spatiotemp,
  author = {Fafa Wang and Shenglun Yi},
  title = {Spatio-temporal Feature Soft Correlation Concatenation Aggregation Structure for Video Action Recognition Networks},
  journal = {ICCK Transactions on Sensing, Communication, and Control},
  year = {2024},
  volume = {1},
  number = {1},
  pages = {60-71},
  doi = {10.62762/TSCC.2024.212751},
  url = {https://www.icck.org/article/abs/TSCC.2024.212751},
  abstract = {The efficient extraction and fusion of video features to accurately identify complex and similar actions has consistently remained a significant research endeavor in the field of video action recognition. While adept at feature extraction, prevailing methodologies for video action recognition frequently exhibit suboptimal performance in the context of complex scenes and similar actions. This shortcoming arises primarily from their reliance on uni-dimensional feature extraction, thereby overlooking the interrelations among features and the significance of multi-dimensional fusion. To address this issue, this paper introduces an innovative framework predicated upon a soft correlation strategy aimed at augmenting the representational capacity of features by implementing multi-level, multi-dimensional feature aggregation and concatenating the temporal features produced by the network. Our end-to-end multi-feature encoding soft correlation concatenation aggregation layer, situated at the temporal feature output terminal of the Video Action Recognition network, proficiently aggregates and integrates the output temporal features. This approach culminates in producing a composite feature that cohesively unifies multi-dimensional information, markedly enhancing the network's competency in differentiating analogous video actions. Empirical findings demonstrate that the approach delineated in this paper bolsters the efficacy of video action recognition networks, achieving a more comprehensive characterization of spatiotemporal features, and yielding superior accuracy and robustness.},
  keywords = {video action recognition, soft correlation, spatio-temporal feature extraction, concatenation aggregation structure, bidirectional LSTM},
  issn = {3068-9287},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
10
Scopus
11
Views
3162
PDF Downloads
619

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Sensing, Communication, and Control
ICCK Transactions on Sensing, Communication, and Control
ISSN: 3068-9287 (Online) | ISSN: 3068-9279 (Print)
Portico
Preserved at
Portico