Enhanced Deepfake Detection Through Multi-Attention Mechanisms: A Comprehensive Framework for Synthetic Media Identification
Article Information
Abstract
The proliferation of deepfake technology poses significant threats to digital media authenticity, necessitating robust intelligent detection systems to combat manipulated content. This paper presents a novel attention-based framework for deepfake detection that systematically integrates multiple complementary attention mechanisms to enhance discriminative feature learning. Our approach combines spatial attention, multi-head self-attention, and channel attention modules with a VGG-16 backbone to capture comprehensive representations across different feature spaces. The spatial attention mechanism focuses on discriminative facial regions, while multi-head self-attention captures long-range spatial dependencies and global contextual relationships. Channel attention further refines feature representations by emphasizing the most informative channels for detection. Extensive experiments on FaceForensics++ and Celeb-DF datasets demonstrate the effectiveness of our progressive attention integration strategy. The proposed framework achieves competitive performance with 92.67% accuracy and 99.30% Area Under the Curve (AUC) on FF++, while maintaining solid generalization capabilities with 82.35% accuracy and 82.7% AUC on the challenging Celeb-DF dataset. Comprehensive ablation studies validate the contribution of each attention component and justify key design choices, including the optimal 3×3 kernel size for spatial attention. Comparison with state-of-the-art methods demonstrates that our approach achieves competitive detection performance while maintaining architectural simplicity and computational efficiency. The modular design of our framework provides interpretability and flexibility for deployment across various computational environments, making it suitable for practical artificial media detection applications.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Mirsky, Y., \& Lee, W. (2021). The creation and detection of deepfakes: A survey. ACM computing surveys (CSUR), 54(1), 1-41.
[CrossRef] [Google Scholar] - Mubarak, R., Alsboui, T., Alshaikh, O., Inuwa-Dutse, I., Khan, S., & Parkinson, S. (2023). A survey on the detection and impacts of deepfakes in visual, audio, and textual formats. Ieee Access, 11, 144497-144529.
[CrossRef] [Google Scholar] - Parashar, A., Rida, I., Parashar, A., & Aski, V. (2022). Protecting the privacy of face by De-Identification Pipeline Based on Deep Learning. In 2022 16th International Conference on Signal-Image Technology & Internet-Based Systems (SITIS) (pp. 409–416).
[CrossRef] [Google Scholar] - Kaur, J., Sharma, K., & Singh, M. P. (2024). Exploring the Depth: Ethical Considerations, Privacy Concerns, and Security Measures in the Era of Deepfakes. In Navigating the World of Deepfake Technology (pp. 141-165). IGI Global.
[CrossRef] [Google Scholar] - Zhang, G., Gao, M., Li, Q., Zhai, W., Zou, G., & Jeon, G. (2023). Disrupting deepfakes via union-saliency adversarial attack. IEEE Transactions on Consumer Electronics, 70(1), 2018–2026.
[CrossRef] [Google Scholar] - Whittaker, L., Mulcahy, R., Letheren, K., Kietzmann, J., & Russell-Bennett, R. (2023). Mapping the deepfake landscape for innovation: A multidisciplinary systematic review and future research agenda. Technovation, 125, 102784.
[CrossRef] [Google Scholar] - Verdoliva, L. (2020). Media forensics and deepfakes: an overview. IEEE journal of selected topics in signal processing, 14(5), 910-932.
[CrossRef] [Google Scholar] - Seow, J. W., Lim, M. K., Phan, R. C., & Liu, J. K. (2022). A comprehensive overview of Deepfake: Generation, detection, datasets, and opportunities. Neurocomputing, 513, 351-371.
[CrossRef] [Google Scholar] - Afchar, D., Nozick, V., Yamagishi, J., & Echizen, I. (2018, December). Mesonet: a compact facial video forgery detection network. In 2018 IEEE international workshop on information forensics and security (WIFS) (pp. 1-7). IEEE.
[CrossRef] [Google Scholar] - Das, A., Das, S., & Dantcheva, A. (2021, December). Demystifying attention mechanisms for deepfake detection. In 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021) (pp. 1-7). IEEE.
[CrossRef] [Google Scholar] - Agarwal, S., & Farid, H. (2021, June). Detecting Deep-Fake Videos from Aural and Oral Dynamics. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 981-989). IEEE.
[CrossRef] [Google Scholar] - Bonettini, N., Cannas, E. D., Mandelli, S., Bondi, L., Bestagini, P., & Tubaro, S. (2021, January). Video face manipulation detection through ensemble of cnns. In 2020 25th international conference on pattern recognition (ICPR) (pp. 5012-5019). IEEE.
[CrossRef] [Google Scholar] - Lee, S., Tariq, S., Kim, J., & Woo, S. S. (2021, June). Tar: Generalized forensic framework to detect deepfakes using weakly supervised learning. In IFIP International conference on ICT systems security and privacy protection (pp. 351-366). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Du, M., Pentyala, S., Li, Y., & Hu, X. (2020, October). Towards generalizable deepfake detection with locality-aware autoencoder. In Proceedings of the 29th ACM international conference on information & knowledge management (pp. 325-334).
[CrossRef] [Google Scholar] - Heidari, A., Jafari Navimipour, N., Dag, H., & Unal, M. (2024). Deepfake detection using deep learning methods: A systematic and comprehensive review. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(2), e1520.
[CrossRef] [Google Scholar] - Yoon, J. H., Panizo-LLedot, A., Camacho, D., & Choi, C. (2024). Triple-modality interaction for deepfake detection on zero-shot identity. Information Fusion, 109, 102424.
[CrossRef] [Google Scholar] - Ciftci, U. A., Demir, I., & Yin, L. (2020). Fakecatcher: Detection of synthetic portrait videos using biological signals. IEEE Transactions on Pattern Analysis and Machine Intelligence.
[CrossRef] [Google Scholar] - Güera, D., & Delp, E. J. (2018, November). Deepfake video detection using recurrent neural networks. In 2018 15th IEEE international conference on advanced video and signal based surveillance (AVSS) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Li, Y., & Lyu, S. (2018). Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656.
[Google Scholar] - Heo, Y. J., Choi, Y. J., Lee, Y. W., & Kim, B. G. (2021). Deepfake detection scheme based on vision transformer and distillation. arXiv preprint arXiv:2104.01353.
[Google Scholar] - Wang, J., Lei, J., Li, S., & Zhang, J. (2025). STA-3D: Combining Spatiotemporal Attention and 3D Convolutional Networks for Robust Deepfake Detection. Symmetry, 17(7), 1037.
[CrossRef] [Google Scholar] - Chen, H. S., Rouhsedaghat, M., Ghani, H., Hu, S., You, S., & Kuo, C. C. J. (2021, July). DefakeHop: A Light-Weight High-Performance Deepfake Detector. In 2021 IEEE International Conference on Multimedia and Expo (ICME) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Wodajo, D., & Atnafu, S. (2021). Deepfake video detection using convolutional vision transformer. arXiv preprint arXiv:2102.11126.
[Google Scholar] - Nguyen, T. T., Nguyen, Q. V. H., Nguyen, D. T., Nguyen, D. T., Huynh-The, T., Nahavandi, S., Nguyen, T. T., Pham, Q.-V., & Nguyen, C. M. (2022). Deep learning for deepfakes creation and detection: A survey. Computer Vision and Image Understanding, 223, 103525.
[CrossRef] [Google Scholar] - Prajapati, P., & Pollett, C. (2022). MRI-GAN: A Generalized Approach to Detect DeepFakes using Perceptual Image Assessment. arXiv preprint arXiv:2203.00108.
[Google Scholar] - Song, W., Guo, S., Gao, M., Li, Q., Zhu, X., & Rida, I. (2025). Deepfake detection via Feature Refinement and Enhancement Network. Image and Vision Computing, 105663.
[CrossRef] [Google Scholar] - Guan, W., Wang, W., Dong, J., & Peng, B. (2024). Improving generalization of deepfake detectors by imposing gradient regularization. IEEE Transactions on Information Forensics and Security, 19, 5345-5356.
[CrossRef] [Google Scholar] - Mittal, T., Bhattacharya, U., Chandra, R., Bera, A., & Manocha, D. (2020). Emotions don't lie: An audio-visual deepfake detection method using affective cues. In Proceedings of the 28th ACM International Conference on Multimedia (pp. 2823–2832).
[CrossRef] [Google Scholar] - Montserrat, D. M., Hao, H., Yarlagadda, S. K., Baireddy, S., Shao, R., Horváth, J., ... & Delp, E. J. (2020, June). Deepfakes Detection with Automatic Face Weighting. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 2851-2859). IEEE.
[CrossRef] [Google Scholar] - De Lima, O., Franklin, S., Basu, S., Karwoski, B., & George, A. (2020). Deepfake detection using spatiotemporal convolutional networks. arXiv preprint arXiv:2006.14749.
[Google Scholar] - Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., & Niessner, M. (2019, October). FaceForensics++: Learning to Detect Manipulated Facial Images. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 1-11). IEEE.
[CrossRef] [Google Scholar] - Li, Y., Yang, X., Sun, P., Qi, H., & Lyu, S. (2020, June). Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3204-3213). IEEE.
[CrossRef] [Google Scholar] - Haliassos, A., Vougioukas, K., Petridis, S., & Pantic, M. (2021). Lips don't lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5039–5049).
[CrossRef] [Google Scholar] - Chen, S., Yao, T., Chen, Y., Ding, S., Li, J., & Ji, R. (2021, May). Local relation learning for face forgery detection. In Proceedings of the AAAI conference on artificial intelligence (Vol. 35, No. 2, pp. 1081-1088).
[CrossRef] [Google Scholar] - Li, L., Bao, J., Zhang, T., Yang, H., Chen, D., Wen, F., & Guo, B. (2020, June). Face X-Ray for More General Face Forgery Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5000-5009). IEEE.
[CrossRef] [Google Scholar] - Sun, K., Yao, T., Chen, S., Ding, S., Li, J., & Ji, R. (2022, June). Dual contrastive learning for general face forgery detection. In Proceedings of the AAAI conference on artificial intelligence (Vol. 36, No. 2, pp. 2316-2324).
[CrossRef] [Google Scholar] - Tan, M., & Le, Q. (2019, May). Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning (pp. 6105-6114). PMLR.
[Google Scholar] - Zhao, H., Wei, T., Zhou, W., Zhang, W., Chen, D., & Yu, N. (2021, June). Multi-attentional Deepfake Detection. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2185-2194). IEEE.
[CrossRef] [Google Scholar] - Zhuang, W., Chu, Q., Tan, Z., Liu, Q., Yuan, H., Miao, C., ... & Yu, N. (2022, October). UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In European conference on computer vision (pp. 391-407). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Qian, Y., Yin, G., Sheng, L., Chen, Z., & Shao, J. (2020, August). Thinking in frequency: Face forgery detection by mining frequency-aware clues. In European conference on computer vision (pp. 86-103). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Li, D., Yang, Y., Song, Y. Z., & Hospedales, T. (2018, April). Learning to generalize: Meta-learning for domain generalization. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32, No. 1).
[CrossRef] [Google Scholar] - Chollet, F. (2017, July). Xception: Deep Learning with Depthwise Separable Convolutions. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1800-1807). IEEE.
[CrossRef] [Google Scholar] - Sun, K., Liu, H., Ye, Q., Gao, Y., Liu, J., Shao, L., & Ji, R. (2021, May). Domain general face forgery detection by learning to weight. In Proceedings of the AAAI conference on artificial intelligence (Vol. 35, No. 3, pp. 2638-2646).
[CrossRef] [Google Scholar] - Usman, M. T., Khan, H., Singh, S. K., Lee, M. Y., & Koo, J. (2024). Efficient deepfake detection via layer-frozen assisted dual attention network for consumer imaging devices. IEEE Transactions on Consumer Electronics.
[CrossRef] [Google Scholar]
Cited By (6)
-
Kislay Raj, Raja Vavekanand, Aditya Singh. Custom soft spatial attention mechanism for DeepFake detection using EfficientNet-B7.
Journal of Computer Virology and Hacking Techniques, 2026 , 22 (1).
[CrossRef] -
S. Sundaresh, Akalya A. .
2026 International Conference on Data Science for Cyber-Physical Systems Resilience using Advanced Applications (ICDCA), 2026 .
[CrossRef] -
Mohammadali Vaezi, Victor Klamert, Mugdim Bublin. Deep Learning for Process Monitoring and Defect Detection of Laser-Based Powder Bed Fusion of Polymers.
Polymers, 2026 , 18 (5).
[CrossRef] -
Sahar Ebadinezhad, Temiloluwa Oluwabukunmi Modupeola. .
2026 8th International Congress on Human-Computer Interaction, Optimization and Robotic Applications (ICHORA), 2026 .
[CrossRef] -
Zainab Ghazanfar, Haill An, Saba Ghazanfar Ali, Bin Sheng, Younhyun Jung. Deep learning techniques for otologic imaging: a systematic review of segmentation methods.
Artificial Intelligence Review, 2026 , 59 (9).
[CrossRef] -
Hessa Almaazmi, Abdulrahman Alblooshi, Manar Alkhatib. .
2026 17th Student Research Conference on Applied Computing (SRC), 2026 .
[CrossRef]
Cite This Article
TY - JOUR AU - Ali, Farhan AU - Ghazanfar, Zainab PY - 2025 DA - 2025/11/24 TI - Enhanced Deepfake Detection Through Multi-Attention Mechanisms: A Comprehensive Framework for Synthetic Media Identification JO - ICCK Transactions on Intelligent Systematics T2 - ICCK Transactions on Intelligent Systematics JF - ICCK Transactions on Intelligent Systematics VL - 2 IS - 4 SP - 248 EP - 258 DO - 10.62762/TIS.2025.756872 UR - https://www.icck.org/article/abs/TIS.2025.756872 KW - deepfake detection KW - multi-head self-attention KW - synthetic media detection KW - facial manipulation detection AB - The proliferation of deepfake technology poses significant threats to digital media authenticity, necessitating robust intelligent detection systems to combat manipulated content. This paper presents a novel attention-based framework for deepfake detection that systematically integrates multiple complementary attention mechanisms to enhance discriminative feature learning. Our approach combines spatial attention, multi-head self-attention, and channel attention modules with a VGG-16 backbone to capture comprehensive representations across different feature spaces. The spatial attention mechanism focuses on discriminative facial regions, while multi-head self-attention captures long-range spatial dependencies and global contextual relationships. Channel attention further refines feature representations by emphasizing the most informative channels for detection. Extensive experiments on FaceForensics++ and Celeb-DF datasets demonstrate the effectiveness of our progressive attention integration strategy. The proposed framework achieves competitive performance with 92.67% accuracy and 99.30% Area Under the Curve (AUC) on FF++, while maintaining solid generalization capabilities with 82.35% accuracy and 82.7% AUC on the challenging Celeb-DF dataset. Comprehensive ablation studies validate the contribution of each attention component and justify key design choices, including the optimal 3×3 kernel size for spatial attention. Comparison with state-of-the-art methods demonstrates that our approach achieves competitive detection performance while maintaining architectural simplicity and computational efficiency. The modular design of our framework provides interpretability and flexibility for deployment across various computational environments, making it suitable for practical artificial media detection applications. SN - 3068-5079 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Ali2025Enhanced,
author = {Farhan Ali and Zainab Ghazanfar},
title = {Enhanced Deepfake Detection Through Multi-Attention Mechanisms: A Comprehensive Framework for Synthetic Media Identification},
journal = {ICCK Transactions on Intelligent Systematics},
year = {2025},
volume = {2},
number = {4},
pages = {248-258},
doi = {10.62762/TIS.2025.756872},
url = {https://www.icck.org/article/abs/TIS.2025.756872},
abstract = {The proliferation of deepfake technology poses significant threats to digital media authenticity, necessitating robust intelligent detection systems to combat manipulated content. This paper presents a novel attention-based framework for deepfake detection that systematically integrates multiple complementary attention mechanisms to enhance discriminative feature learning. Our approach combines spatial attention, multi-head self-attention, and channel attention modules with a VGG-16 backbone to capture comprehensive representations across different feature spaces. The spatial attention mechanism focuses on discriminative facial regions, while multi-head self-attention captures long-range spatial dependencies and global contextual relationships. Channel attention further refines feature representations by emphasizing the most informative channels for detection. Extensive experiments on FaceForensics++ and Celeb-DF datasets demonstrate the effectiveness of our progressive attention integration strategy. The proposed framework achieves competitive performance with 92.67\% accuracy and 99.30\% Area Under the Curve (AUC) on FF++, while maintaining solid generalization capabilities with 82.35\% accuracy and 82.7\% AUC on the challenging Celeb-DF dataset. Comprehensive ablation studies validate the contribution of each attention component and justify key design choices, including the optimal 3×3 kernel size for spatial attention. Comparison with state-of-the-art methods demonstrates that our approach achieves competitive detection performance while maintaining architectural simplicity and computational efficiency. The modular design of our framework provides interpretability and flexibility for deployment across various computational environments, making it suitable for practical artificial media detection applications.},
keywords = {deepfake detection, multi-head self-attention, synthetic media detection, facial manipulation detection},
issn = {3068-5079},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico