COMPARATIVE EVALUATION OF DEEPFAKE DETECTION MODELS FOR SECURE REMOTE BIOMETRIC VERIFICATION
DOI:
https://doi.org/10.28925/2663-4023.2026.34.1349Keywords:
deepfake, deepfake detection, image-to-video animation, GenConViT, Swin Transformer, model benchmarking, biometric verification, cybersecurityAbstract
This study examines how remote biometric verification systems hold up against next-generation video deepfakes. Financial and digital platforms increasingly use video-based onboarding, but these pipelines are vulnerable to image-to-video animation, which generates video from a single static image without leaving the spatial splicing or blending artifacts that older detectors look for. To evaluate this threat, we developed a balanced dataset of 51 videos, consisting of 26 authentic videos and 25 synthetic videos generated via face-merging and subsequent image-to-video animation. We benchmarked two commercial SaaS platforms (TruthScan and Hive Moderation) against six open-source deep learning models: EfficientNet-B7, a CLIP-based detector, GANomaly, two Vision Transformers, and GenConViT. Classical convolutional networks such as EfficientNet-B7 did not generalize to this synthesis method: recall fell to 0.240. The hybrid GenConViT model gave the most balanced performance of the open-source systems, with an accuracy of 0.843, an F1-score of 0.825 and a specificity of 0.923. That accuracy is close to the commercial platforms, which scored between 0.882 and 0.901.
These results indicate that hybrid convolutional-transformer architectures generalize to boundary-free synthesis in a way that purely convolutional detectors do not, and that an open-source model can approach commercial detection quality without transmitting biometric data to a third party.
Downloads
References
Westerlund, M. (2019). The emergence of deepfake technology: A review. Technology Innovation Management Review, 9(11), 39–52. https://doi.org/10.22215/timreview/1282
Mirsky, Y., & Lee, W. (2021). The creation and detection of deepfakes: A survey. ACM Computing Surveys, 54(1), Article 7, 1–41. https://doi.org/10.1145/3425780
European Union Agency for Cybersecurity. (2023). ENISA threat landscape 2023. ENISA. https://www.enisa.europa.eu/publications/enisa-threat-landscape-2023
Tolosana, R., Vera-Rodriguez, R., Fierrez, J., Morales, A., & Ortega-Garcia, J. (2020). Deepfakes and beyond: A survey of face manipulation and fake detection. Information Fusion, 64, 131–148. https://doi.org/10.1016/j.inffus.2020.06.014
Pearson, J. (2022). Deepfake video of Zelenskyy could be “tip of the iceberg” in info war, experts warn. NPR. https://www.npr.org/2022/03/16/1087062648/deepfake-video-zelenskyy-experts-war-manipulation-ukraine-russia
Chen, H., & Magramo, K. (2024). Finance worker pays out $25 million after video call with deepfake “chief financial officer.” CNN. https://edition.cnn.com/2024/02/04/asia/deepfake-cfo-scam-hong-kong-intl-hnk/index.html
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27. https://papers.nips.cc/paper/5423-generative-adversarial-nets
Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1812.04948
Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33. https://arxiv.org/abs/2006.11239
Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., & Nießner, M. (2019). FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). https://arxiv.org/abs/1901.08971
Sarker, I. H. (2021). Deep cybersecurity: A comprehensive overview from neural network and deep learning perspective. SN Computer Science, 2, Article 154. https://doi.org/10.1007/s42979-021-00545-3
Verdoliva, L. (2020). Media forensics and DeepFakes: An overview. IEEE Journal of Selected Topics in Signal Processing, 14(5), 910–932. https://doi.org/10.1109/JSTSP.2020.3002101
Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., & Ferrer, C. C. (2020). The DeepFake Detection Challenge (DFDC) dataset. arXiv. https://arxiv.org/abs/2006.07397
Seferbekov, S. (2020). Deepfake Detection Challenge 1st place solution. Kaggle. https://www.kaggle.com/c/deepfake-detection-challenge
Yermakov, A., Cech, J., & Matas, J. (2025). Unlocking the hidden potential of CLIP in generalizable deepfake detection. arXiv. https://arxiv.org/abs/2503.19683
Wodajo, D., Atnafu, S., & Akhtar, Z. (2023). Deepfake video detection using generative convolutional vision transformer. arXiv. https://arxiv.org/abs/2307.07036
DeepStrike. (2025). Deepfake statistics 2025: AI fraud data & trends. https://deepstrike.io/blog/deepfake-statistics-2025
Akcay, S., Atapour-Abarghouei, A., & Breckon, T. P. (2018). GANomaly: Semi-supervised anomaly detection via adversarial training. arXiv. https://arxiv.org/abs/1805.06725
prithivMLmods. (2024). Deep-Fake-Detector-Model [Machine learning model]. Hugging Face. https://huggingface.co/prithivMLmods/Deep-Fake-Detector-Model
prithivMLmods. (2025). Deep-Fake-Detector-v2-Model [Machine learning model]. Hugging Face. https://huggingface.co/prithivMLmods/Deep-Fake-Detector-v2-Model
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Назарій Джалюк, Володимир Хома, Юрій Хома

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.