Skip to main content
Log in

Spatial-temporal transformer network for protecting person-of-interest from deepfaking

  • Regular Paper
  • Published:
Multimedia Systems Aims and scope Submit manuscript

Abstract

The rampant use of forgery techniques poses a significant threat to the security of celebrities’ identities. Although current deepfake detection methods have shown effectiveness when dealing with specific public face forgery datasets, their reliability diminishes when applied to open data. Moreover, these methods are susceptible to re-compression and mainly rely on pixel-level abnormalities in forgery faces. In this study, we present a novel approach to detecting face forgery by leveraging individual speaking patterns of facial expressions and head movements. Our method utilizes potential motion patterns and inter-frame variations to effectively differentiate between fake and real videos. We propose an end-to-end dual-branch detection network, named the spatial-temporal transformer (STT), which aims to safeguard the identity of the person-of-interest (POI) from deepfaking. The STT incorporates the spatial transformer (ST) to establish the connection between facial expressions and head movements, while the temporal transformer (TT) exploits inconsistencies in facial attribute changes. Additionally, we introduce a central compression loss to enhance the detection performance. Extensive experiments are conducted to evaluate the effectiveness of the STT, and the results demonstrate its superiority over other SOTA methods in detecting forgery videos involving POIs. Furthermore, our network exhibits resilience to pixel-level re-compression perturbations, making it a robust solution in the face of evolving forgery techniques.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5

Similar content being viewed by others

Data availability

The data that support the findings of this study are available at https://github.com/cccvl/DeepPOI.

References

  1. Deepfacelab. https://github.com/iperov/DeepFaceLab/?utm_source=catalyzex.com. Accessed 2022-1

  2. Faceapp. https://apps.apple.com/gb/app/faceapp-ai-face-editor/id1180884341. Accessed 2022-1

  3. Tolosana, R., Vera-Rodriguez, R., Fierrez, J., Morales, A., Ortega-Garcia, J.: Deepfakes and beyond: a survey of face manipulation and fake detection. Inf. Fusion 64, 131–148 (2020)

    Article  MATH  Google Scholar 

  4. Zhou, P., Han, X., Morariu, V.I., Davis, L.S.: Two-stream neural networks for tampered face detection. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1831–1839. IEEE (2017)

  5. Matern, F., Riess, C., Stamminger, M.: Exploiting visual artifacts to expose deepfakes and face manipulations. In: 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW), pp. 83–92. IEEE (2019)

  6. Frank, J., Eisenhofer, T., Schönherr, L., Fischer, A., Kolossa, D., Holz, T.: Leveraging frequency analysis for deep fake image recognition. In: International Conference on Machine Learning, pp. 3247–3258. PMLR (2020)

  7. Masi, I., Killekar, A., Mascarenhas, R.M., Gurudatt, S.P., AbdAlmageed, W.: Two-branch recurrent network for isolating deepfakes in videos. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VII 16, pp. 667–684. Springer (2020)

  8. Liu, H., Li, X., Zhou, W., Chen, Y., He, Y., Xue, H., Zhang, W., Yu, N.: Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 772–781 (2021)

  9. Khan, S.A., Dai, H.: Video transformer for deepfake detection with incremental learning. In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 1821–1828 (2021)

  10. Chen, L., Zhang, Y., Song, Y., Liu, L., Wang, J.: Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18710–18719 (2022)

  11. Williams, G., Taylor, G., Smolskiy, K., Bregler, C.: Body motion analysis for multi-modal identity verification. In: 2010 20th International Conference on Pattern Recognition, pp. 2198–2201. IEEE (2010)

  12. Yang, X., Li, Y., Lyu, S.: Exposing deep fakes using inconsistent head poses. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 8261–8265. IEEE (2019)

  13. Bappy, J.H., Simons, C., Nataraj, L., Manjunath, B., Roy-Chowdhury, A.K.: Hybrid lstm and encoder–decoder architecture for detection of image forgeries. IEEE Trans. Image Process. 28(7), 3286–3300 (2019)

    Article  MathSciNet  Google Scholar 

  14. Huang, Y., Zhang, W., Wang, J.: Deep Frequent Spatial Temporal Learning for Face Anti-Spoofing. arXiv preprint arXiv:2002.03723 (2020)

  15. Qian, Y., Yin, G., Sheng, L., Chen, Z., Shao, J.: Thinking in frequency: Face forgery detection by mining frequency-aware clues. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XII, pp. 86–103. Springer (2020)

  16. Jeong, Y., Kim, D., Min, S., Joe, S., Gwon, Y., Choi, J.: Bihpf: bilateral high-pass filters for robust deepfake detection. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 48–57 (2022)

  17. Chen, S., Yao, T., Chen, Y., Ding, S., Li, J., Ji, R.: Local relation learning for face forgery detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 1081–1088 (2021)

  18. Zhang, B., Li, S., Feng, G., Qian, Z., Zhang, X.: Patch diffusion: a general module for face manipulation detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 3243–3251 (2022)

  19. Li, Y., Lyu, S.: Exposing Deepfake Videos by Detecting Face Warping Artifacts. arXiv preprint arXiv:1811.00656 (2018)

  20. Li, Y., Chang, M.-C., Lyu, S.: In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In: 2018 IEEE International Workshop on Information Forensics and Security (WIFS), pp. 1–7. IEEE (2018)

  21. Ciftci, U.A., Demir, I., Yin, L.: How do the hearts of deep fakes beat? deep fake source detection via interpreting residuals with biological signals. In: 2020 IEEE International Joint Conference on Biometrics (IJCB), pp. 1–10. IEEE (2020)

  22. Ciftci, U.A., Demir, I., Yin, L.: FakeCatcher: Detection of Synthetic Portrait Videos using Biological Signals. In: IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1 (2020). https://doi.org/10.1109/TPAMI.2020.3009287

  23. Qi, H., Guo, Q., Juefei-Xu, F., Xie, X., Ma, L., Feng, W., Liu, Y., Zhao, J.: Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 4318–4327 (2020)

  24. Chugh, K., Gupta, P., Dhall, A., Subramanian, R.: Not made for each other-audio-visual dissonance-based deepfake detection and localization. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 439–447 (2020)

  25. Zhang, J., Ni, J., Xie, H.: Deepfake videos detection using self-supervised decoupling network. In: 2021 IEEE International Conference on Multimedia and Expo (ICME), pp. 1–6. IEEE (2021)

  26. Cheng, H., Guo, Y., Wang, T., Li, Q., Ye, T., Nie, L.: Voice-Face Homogeneity Tells Deepfake. arXiv preprint arXiv:2203.02195 (2022)

  27. Cozzolino, D., Nießner, M., Verdoliva, L.: Audio-Visual Person-of-Interest Deepfake Detection. arXiv preprint arXiv:2204.03083 (2022)

  28. Zheng, Y., Bao, J., Chen, D., Zeng, M., Wen, F.: Exploring temporal coherence for more general video face forgery detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15044–15054 (2021)

  29. Gu, Z., Chen, Y., Yao, T., Ding, S., Li, J., Ma, L.: Delving into the local: Dynamic inconsistency learning for deepfake video detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 744–752 (2022)

  30. Roy, R., Joshi, I., Das, A., Dantcheva, A.: 3d cnn architectures and attention mechanisms for deepfake detection. In: Handbook of Digital Face Manipulation and Detection: From DeepFakes to Morphing Attacks, pp. 213–234. Springer, Cham (2022)

  31. de Lima, O., Franklin, S., Basu, S., Karwoski, B., George, A.: Deepfake Detection Using Spatiotemporal Convolutional Networks. arXiv preprint arXiv:2006.14749 (2020)

  32. Amerini, I., Galteri, L., Caldelli, R., Del Bimbo, A.: Deepfake video detection through optical flow based cnn. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (2019)

  33. Hu, J., Liao, X., Liang, J., Zhou, W., Qin, Z.: Finfer: Frame inference-based deepfake detection for high-visual-quality videos. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 951–959 (2022)

  34. Sabir, E., Cheng, J., Jaiswal, A., AbdAlmageed, W., Masi, I., Natarajan, P.: Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI) 3(1), 80–87 (2019)

    Google Scholar 

  35. Dong, X., Bao, J., Chen, D., Zhang, T., Zhang, W., Yu, N., Chen, D., Wen, F., Guo, B.: Protecting celebrities from deepfake with identity consistency transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9468–9478 (2022)

  36. Agarwal, S., Farid, H., Gu, Y., He, M., Nagano, K., Li, H.: Protecting world leaders against deep fakes. In: CVPR Workshops, vol. 1, p. 38 (2019)

  37. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems, vol. 30 (2017)

  38. Yan, L., Wang, F., Leng, L., Teoh, A.B.J.: Toward comprehensive and effective palmprint reconstruction attack. Pattern Recogn. 155, 110655 (2024)

    Article  Google Scholar 

  39. Pianese, A., Cozzolino, D., Poggi, G., Verdoliva, L.: Deepfake audio detection by speaker verification. In: 2022 IEEE International Workshop on Information Forensics and Security (WIFS), pp. 1–6 (2022). https://doi.org/10.1109/WIFS55849.2022.9975428

  40. Microexpression: Facial Action Coding System: Manual. Agriculture (1978)

  41. Baltrušaitis, T., Robinson, P., Morency, L.-P.: Openface: an open source facial behavior analysis toolkit. In: 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1–10. IEEE (2016)

  42. Oh, T.-H., Jaroensri, R., Kim, C., Elgharib, M., Durand, F., Freeman, W.T., Matusik, W.: Learning-based video motion magnification. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 633–648 (2018)

  43. Wu, H.-Y., Rubinstein, M., Shih, E., Guttag, J., Durand, F., Freeman, W.: Eulerian video magnification for revealing subtle changes in the world. ACM Trans. Graph. 31(4), 1–8 (2012)

    Article  MATH  Google Scholar 

  44. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)

  45. Sun, K., Liu, H., Ye, Q., Gao, Y., Liu, J., Shao, L., Ji, R.: Domain general face forgery detection by learning to weight. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 2638–2646 (2021)

  46. Chen, R., Chen, X., Ni, B., Ge, Y.: Simswap: an efficient framework for high fidelity face swapping. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 2003–2011 (2020)

  47. Prajwal, K., Mukhopadhyay, R., Namboodiri, V.P., Jawahar, C.: A lip sync expert is all you need for speech to lip generation in the wild. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 484–492 (2020)

  48. Siarohin, A., Lathuilière, S., Tulyakov, S., Ricci, E., Sebe, N.: First order motion model for image animation. In: Advances in Neural Information Processing Systems, vol. 32 (2019)

  49. Zhang, K., Zhang, Z., Li, Z., Qiao, Y.: Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process. Lett. 23(10), 1499–1503 (2016)

    Article  MATH  Google Scholar 

  50. Rossler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., Nießner, M.: Faceforensics++: learning to detect manipulated facial images. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1–11 (2019)

  51. Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., Ferrer, C.C.: The Deepfake Detection Challenge (DFDC) Dataset. arXiv preprint arXiv:2006.07397 (2020)

  52. Li, Y., Yang, X., Sun, P., Qi, H., Lyu, S.: Celeb-df: A large-scale challenging dataset for deepfake forensics. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3207–3216 (2020)

  53. Zi, B., Chang, M., Chen, J., Ma, X., Jiang, Y.-G.: Wilddeepfake: a challenging real-world dataset for deepfake detection. In: Proceedings of the 28th ACM International Conference on Multimedia, pp. 2382–2390 (2020)

  54. Korshunov, P., Marcel, S.: Deepfakes: A New Threat to Face Recognition? Assessment and Detection. arXiv preprint arXiv:1812.08685 (2018)

  55. Jiang, L., Li, R., Wu, W., Qian, C., Loy, C.C.: Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2889–2898 (2020)

  56. Kingma, D.P., Ba, J.: Adam: A Method for Stochastic Optimization. arXiv preprint arXiv:1412.6980 (2014)

  57. Wang, J., Wu, Z., Ouyang, W., Han, X., Chen, J., Jiang, Y.-G., Li, S.-N.: M2tr: Multi-modal multi-scale transformers for deepfake detection. In: Proceedings of the 2022 International Conference on Multimedia Retrieval, pp. 615–623 (2022)

Download references

Acknowledgements

This study is supported in part by the National Key Research and Development Plan of China (No. 2021YFF0901600) and the National Natural Science Foundation of China (No. 61971016).

Author information

Authors and Affiliations

Authors

Contributions

Dingyu Lu and Dongming Zhang wrote the main manuscript text and Jing Zhang and Guoqing edited the manuscript. Dongming Zhang and Zihou Liu revised the manuscript according to the reviewers’ comments. All authors reviewed the manuscript.

Corresponding author

Correspondence to Dongming Zhang.

Ethics declarations

Conflict of interest

The authors declare no conflict of interest.

Additional information

Communicated by Stefanos Vrochidis.

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Dingyu Lu: Work done during an internship at SKLCCC.

Rights and permissions

Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.

Reprints and permissions

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Lu, D., Liu, Z., Zhang, D. et al. Spatial-temporal transformer network for protecting person-of-interest from deepfaking. Multimedia Systems 31, 59 (2025). https://doi.org/10.1007/s00530-024-01655-8

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1007/s00530-024-01655-8

Keywords

Profiles

  1. Dongming Zhang
  2. Jing Zhang