
Slide

Centre Interdisciplinaire
de Recherche et d’Innovation
en Cybersécurité et Société
de Recherche et d’Innovation
en Cybersécurité et Société
1.
Jalleli, O.; Zhu, Y.; Falk, T. H.
Audio-Visual Cross-Attention for Improved Deepfake Video Detection and Forgery Localization Article d'actes
Dans: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern., p. 75–79, Institute of Electrical and Electronics Engineers Inc., 2025, ISBN: 1062922X (ISSN); 979-833153358-8 (ISBN), (Journal Abbreviation: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.).
Résumé | Liens | BibTeX | Étiquettes: Artificial intelligence, Audio acoustics, Audio signal processing, Audio-visual, Computer vision, Crossmodal attention, Deepfake detection, Forgery, Generative AI, Localisation, Multi-modal, Video detection, Video forgeries, Visual modalities
@inproceedings{jalleliAudioVisualCrossAttentionImproved2025,
title = {Audio-Visual Cross-Attention for Improved Deepfake Video Detection and Forgery Localization},
author = {O. Jalleli and Y. Zhu and T. H. Falk},
url = {https://www.scopus.com/pages/publications/105033150623?origin=resultslist},
doi = {10.1109/SMC58881.2025.11342953},
isbn = {1062922X (ISSN); 979-833153358-8 (ISBN)},
year = {2025},
date = {2025-01-01},
booktitle = {Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.},
pages = {75–79},
publisher = {Institute of Electrical and Electronics Engineers Inc.},
abstract = {With the emergence of multi-modal generative models, synthesized videos are becoming increasingly realistic, making the detection of deepfakes extremely challenging. While several video deepfake detection models have shown promising performance, their focus has been primarily on the visual modality. To overcome this limitation, we propose a dualstream framework that fuses visual and auditory information via cross-attention computed between embeddings extracted from pre-trained video and audio encoders. Additionally, we design a weakly-supervised forgery localization head that infers frame-level forgery scores from coarse segment-level labels, minimizing the need for fine-grained annotations and allowing for forgery location characterization. In this paper, we describe our preliminary results showing the proposed model outperforming state-of-the-art detectors on both frame-level localization and sequence-level deepfake detection tasks. Ongoing work focuses on investigating the complementarity between the visual and auditory modalities to improve model robustness and explainability. © 2025 IEEE.},
note = {Journal Abbreviation: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.},
keywords = {Artificial intelligence, Audio acoustics, Audio signal processing, Audio-visual, Computer vision, Crossmodal attention, Deepfake detection, Forgery, Generative AI, Localisation, Multi-modal, Video detection, Video forgeries, Visual modalities},
pubstate = {published},
tppubtype = {inproceedings}
}
With the emergence of multi-modal generative models, synthesized videos are becoming increasingly realistic, making the detection of deepfakes extremely challenging. While several video deepfake detection models have shown promising performance, their focus has been primarily on the visual modality. To overcome this limitation, we propose a dualstream framework that fuses visual and auditory information via cross-attention computed between embeddings extracted from pre-trained video and audio encoders. Additionally, we design a weakly-supervised forgery localization head that infers frame-level forgery scores from coarse segment-level labels, minimizing the need for fine-grained annotations and allowing for forgery location characterization. In this paper, we describe our preliminary results showing the proposed model outperforming state-of-the-art detectors on both frame-level localization and sequence-level deepfake detection tasks. Ongoing work focuses on investigating the complementarity between the visual and auditory modalities to improve model robustness and explainability. © 2025 IEEE.



