
Slide

Centre Interdisciplinaire
de Recherche et d’Innovation
en Cybersécurité et Société
de Recherche et d’Innovation
en Cybersécurité et Société
1.
Pimentel, A.; Zhu, Y.; Falk, T. H.
Partial Audio Deepfake Detection: Are We Really Detecting Synthetic Media or Just Dataset Content Biases? Article d'actes
Dans: IEEE Conf. Artif. Intell., CAI, p. 1846–1851, Institute of Electrical and Electronics Engineers Inc., 2026, ISBN: 979-833156039-3 (ISBN), (Journal Abbreviation: IEEE Conf. Artif. Intell., CAI).
Résumé | Liens | BibTeX | Étiquettes: Audio signal processing, Computational linguistics, Condition, Detection models, Information integrity, Language model, Large datasets, Learning models, Optimistics, Performance, Self-supervised learning, Speech communication, Speech models, Speech recognition, Speech signals, Synthetic media, Transcription
@inproceedings{pimentelPartialAudioDeepfake2026,
title = {Partial Audio Deepfake Detection: Are We Really Detecting Synthetic Media or Just Dataset Content Biases?},
author = {A. Pimentel and Y. Zhu and T. H. Falk},
url = {https://www.scopus.com/pages/publications/105042045476?origin=resultslist},
doi = {10.1109/CAI68641.2026.11536622},
isbn = {979-833156039-3 (ISBN)},
year = {2026},
date = {2026-01-01},
booktitle = {IEEE Conf. Artif. Intell., CAI},
pages = {1846–1851},
publisher = {Institute of Electrical and Electronics Engineers Inc.},
abstract = {In recent years, the generation of highly realistic audio deepfakes has raised significant concerns regarding privacy and information integrity. While most research has focused on fully bonafide or spoofed speech, partial deep-fakes, where only segments of an utterance are manipulated, remain less explored. Given the nature of the task, existing datasets are relying on large language models to manipulate bonafide speech signals into partial deepfakes by altering, deleting, or replacing segments with synthetic content. These manipulations may alter the semantics and sentiment of the generated content, creating biases that can be captured by deepfake detection models, leading to overly optimistic performance and poor generalizability. In this work, we investigate the presence of such biases in two popular partial deepfake datasets, namely AV-Deepfake1M and PartialEdit. We explore four self-supervised learning models, two relying only on text transcriptions (RoBERTa and RoBERTa-sentiment) and two relying on universal speech representations (WavLM Large and Wav2vec2-XLSR). Our experiments show that, while speech models achieve higher in-domain performance, they do not generalize to out-of-domain conditions. In turn, models trained on only transcribed speech can effectively distinguish manipulated content, achieving up to 97.8% AUC in-domain and nearly 70% AUC out-of-domain. These results suggest that linguistic patterns and dataset confounds may indeed be biasing partial deepfake detection models, leading to poor generalizability. © 2026 IEEE.},
note = {Journal Abbreviation: IEEE Conf. Artif. Intell., CAI},
keywords = {Audio signal processing, Computational linguistics, Condition, Detection models, Information integrity, Language model, Large datasets, Learning models, Optimistics, Performance, Self-supervised learning, Speech communication, Speech models, Speech recognition, Speech signals, Synthetic media, Transcription},
pubstate = {published},
tppubtype = {inproceedings}
}
In recent years, the generation of highly realistic audio deepfakes has raised significant concerns regarding privacy and information integrity. While most research has focused on fully bonafide or spoofed speech, partial deep-fakes, where only segments of an utterance are manipulated, remain less explored. Given the nature of the task, existing datasets are relying on large language models to manipulate bonafide speech signals into partial deepfakes by altering, deleting, or replacing segments with synthetic content. These manipulations may alter the semantics and sentiment of the generated content, creating biases that can be captured by deepfake detection models, leading to overly optimistic performance and poor generalizability. In this work, we investigate the presence of such biases in two popular partial deepfake datasets, namely AV-Deepfake1M and PartialEdit. We explore four self-supervised learning models, two relying only on text transcriptions (RoBERTa and RoBERTa-sentiment) and two relying on universal speech representations (WavLM Large and Wav2vec2-XLSR). Our experiments show that, while speech models achieve higher in-domain performance, they do not generalize to out-of-domain conditions. In turn, models trained on only transcribed speech can effectively distinguish manipulated content, achieving up to 97.8% AUC in-domain and nearly 70% AUC out-of-domain. These results suggest that linguistic patterns and dataset confounds may indeed be biasing partial deepfake detection models, leading to poor generalizability. © 2026 IEEE.



