
Slide

Centre Interdisciplinaire
de Recherche et d’Innovation
en Cybersécurité et Société
de Recherche et d’Innovation
en Cybersécurité et Société
1.
Zhu, Y.; Davoust, A.; Falk, T. H.
DeepSick: Deceiving Voice-Based Diagnostic Models with Synthetic Multilingual Pathological Speech Signals Article d'actes
Dans: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern., p. 69–74, Institute of Electrical and Electronics Engineers Inc., 2025, ISBN: 1062922X (ISSN); 979-833153358-8 (ISBN), (Journal Abbreviation: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.).
Résumé | Liens | BibTeX | Étiquettes: COVID-19, Detection models, Diagnosis, Diagnostic model, Diagnostic systems, Generative model, Health assessments, Pathological conditions, Pathological speech signals, Scalable solution, Speech communication, Speech recognition, Speech synthesis, State of the art, Voice model
@inproceedings{zhuDeepSickDeceivingVoiceBased2025,
title = {DeepSick: Deceiving Voice-Based Diagnostic Models with Synthetic Multilingual Pathological Speech Signals},
author = {Y. Zhu and A. Davoust and T. H. Falk},
url = {https://www.scopus.com/pages/publications/105033143787?origin=resultslist},
doi = {10.1109/SMC58881.2025.11343240},
isbn = {1062922X (ISSN); 979-833153358-8 (ISBN)},
year = {2025},
date = {2025-01-01},
booktitle = {Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.},
pages = {69–74},
publisher = {Institute of Electrical and Electronics Engineers Inc.},
abstract = {Voice-based diagnostic systems offer a scalable solution for remote health assessment. However, recent advances in generative voice models may enable malicious manipulation of voice samples to simulate or conceal disease-related speech characteristics, which poses new risks to diagnostic systems. This paper investigates the vulnerability of diagnostic and detection models to such types of "deepfake"attacks. We show that it is possible to train a generative model to convert between healthy voices and pathological ones, which in turn, can successfully deceive existing diagnostic systems. Here, focus is placed on COVID-19 infection and respiratory abnormalities, but the method can be applied across different pathological conditions affecting vocal attributes. We also benchmark four state-of-the-art synthesized voice detection models on both real and generated pathological speech from three datasets. Our results show that current synthetic voice detectors, typically trained on healthy speech data, perform poorly on generated pathological samples. While fine-tuning with real pathological voices improves detection, a substantial performance gap remains. This work provides initial insights on an emerging threat to remote voice diagnostic systems that needs further work. © 2025 IEEE.},
note = {Journal Abbreviation: Conf. Proc. IEEE Int. Conf. Syst. Man Cybern.},
keywords = {COVID-19, Detection models, Diagnosis, Diagnostic model, Diagnostic systems, Generative model, Health assessments, Pathological conditions, Pathological speech signals, Scalable solution, Speech communication, Speech recognition, Speech synthesis, State of the art, Voice model},
pubstate = {published},
tppubtype = {inproceedings}
}
Voice-based diagnostic systems offer a scalable solution for remote health assessment. However, recent advances in generative voice models may enable malicious manipulation of voice samples to simulate or conceal disease-related speech characteristics, which poses new risks to diagnostic systems. This paper investigates the vulnerability of diagnostic and detection models to such types of "deepfake"attacks. We show that it is possible to train a generative model to convert between healthy voices and pathological ones, which in turn, can successfully deceive existing diagnostic systems. Here, focus is placed on COVID-19 infection and respiratory abnormalities, but the method can be applied across different pathological conditions affecting vocal attributes. We also benchmark four state-of-the-art synthesized voice detection models on both real and generated pathological speech from three datasets. Our results show that current synthetic voice detectors, typically trained on healthy speech data, perform poorly on generated pathological samples. While fine-tuning with real pathological voices improves detection, a substantial performance gap remains. This work provides initial insights on an emerging threat to remote voice diagnostic systems that needs further work. © 2025 IEEE.



