
Slide

Centre Interdisciplinaire
de Recherche et d’Innovation
en Cybersécurité et Société
de Recherche et d’Innovation
en Cybersécurité et Société
1.
Moradi, A.; Falk, T. H.
Benchmarking Foundation Models for Cross-Domain Speaker Profiling Article d'actes
Dans: IEEE Conf. Artif. Intell., CAI, p. 78–84, Institute of Electrical and Electronics Engineers Inc., 2026, ISBN: 979-833156039-3 (ISBN), (Journal Abbreviation: IEEE Conf. Artif. Intell., CAI).
Résumé | Liens | BibTeX | Étiquettes: Benchmarking, Continuous speech recognition, Cross-domain, Forecasting, Formant frequency, Foundation models, Learning systems, Linguistics, Multi-attributes, Multi-task learning, Paralinguistic, Performance, Speaker identification, Speaker verification, Speech communication, State of the art, Verification task
@inproceedings{moradiBenchmarkingFoundationModels2026,
title = {Benchmarking Foundation Models for Cross-Domain Speaker Profiling},
author = {A. Moradi and T. H. Falk},
url = {https://www.scopus.com/pages/publications/105042046048?origin=resultslist},
doi = {10.1109/CAI68641.2026.11536565},
isbn = {979-833156039-3 (ISBN)},
year = {2026},
date = {2026-01-01},
booktitle = {IEEE Conf. Artif. Intell., CAI},
pages = {78–84},
publisher = {Institute of Electrical and Electronics Engineers Inc.},
abstract = {Speech conveys both linguistic and paralinguistic content. While pre-trained speech foundation models have been widely explored for linguistic tasks, such as speech recognition, and for speaker identification and verification tasks, very limited work has been done to test their usefulness for multi-attribute speaker profiling, i.e., simultaneous prediction of biological sex, age, and height from speech. In this paper, we aim to benchmark the performance of four state-of-the-art self-supervised foundation models, namely, WavLM, Wav2vec2, HuBERT, and XLSR-53, under both within- and cross-domain conditions. Each model is employed as a frozen pre-trained feature extractor, with lightweight task-specific heads trained jointly in a multi-task learning framework, and their feature extraction latency is empirically analyzed to assess practical deployment considerations. Experiments on two datasets show that WavLM achieves consistently strong within- and cross-domain performance for biological sex prediction, while Wav2vec2 and XLSR-53 exhibit more consistent performance for the age and height regression tasks in cross-domain settings. Interpretability analysis based on correlations with key acoustic features show (1) WavLM internal representations correlating highly with pitch and formant frequencies, corroborating the improved performance on biological sex prediction, and (2) XLSR-53 correlating highly with the second formant frequency, corroborating the results obtained for age and height. Overall, our analysis shows that pre-trained speech foundation models could serve as useful tools for cross-domain speaker profiling tasks. While no model stood out as a clear winner across all tested physical traits, future work could explore the use of ensemble methods for improved generalizability. © 2026 IEEE.},
note = {Journal Abbreviation: IEEE Conf. Artif. Intell., CAI},
keywords = {Benchmarking, Continuous speech recognition, Cross-domain, Forecasting, Formant frequency, Foundation models, Learning systems, Linguistics, Multi-attributes, Multi-task learning, Paralinguistic, Performance, Speaker identification, Speaker verification, Speech communication, State of the art, Verification task},
pubstate = {published},
tppubtype = {inproceedings}
}
Speech conveys both linguistic and paralinguistic content. While pre-trained speech foundation models have been widely explored for linguistic tasks, such as speech recognition, and for speaker identification and verification tasks, very limited work has been done to test their usefulness for multi-attribute speaker profiling, i.e., simultaneous prediction of biological sex, age, and height from speech. In this paper, we aim to benchmark the performance of four state-of-the-art self-supervised foundation models, namely, WavLM, Wav2vec2, HuBERT, and XLSR-53, under both within- and cross-domain conditions. Each model is employed as a frozen pre-trained feature extractor, with lightweight task-specific heads trained jointly in a multi-task learning framework, and their feature extraction latency is empirically analyzed to assess practical deployment considerations. Experiments on two datasets show that WavLM achieves consistently strong within- and cross-domain performance for biological sex prediction, while Wav2vec2 and XLSR-53 exhibit more consistent performance for the age and height regression tasks in cross-domain settings. Interpretability analysis based on correlations with key acoustic features show (1) WavLM internal representations correlating highly with pitch and formant frequencies, corroborating the improved performance on biological sex prediction, and (2) XLSR-53 correlating highly with the second formant frequency, corroborating the results obtained for age and height. Overall, our analysis shows that pre-trained speech foundation models could serve as useful tools for cross-domain speaker profiling tasks. While no model stood out as a clear winner across all tested physical traits, future work could explore the use of ensemble methods for improved generalizability. © 2026 IEEE.



