Beyond the black box: Explainable deepfake detection through prosodic analysis of the Spanish discourse marker «¿no?»

Abstract

The increasing sophistication of audio deepfake generation by current algorithms compounds the difficulty of developing models capable of distinguishing between human and synthetic speech, one of the foremost challenges in contemporary forensic phonetics. Despite the widespread use of end-to-end detection systems, handcrafted feature-based models remain valuable due to their greater explainability and reduced opacity. The present work explores the use of high-level acoustic-prosodic features (mean F0, dynamic F0 contour, and duration) as linguistically interpretable correlates that allow, on the one hand, an assessment of the importance of incorporating the suprasegmental component into deepfake detection models and, on the other, the interpretation of results on forensic grounds. Focusing on the Spanish discourse marker ¿no?, we extracted 14 prosodic variables from 272 tokens drawn from the VoxCeleb-ESP spontaneous speech corpus and their zero-shot cloned counterparts generated with the Qwen TTS system. Following Elastic Net feature selection, we trained six supervised machine learning classifiers and demonstrate that a Random Forest model achieves an AUC of 0.948 and an F1-score of 87.8% on an independent validation set. The results show that cloned voices systematically exaggerate the final portion of the marker’s intonation contour, producing a markedly higher pitch target. Furthermore, statistical models reveal that the duration of cloned audio samples is significantly longer (p < 0.009) than that of their bonafide counterparts. These findings are attributed to a read-speech bias in the cloning model’s training data. It is therefore concluded that further exploration of discourse markers and suprasegmental variables in forensic deepfake detection contexts is well warranted

Type
Publication
Normas. Revista de estudios lingüísticos hispánicos, 16 (1), pp. 1–20