Journal of Information Security Reserach ›› 2026, Vol. 12 ›› Issue (9): 813-822.DOI: 10.12379/j.issn.2096-1057.2026.09.04

Previous Articles     Next Articles

A Comprehensive Review of Deepfake Audio Detection Methodologies and Their Applications

Qiu Xinpeng, Wang Jimin   

  • Online:2026-09-02 Published:2026-09-02

深度伪造音频检测方法及其应用研究综述

邱昕鹏,王继民   

  • 作者简介:邱昕鹏,王继民

Abstract: With the rapid advancement of generative AI technologies, speech synthesized through text to speech and voice conversion has become increasingly realistic,posing a serious threat to public safety. This paper systematically reviews deepfake audio detection methods and their applications across various domains which aids in clarifying the technical developments and application landscape within the field of deepfake audio detection, thereby providing theoretical support and practical references for public safety governance. This study conducted a comprehensive literature review based on an analysis of 70 selected publications from the China National Knowledge Infrastructure (CNKI) and Web of Science databases, covering the period from 2017 to 2025. The research systematically examined progress in deepfake audio technology, experimental datasets, evaluation metrics, and detection methodologies. Furthermore, it synthesized the application status of these detection methods across various domains, including online public opinion monitoring, user privacy protection, multimodal misinformation governance, and voice fraud detection. Findings indicate that current deepfake audio detection primarily relies on end-to-end architectures, which demonstrate superior performance yet limited interpretability, and have been successfully deployed in multiple task scenarios within public security governance. Nevertheless, challenges remain, such as insufficient diversity in deepfake audio datasets, limited generalizability and interpretability of detection techniques, and the need for broader application scenarios. Future research should prioritize the development of more diverse datasets, enhance the generalizability of detection methods, and expand into more varied application domains.

Key words: deepfake audio, audio forgery detection, applied research, research overview, public security

摘要: 随着生成式人工智能技术的快速发展,依托语音合成、语音转换等技术生成的伪造音频逼真度不断提升,已对公共安全构成严重威胁。因此,对深度伪造音频检测方法及其应用开展系统梳理与总结,有助于清晰地把握该领域的技术发展脉络与应用现状,为公共安全治理工作提供理论支撑与应用参考。本文以中国知网、Web of Science等数据库为数据源,筛选出2017~2025年间发表的70篇相关文献开展综述研究,系统梳理了深度伪造音频技术、实验数据集、评价指标以及深度伪造音频检测方法的研究进展,总结了现有检测方法在网络舆情监测、用户隐私保护、多模态虚假信息治理、语音欺诈检测等场景的应用现状。研究表明,当前深度伪造音频检测领域以端到端架构的研究为主流方向,该类方法检测性能较优但可解释性较弱,目前已在公共安全治理领域的多个任务场景实现落地应用。但该领域当前仍存在诸多问题,主要包括深度伪造音频数据集多样性不足、检测技术的泛化能力与可解释性偏弱、应用场景有待进一步拓宽等。未来研究亟须探索构建多样化数据集,持续提升检测方法的泛化性能,进一步拓展多元化应用场景。

关键词: 深度伪造音频, 音频伪造检测, 应用研究, 研究综述, 公共安全

CLC Number: