| [1] 冯畅,吴晓龙,赵熠扬,等. 生成式伪造语音安全问题与解决方案[J]. 信息安全研究,2024,10(2):122-129.
[2] 许旻辰,屈丹,司念文,等. 社交媒体虚假信息检测技术研究综述[EB/OL]. 2025[2026-07-12]. https://link.cnki.net/doi/10.19678/j.issn.1000-3428.0070287.
[3] Ren Y, Ruan Y, Tan X, et al. FastSpeech: fast, robust and controllable text to speech[EB/OL]. 2019[2026-07-12]. https://arxiv.org/pdf/1905.09263.
[4] Ren Y, Hu C, Tan X, et al. Fastspeech 2: Fast and High-Quality End-To-End Text To Speech[EB/OL]. 2020[2026-07-12]. https://arxiv.org/pdf/2006.04558.
[5] Wu D Y, Lee H Y. One-Shot Voice Conversion by Vector Quantization[C]//ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway, NJ: IEEE, 2020: 7734-7738.
[6] Xu L, Zhong R, Liu Y, et al. Flow-VAE VC: end-to-end flow framework with contrastive loss for zero-shot voice conversion[C]//Interspeech 2023. Baixas, France: ISCA, 2023: 2293-2297.
[7] Yamagishi J, Veaux C, MacDonald K. CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92)[EB/OL]. 2017[2026-07-12]. https://datashare.ed.ac.uk/items/7f6ee35e-2626-4b13-b709-947579036792.
[8] Frank J , Schnherr L. Wavefake: A data set to facilitate audio deepfake detection[EB/OL]. 2021[2026-07-12]. https://arxiv.org/pdf/2111.02813.
[9] Yi J, Tao J, Fu R, et al. Add 2023: the second audio deepfake detection challenge[EB/OL]. 2024[2026-07-12]. https://arxiv.org/pdf/2408.04967.
[10] Müller N M, Czempin P, Dieckmann F, et al. Does audio deepfake detection generalize?[EB/OL]. 2022[2026-07-12]. https://arxiv.org/pdf/2203.16263.
[11] Zhang L, Wang X, Cooper E, et al. The partialspoof database and countermeasures for the detection of short fake speech segments embedded in an utterance[EB/OL]. 2022[2026-07-12]. https://ieeexplore.ieee.org/document/10003971.
[12] Yi J, Fu R, Tao J, et al. Add 2022: the first audio deep synthesis detection challenge[C]//ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway, NJ: IEEE, 2022: 9216-9220.
[13] Yi J, Bai Y, Tao J, et al. Half-truth: A partially fake audio detection dataset[EB/OL]. 2022[2026-07-12]. https://arxiv.org/pdf/2104.03617.
[14] Zhang B, Cui H, Nguyen V, et al. Audio deepfake detection: What has been achieved and what lies ahead[J]. Sensors, 2025, 25(7): 1989.
[15] Wu Z, Kinnunen T, Evans N, et al. ASVspoof 2015: Automatic speaker verification spoofing and countermeasures challenge evaluation plan[EB/OL].(2014)[2026-07-12]. https://arxiv.org/pdf/1503.02907.
[16] Kinnunen T, Lee K A, Delgado H, et al. t-DCF: a detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification[EB/OL]. 2018[2026-07-12]. https://arxiv.org/pdf/1804.09618.
[17] Zhang Y, Wang W, Zhang P. The effect of silence and dual-band fusion in anti-spoofing system[C]//Interspeech 2021. Baixas, France ISCA, 2021: 4279-4283.
[18] Zhang Z, Yi X, Zhao X. Fake speech detection using residual network with transformer encoder[C]//Proceedings of the 2021 ACM workshop on information hiding and multimedia security. New York: Association for Computing Machinery, 2021: 13-22.
[19] Li M, Ahmadiadli Y, Zhang X-P. Robust deepfake audio detection via bi-level optimization[C]// 2023 IEEE 25th International Workshop on Multimedia Signal Processing (MMSP). Piscataway, NJ: IEEE, 2023: 1-6.
[20] Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0: A framework for self-supervised learning of speech representations[EB/OL]. 2020[2026-07-12]. https://arxiv.org/pdf/2006.11477.
[21] Reynolds D A, Rose R C. Robust text-independent speaker identification using Gaussian mixture speaker models[J]. IEEE Transactions on Speech and Audio Process, 1995, 3(1): 72-83.
[22] Wang C, He J, Yi J, et al. Multi-scale permutation entropy for audio deepfake detection[C]// ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Korea: IEEE, 2024: 1406-1410.
[23] Tak H, Patino J, Todisco M, et al. End-to-end anti-spoofing with rawnet2[C]//ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway, NJ: IEEE, 2021: 6369-6373.
[24] Tak H, Jung J-w, Patino J, et al. Graph Attention Networks for Anti-Spoofing[C]//Interspeech 2021. Baixas, France: ISCA, 2021: 2356-2360.
[25] Mubarak R, Alsboui T, Alshaikh O, et al. A survey on the detection and impacts of deepfakes in visual, audio, and textual formats[EB/OL]. 2023[2026-07-12]. https://ieeexplore.ieee.org/document/10365143.
[26] Cai Z, Li M. Integrating frame-level boundary detection and deepfake detection for locating manipulated regions in partially spoofed audio forgery attacks[EB/OL]. 2023[2026-07-12]. https://www.sciencedirect.com/science/article/abs/pii/S088523082300116X.
[27] Li X, Li K, Zheng Y, et al. Safeear: Content privacy-preserving audio deepfake detection[C]// Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. New York: Association for Computing Machinery, 2024: 3585-3599.
[28] Zhou Y, Lim S-N. Joint audio-visual deepfake detection[C]//Proceedings of the IEEE/CVF international conference on computer vision. Piscataway, NJ: IEEE, 2021: 14800-14809.
[29] Mittal T, Bhattacharya U, Chandra R, et al. Emotions don't lie: An audio-visual deepfake detection method using affective cues[C]//Proceedings of the 28th ACM international conference on multimedia. New York: Association for Computing Machinery, 2020: 2823-2832.
[30] Liu X, Sahidullah M, Lee K A, et al. Speaker-aware anti-spoofing[EB/OL]. 2023[2026-07-12]. https://deepfake-total.com/related_work/2303.01126.
[31] Liu R, Zhang J, Gao G, et al. Betray oneself: A novel audio deepfake detection model via mono-to-stereo conversion[EB/OL]. 2023[2026-07-12]. https://arxiv.org/abs/2305.16353. |