Selected attack sample
Seed-VC · Mandarin · Male · 160 ms
The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.
Explore real-time voice-cloning attacks with synchronous detection, followed by streaming deepfake detection during live communication. Additional comparisons cover five generation models and three communication platforms.
01 / Paper overview
Speech synthesis has made voice interaction increasingly convenient, but has also heightened the risk of voice fraud. In real-time communication scenarios such as online meetings, speech deepfakes can directly influence decision-making and cause financial losses. However, existing detection methods have yet to systematically address the requirements of speech deepfake detection in real-time communication, such as streaming speech processing, real-time responsiveness, and robustness to transmission distortions. In this paper, we propose StreamFake, a low-overhead and robust framework for word-level streaming speech deepfake detection in real-time communication. We observe that streaming automatic speech recognition is already widely deployed in real-time communication and that its encoder produces rich discriminative representations. Further analysis shows no substantial optimization conflict between speech recognition and deepfake detection. Based on these insights, we unify the two tasks within a streaming framework. Specifically, StreamFake reuses the streaming encoder, aggregates word-level acoustic representations according to the temporal boundaries emitted by the content decoder, and performs synchronous authenticity classification using a lightweight detection decoder. This design introduces only 0.22M additional parameters without degrading the original speech recognition performance. Extensive experiments show that StreamFake achieves state-of-the-art performance under low-latency settings. It also consistently outperforms existing baselines on black-box voice-cloning APIs and in real-world communication environments, further demonstrating its practical value.
03 / Attack demo
StreamFake detects speech from real speakers and real-time voice clones as it unfolds. Samples span Mandarin and English, male and female speakers, and four streaming chunk sizes.
Selected attack sample
The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.
04 / Detection demo
Communication audio is captured from Google Meet in real time while StreamFake produces evolving word-level authenticity decisions. The samples use Gemini 3.1 Flash TTS and compare four streaming chunk sizes.
Selected detection sample · Gemini 3.1 Flash TTS · Google Meet
05 / Generator demos
Compare StreamFake across five generation methods sourced from Web voice clone under the same Google Meet setup and 160 ms streaming chunk size.
Source · Web voice clone · Google Meet · 160 ms chunk size
06 / Platform demos
Compare StreamFake using audio captured from Google Meet, Zoom, and Jitsi. The samples use Gemini 3.1 Flash TTS and a fixed 160 ms streaming chunk size.
Selected platform sample · Gemini 3.1 Flash TTS · 160 ms chunk size