Selected attack sample · Gemini 3.1 Flash TTS
Mandarin · Male · 160 ms
The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.
Explore real-time voice-cloning attacks with synchronous detection, followed by streaming deepfake detection during live communication. Additional comparisons cover five generation models and three communication platforms.
01 / Paper overview
Speech synthesis has made voice interaction increasingly convenient, but has also heightened the risk of voice fraud. In real-time communication scenarios such as online meetings, speech deepfakes can directly influence decision-making and cause financial losses. However, existing detection methods have yet to systematically address the requirements of speech deepfake detection in real-time communication, such as streaming speech processing, real-time responsiveness, and robustness to transmission distortions. In this paper, we propose StreamFake, a low-overhead and robust framework for word-level streaming speech deepfake detection in real-time communication. We observe that streaming automatic speech recognition is already widely deployed in real-time communication and that its encoder produces rich discriminative representations. Further analysis shows no substantial optimization conflict between speech recognition and deepfake detection. Based on these insights, we unify the two tasks within a streaming framework. Specifically, StreamFake reuses the streaming encoder, aggregates word-level acoustic representations according to the temporal boundaries emitted by the content decoder, and performs synchronous authenticity classification using a lightweight detection decoder. This design introduces only 0.22M additional parameters without degrading the original speech recognition performance. Extensive experiments show that StreamFake achieves state-of-the-art performance under low-latency settings. It also consistently outperforms existing baselines on black-box voice-cloning APIs and in real-world communication environments, further demonstrating its practical value.
03 / Attack demo
StreamFake detects speech from real speakers and real-time voice clones as it unfolds. Samples span Mandarin and English, male and female speakers, and four streaming chunk sizes.
Selected attack sample · Gemini 3.1 Flash TTS
The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.
04 / Detection demo
This demo uses Gemini 3.1 Flash TTS for real-time voice cloning and Google Meet as the communication platform to show how different streaming chunk sizes affect detection performance.
Selected detection sample · Gemini 3.1 Flash TTS · Google Meet
05 / Generator demos
Compare word-level streaming detection across five black-box voice-cloning APIs under the same 160 ms configuration on Google Meet.
Selected generation sample · Google Meet · 160 ms chunk size
06 / Platform demos
We present detection demos on three real-world communication platforms, including Google Meet, Zoom, and Jitsi, using Gemini 3.1 Flash TTS and a default streaming chunk size of 160 ms.
Selected platform sample · Gemini 3.1 Flash TTS · 160 ms chunk size