StreamFake: Word-level Streaming Speech Deepfake Detection

Explore real-time voice-cloning attacks with synchronous detection, followed by streaming deepfake detection during live communication. Additional comparisons cover five generation models and three communication platforms.

Teaser: Gemini 3.1 Flash TTS · English · Female · Google Meet · 1920 ms streaming chunk size.

01 / Paper overview

Abstract

Speech synthesis has made voice interaction increasingly convenient, but has also heightened the risk of voice fraud. In real-time communication scenarios such as online meetings, speech deepfakes can directly influence decision-making and cause financial losses. However, existing detection methods have yet to systematically address the requirements of speech deepfake detection in real-time communication, such as streaming speech processing, real-time responsiveness, and robustness to transmission distortions. In this paper, we propose StreamFake, a low-overhead and robust framework for word-level streaming speech deepfake detection in real-time communication. We observe that streaming automatic speech recognition is already widely deployed in real-time communication and that its encoder produces rich discriminative representations. Further analysis shows no substantial optimization conflict between speech recognition and deepfake detection. Based on these insights, we unify the two tasks within a streaming framework. Specifically, StreamFake reuses the streaming encoder, aggregates word-level acoustic representations according to the temporal boundaries emitted by the content decoder, and performs synchronous authenticity classification using a lightweight detection decoder. This design introduces only 0.22M additional parameters without degrading the original speech recognition performance. Extensive experiments show that StreamFake achieves state-of-the-art performance under low-latency settings. It also consistently outperforms existing baselines on black-box voice-cloning APIs and in real-world communication environments, further demonstrating its practical value.

02 / Method

Word-level streaming deepfake detection

StreamFake architecture showing the streaming speech encoder, incremental content decoder, representation aggregation, and deepfake detection decoder.
Incremental decoding supplies content events; the detector aggregates the corresponding streaming speech representations and produces word-level authenticity decisions.

03 / Attack demo

Real-time voice-cloning attack and detection

StreamFake detects speech from real speakers and real-time voice clones as it unfolds. Samples span Mandarin and English, male and female speakers, and four streaming chunk sizes.

Selected attack sample

Seed-VC · Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
A
Microphone inputOriginal captured speech
B
Cloned outputComplete generated speech

The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.

Word-level labels Real Fake

04 / Detection demo

Streaming detection in real-time communication

Communication audio is captured from Google Meet in real time while StreamFake produces evolving word-level authenticity decisions. The samples use Gemini 3.1 Flash TTS and compare four streaming chunk sizes.

Selected detection sample · Gemini 3.1 Flash TTS · Google Meet

Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
Captured communication audio Audio captured from Google Meet
Word-level labels Real Fake

05 / Generator demos

Voice-cloning models

Compare StreamFake across five generation methods sourced from Web voice clone under the same Google Meet setup and 160 ms streaming chunk size.

Source · Web voice clone · Google Meet · 160 ms chunk size

Gemini 3.1 Flash TTS Preview · Mandarin · Male

Sample 01 of 20
Language
Generation method
Voice
Preparing selected sample…
Captured communication audio Audio captured from Google Meet
Word-level labels Real Fake

06 / Platform demos

Real-time communication platforms

Compare StreamFake using audio captured from Google Meet, Zoom, and Jitsi. The samples use Gemini 3.1 Flash TTS and a fixed 160 ms streaming chunk size.

Selected platform sample · Gemini 3.1 Flash TTS · 160 ms chunk size

Google Meet · Mandarin · Male

Sample 01 of 12
Language
Platform
Voice
Preparing selected sample…
Captured communication audio Audio captured from the selected platform
Word-level labels Real Fake