StreamFake: Word-level Streaming Speech Deepfake Detection

Explore real-time voice-cloning attacks with synchronous detection, followed by streaming deepfake detection during live communication. Additional comparisons cover five generation models and three communication platforms.

Teaser: Gemini 3.1 Flash TTS · English · Female · Google Meet · 1920 ms streaming chunk size.

01 / Paper overview

Abstract

Speech synthesis has made voice interaction increasingly convenient, but has also heightened the risk of voice fraud. In real-time communication scenarios such as online meetings, speech deepfakes can directly influence decision-making and cause financial losses. However, existing detection methods have yet to systematically address the requirements of speech deepfake detection in real-time communication, such as streaming speech processing, real-time responsiveness, and robustness to transmission distortions. In this paper, we propose StreamFake, a low-overhead and robust framework for word-level streaming speech deepfake detection in real-time communication. We observe that streaming automatic speech recognition is already widely deployed in real-time communication and that its encoder produces rich discriminative representations. Further analysis shows no substantial optimization conflict between speech recognition and deepfake detection. Based on these insights, we unify the two tasks within a streaming framework. Specifically, StreamFake reuses the streaming encoder, aggregates word-level acoustic representations according to the temporal boundaries emitted by the content decoder, and performs synchronous authenticity classification using a lightweight detection decoder. This design introduces only 0.22M additional parameters without degrading the original speech recognition performance. Extensive experiments show that StreamFake achieves state-of-the-art performance under low-latency settings. It also consistently outperforms existing baselines on black-box voice-cloning APIs and in real-world communication environments, further demonstrating its practical value.

02 / Method

Word-level streaming deepfake detection

StreamFake architecture showing the streaming speech encoder, incremental content decoder, representation aggregation, and deepfake detection decoder.
Incremental decoding supplies content events; the detector aggregates the corresponding streaming speech representations and produces word-level authenticity decisions.

03 / Attack demo

Real-time voice-cloning attack and detection

StreamFake detects speech from real speakers and real-time voice clones as it unfolds. Samples span Mandarin and English, male and female speakers, and four streaming chunk sizes.

Selected attack sample · Gemini 3.1 Flash TTS

Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
A
Microphone inputOriginal captured speech
B
Cloned outputComplete generated speech

The audio contains the complete output after real-time voice cloning, while the video preserves the original speech captured by the microphone.

Real Fake

04 / Detection demo

Streaming detection during live communication

This demo uses Gemini 3.1 Flash TTS for real-time voice cloning and Google Meet as the communication platform to show how different streaming chunk sizes affect detection performance.

Selected detection sample · Gemini 3.1 Flash TTS · Google Meet

Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
Real Fake

05 / Generator demos

Commercial black-box voice-cloning APIs

Compare word-level streaming detection across five black-box voice-cloning APIs under the same 160 ms configuration on Google Meet.

Selected generation sample · Google Meet · 160 ms chunk size

Gemini 3.1 Flash TTS Preview · Mandarin · Male

Sample 01 of 20
Language
Voice-cloning API
Voice
Preparing selected sample…

06 / Platform demos

Commercial real-time communication platforms

We present detection demos on three real-world communication platforms, including Google Meet, Zoom, and Jitsi, using Gemini 3.1 Flash TTS and a default streaming chunk size of 160 ms.

Selected platform sample · Gemini 3.1 Flash TTS · 160 ms chunk size

Google Meet · Mandarin · Male

Sample 01 of 12
Language
Platform
Voice
Preparing selected sample…