Qwen3.8-LiveTranslate targets real-time interpretation. It uses an interleaved process for incoming audio, generated translation, and spoken output, adding live speaker separation, synchronized bilingual display, and longer-context disambiguation. Qwen reports average lag falling from 2.8 seconds in the prior generation to 2.3 seconds on its evaluation. That figure belongs to specific tests and does not guarantee latency for every network and language pair. Read Qwen’s architecture and language details.

What is new
The model supports 60 languages for audio input and text output, with speech output in 29 languages. In a multi-speaker conversation, it attempts to attribute each sentence to a speaker and preserve that person’s timbre more consistently in synthesized speech. Original text and translation can appear together for checking. Longer context helps resolve names, pronouns, and repeated terminology using earlier turns.
These features may help with remote meetings, interviews, lessons, and bilingual events, but automatic diarization can still be wrong. Overlapping speech, uneven microphone distance, and echo may cause attribution errors. Voice-cloning features require the speaker’s permission; tell participants whether their audio will be retained.
Build a meeting test
- Use an authorized recording or demo conversation with turn-taking, proper nouns, and brief interruptions.
- Set source and target languages, enable speaker separation and bilingual display, and check alignment sentence by sentence.
- Compare terminology, speaker names, and references before and after applying longer context.
- Measure end-to-end delay, omissions, mistranslations, and speaker attribution. Have a person review decisions, amounts, and contract language.