Qwen's Voice Split: What You Can Rent and What You Can Run

I’ve been getting a lot of questions about “Qwen’s new voice model,” and the honest answer is that the phrase can mean two things: the cloud-only Qwen-Audio-3.1 family that Alibaba rolled out between September 20 and 23, or Qwen3-TTS and Qwen3-ASR, the newest open-weight voice models, released in January 2026 and ready to run on your own machine. Both are worth knowing about, so let’s look at what you can rent, what you can run, and how to put a fully local voice loop together.

A split illustration with a glowing cloud of sound waves on one side and a laptop and phone running a voice waveform on the other, connected by a soft arc of light

The track you rent: Qwen-Audio-3.1

Qwen-Audio-3.1 is a line of five hosted models: ASR, ASR-Next, TTS, TTS-Next and Realtime. Alibaba’s Model Studio listed qwen-audio-3.1-realtime-plus on September 20, and the full line was announced at Alibaba’s Yunqi conference on September 23. All five are API-only. There are no open weights, and Alibaba hasn’t disclosed parameter counts.

Realtime is the headline. The QwenCloud docs describe a full-duplex connection with streaming input and output, so the model listens while it speaks and handles interruptions. It supports function calling, web search, cloned voices and eight new system voices, over WebSocket, WebRTC or Alibaba’s AOQ protocol. Alibaba also says it switches languages mid-call and reads emotion and intent. The model page lists a 262,144-token context window, and the docs note that a session keeps up to 50 audio turns or 300 seconds of cumulative audio before older history is dropped.

TTS-Next is the creative one. Alibaba calls it an “AudioGen” model that produces speech, sound effects and ambient sound in a single pass. Its model page lists up to 3,000 input characters, up to 240 seconds of output for podcasts, up to three 30-second reference clips, Chinese and English, and non-streaming output.

ASR covers 30 languages and 16 Chinese dialects with speaker diarization, according to Qwen’s 3.1 ASR page. In Qwen’s own comparison, its message ASR shows about 100 ms to first character, against 570 ms for Doubao Streaming ASR 2.0.

On price, Alibaba says it cut rates by about 85% for Realtime, about 70% for TTS and up to 95% for ASR, as reported by OrcaRouter. Model Studio lists Realtime-plus at $6.4 per million audio-input tokens and $24 per million audio-output tokens in the Singapore region. OrcaRouter reports Beijing-region rates of about ¥40 and ¥150 per million audio-input and combined output tokens.

Every 3.1 benchmark, latency and price figure here comes from Alibaba. I haven’t found an independent result for any 3.1 model yet. The closest outside signal belongs to the earlier Qwen-Audio-3.0-TTS-Plus from July, which OrcaRouter reports in second place on the Artificial Analysis TTS arena at an Elo of 1,259, behind Cartesia Sonic 3.6 at 1,273.

The track you run: Qwen3-TTS and Qwen3-ASR

Qwen3-TTS comes in 0.6B and 1.7B sizes under Apache-2.0. Its technical report describes a 12 Hz speech tokenizer, 10 languages, 3-second voice cloning and voice design from text descriptions. Qwen reports first-packet latency as low as 97 ms for the 0.6B model and 101 ms for the 1.7B, measured on its internal vLLM setup. Developers have clearly taken to it: the 1.7B-Base weights show over 3 million downloads on Hugging Face.

Its partner is Qwen3-ASR, also in 0.6B and 1.7B sizes and also Apache-2.0.

The local tooling is already in good shape:

  • llama.cpp: ggml-org publishes GGUFs for Qwen3-TTS 1.7B-Base and for Qwen3-ASR.
  • MLX: mlx-audio runs mlx-community conversions on Apple Silicon.
  • vLLM: the Qwen3-TTS README says vLLM-Omni has day-0 support, currently for offline inference.
  • LiteRT-LM: a community port of Qwen3-ASR-1.7B reports transcribing 6.4 to 11 second clips in about 2.0 to 2.6 seconds on a Galaxy S26 CPU.

The phone story is thinner for speech output. I couldn’t find a LiteRT or ONNX build of Qwen3-TTS, Ollama doesn’t serve audio output, and I haven’t seen measured phone numbers for Qwen3-TTS.

If you’re on a Mac, this line from the mlx-audio README is a lovely first step:

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello, world!' --voice Vivian

A fully local voice loop

Here’s a suggested setup I’d sketch for a private voice assistant. Treat it as a starting point to experiment with:

  1. Listen: Qwen3-ASR (0.6B or 1.7B) turns speech into text, via GGUF, MLX or the LiteRT port.
  2. Think: a small local LLM reads the transcript and writes a reply. If the assistant needs your notes or files, add a local embedder for retrieval, like the one in my EmbeddingGemma 2 post.
  3. Speak: Qwen3-TTS 0.6B turns the reply into audio, via GGUF or mlx-audio.

This is a turn-taking loop, so it won’t give you Realtime’s full-duplex feel out of the box. If you want an agent that keeps talking while tools run, Alibaba’s qwen-audio-agent is an Apache-2.0 realtime voice-agent runtime built for exactly that. It ships no model weights, so you’d pair it with the models you choose.

What to rent and what to run

  • Rent Qwen-Audio-3.1 when you need full-duplex conversation with interruptions, long context, built-in tools or one-pass soundscapes, and you’re comfortable with audio going to the cloud.
  • Run Qwen3-ASR and Qwen3-TTS when privacy, offline use, cost control or an open license matter most. Voice cloning and 10 languages in a 0.6B model is a great deal of capability to own.

Other open options worth a look

I checked these licenses on their Hugging Face cards:

Go explore

Voice is one of the most human ways to use AI, and you can now build a capable loop on your own hardware with open, permissively licensed models. Open the Qwen3-TTS and Qwen3-ASR model cards, read the Qwen-Audio-3.1 docs, and try the local pieces yourself. Start with a single sentence through mlx-audio or llama.cpp, then wire up the full loop and see how it feels to talk to something running entirely on your own machine.