AssemblyAI adds Qwen3.5 4B to LLM Gateway for fast voice-transcript rewriting
In short: AssemblyAI launched a hosted deployment of the open-source Qwen3.5 4B model on its LLM Gateway, optimized specifically for rewriting raw speech-to-text output into clean, formatted text. On voice rewrite tasks like dictation cleanup and transcript formatting, the model averaged 612ms response time—1.9x faster than GPT-4.1—at 94% lower cost per hour of audio processed. The model is priced at $0.10/$0.50 per million tokens (prompt/completion) and supports only max_tokens, temperature, and stream parameters, with no tool calling.
This summary was generated automatically by AI from AssemblyAI's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1New qwen3.5-4b-32k-fast model hosted by AssemblyAI on its own GPUs, 32k context window
- 2Benchmarked at 612ms average latency vs GPT-4.1's 1,138ms on voice rewrite tasks
- 394% lower cost per hour of audio vs GPT-4.1; pricing $0.10/$0.50 per million tokens (prompt/completion)
- 4No tool calling or structured output support—only max_tokens, temperature, stream
- 5For agent workloads with tool calling needs, AssemblyAI recommends its separate qwen3-next-80b-a3b model instead
- 6Available now via AssemblyAI API/Playground, accessible through the same OpenAI-compatible endpoint as Claude, GPT, and Gemini on LLM Gateway
| Parameter | Before | Now |
|---|---|---|
| Context window | Not specified | 32k |
| Input price | Not specified | $0.10 per million tokens |
| Output price | Not specified | $0.50 per million tokens |
| Latency (voice rewrite avg) | GPT-4.1: 1,138 ms | 612 ms |
| Cost per hour of audio | GPT-4.1: $0.1546/hr | $0.0092/hr |
| Tool calling | Not specified | Not supported |
| Supported parameters | Not specified | max_tokens, temperature, stream only |
Why it matters
Many voice products run the lightweight 'clean up the transcript' step on expensive general-purpose models, adding unnecessary latency and cost to every utterance. A small model tuned specifically for this task can cut both significantly, which matters for any real-time voice pipeline where the rewrite step runs on every turn.
What it means for AI agents and contact centers
If your company runs real-time voice agents or dictation features, this rewrite step—removing filler words, resolving self-corrections, reformatting raw STT output—happens on every utterance and directly affects perceived latency. Worth benchmarking this model against your current rewrite/formatting model for latency, cost per hour of audio, and output quality on your own transcripts, including Lithuanian-language audio, before deciding whether to keep using a larger general-purpose model for this specific step.
🧪 Worth Testing
The claimed latency and cost advantages for a common voice-pipeline task (transcript rewrite/formatting) are substantial enough to justify a side-by-side test against whatever model currently handles this step, especially since free credits are available with no credit card required.
Sources
- AssemblyAIOfficialPrimary sourceOriginal article →„Introducing Qwen3.5 4B on LLM Gateway—optimized by AssemblyAI for the fast rewrite tasks at the center of voice products“
- Published by source
- —
- Found by our system
- 6 Oct 2026, 20:35
- Summary generated
- 6 Oct 2026, 20:37
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy