AssemblyAI adds Qwen3.5 4B to LLM Gateway for fast voice-transcript rewriting
AssemblyAI launched a hosted deployment of the open-source Qwen3.5 4B model on its LLM Gateway, optimized specifically for rewriting raw speech-to-text output into clean, formatted text. On voice rewrite tasks like dictation cleanup and transcript formatting, the model averaged 612ms response time—1.9x faster than GPT-4.1—at 94% lower cost per hour of audio processed. The model is priced at $0.10/$0.50 per million tokens (prompt/completion) and supports only max_tokens, temperature, and stream parameters, with no tool calling.