Skip to content
Wednesday, 7 October 2026
Tenesys AI News
Subscribe

Qwen3.5 4B (qwen3.5-4b-32k-fast)

Change history

  1. 6 Oct 2026
    AssemblyAI adds Qwen3.5 4B to LLM Gateway for fast voice-transcript rewriting
    ParameterBeforeNow
    Context windowNot specified32k
    Input priceNot specified$0.10 per million tokens
    Output priceNot specified$0.50 per million tokens
    Latency (voice rewrite avg)GPT-4.1: 1,138 ms612 ms
    Cost per hour of audioGPT-4.1: $0.1546/hr$0.0092/hr
    Tool callingNot specifiedNot supported
    Supported parametersNot specifiedmax_tokens, temperature, stream only

News

Voice AI· Important· 🧪 Worth Testing

AssemblyAI adds Qwen3.5 4B to LLM Gateway for fast voice-transcript rewriting

AssemblyAI launched a hosted deployment of the open-source Qwen3.5 4B model on its LLM Gateway, optimized specifically for rewriting raw speech-to-text output into clean, formatted text. On voice rewrite tasks like dictation cleanup and transcript formatting, the model averaged 612ms response time—1.9x faster than GPT-4.1—at 94% lower cost per hour of audio processed. The model is priced at $0.10/$0.50 per million tokens (prompt/completion) and supports only max_tokens, temperature, and stream parameters, with no tool calling.

AssemblyAI
Qwen3.5 4B (qwen3.5-4b-32k-fast) · TENESYS AI NEWS