Skip to content
Wednesday, 7 October 2026
Tenesys AI News
Subscribe

All News

TENESYS AI NEWS tracks the most important AI news and explains what changed, why it matters and whether a technology is worth testing.

#Latency — 2 articles ✕

Voice AI· Important· 🧪 Worth Testing

AssemblyAI launches Sync API: full transcript in a single HTTP call, ~134ms latency

AssemblyAI introduced the Sync API, a new transport for transcribing short audio clips: one HTTP POST request returns a finished Universal-3.5 Pro transcript in the same response, at roughly 134ms median latency. It fills the gap between the Async API (submit-and-poll, 5-6 seconds of added latency) and the Realtime API (WebSocket, built for ongoing sessions), targeting workloads like dictation, voice-agent turn transcription, IVR, and push-to-talk. Clips from 80ms to 2 minutes and up to 40MB are supported, with word error rate of 1.59% on short-form audio. Pricing is $0.45/hr, the same rate as Universal-3.5 Pro Realtime.

AssemblyAI
Voice AI· Important· 🧪 Worth Testing

AssemblyAI adds Qwen3.5 4B to LLM Gateway for fast voice-transcript rewriting

AssemblyAI launched a hosted deployment of the open-source Qwen3.5 4B model on its LLM Gateway, optimized specifically for rewriting raw speech-to-text output into clean, formatted text. On voice rewrite tasks like dictation cleanup and transcript formatting, the model averaged 612ms response time—1.9x faster than GPT-4.1—at 94% lower cost per hour of audio processed. The model is priced at $0.10/$0.50 per million tokens (prompt/completion) and supports only max_tokens, temperature, and stream parameters, with no tool calling.

AssemblyAI