AssemblyAI launches Sync API: full transcript in a single HTTP call, ~134ms latency
In short: AssemblyAI introduced the Sync API, a new transport for transcribing short audio clips: one HTTP POST request returns a finished Universal-3.5 Pro transcript in the same response, at roughly 134ms median latency. It fills the gap between the Async API (submit-and-poll, 5-6 seconds of added latency) and the Realtime API (WebSocket, built for ongoing sessions), targeting workloads like dictation, voice-agent turn transcription, IVR, and push-to-talk. Clips from 80ms to 2 minutes and up to 40MB are supported, with word error rate of 1.59% on short-form audio. Pricing is $0.45/hr, the same rate as Universal-3.5 Pro Realtime.
This summary was generated automatically by AI from AssemblyAI's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1New Sync API endpoint (POST https://sync.assemblyai.com/transcribe) returns finished transcripts in one HTTP call, no polling or WebSocket
- 2Runs on Universal-3.5 Pro, the same model used in async and realtime products
- 3~134ms p50 latency on a 2-second clip, versus 5-6 seconds via async polling
- 4Supports clips from 80ms to 2 minutes, files up to 40MB, 18 languages
- 5Accepts conversation_context for prior turns, custom prompts, and word_boost keyterms
- 6Word error rate of 1.59% on short-form audio in benchmarks
- 7Priced at $0.45/hr, same as Universal-3.5 Pro Realtime
| Parameter | Before | Now |
|---|---|---|
| Latency (2s clip) | 5-6s (Async submit-and-poll) | ~134ms (Sync API, p50) |
| Word error rate (short-form) | Not specified | 1.59% |
| Price | Not specified | $0.45/hr |
| Clip length supported | Not specified | 80ms to 2 minutes |
| File size limit | Not specified | Up to 40 MB |
| Languages | Not specified | 18 (same as Universal-3.5 Pro) |
Why it matters
Many voice-AI workloads involve short, already-complete audio clips (a finished user turn, a push-to-talk message, an IVR utterance) where neither async polling nor a persistent WebSocket is a good fit. The Sync API removes that architectural mismatch, letting developers get flagship-accuracy transcripts back inside a single request fast enough to feel instant, which matters directly for turn-based conversational agents and dictation apps.
What it means for AI agents and contact centers
If your voice agent stack already handles turn detection, Sync API could replace a WebSocket session or async job for each utterance with a single stateless HTTP call per turn, simplifying infrastructure (no reconnect logic, no session affinity) while cutting latency compared to submit-and-poll. Worth benchmarking against the current STT setup for Lithuanian accuracy (confirm it's among the 18 supported languages), per-call latency under real call conditions, and cost per minute versus the existing pipeline, especially for IVR and short-utterance routing where dead air during polling currently hurts caller experience.
🧪 Worth Testing
Directly applicable to turn-based voice agents, IVR, and dictation-style features: sub-300ms latency with flagship accuracy and a simple stateless HTTP integration makes it a strong candidate to benchmark against current STT latency, accuracy and cost.
Sources
- AssemblyAIOfficialPrimary sourceOriginal article →„Introducing the Sync API: Finished Transcripts in a Single API Call“
- Published by source
- —
- Found by our system
- 6 Oct 2026, 20:35
- Summary generated
- 6 Oct 2026, 20:38
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy