Skip to content
Wednesday, 7 October 2026
Tenesys AI News
Subscribe
Voice AI· Important· 🧪 Worth Testing

AssemblyAI launches Sync API: full transcript in a single HTTP call, ~134ms latency

In short: AssemblyAI introduced the Sync API, a new transport for transcribing short audio clips: one HTTP POST request returns a finished Universal-3.5 Pro transcript in the same response, at roughly 134ms median latency. It fills the gap between the Async API (submit-and-poll, 5-6 seconds of added latency) and the Realtime API (WebSocket, built for ongoing sessions), targeting workloads like dictation, voice-agent turn transcription, IVR, and push-to-talk. Clips from 80ms to 2 minutes and up to 40MB are supported, with word error rate of 1.59% on short-form audio. Pricing is $0.45/hr, the same rate as Universal-3.5 Pro Realtime.

Source: AssemblyAIAssemblyAIUniversal-3.5 ProOriginal article ↗

This summary was generated automatically by AI from AssemblyAI's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1New Sync API endpoint (POST https://sync.assemblyai.com/transcribe) returns finished transcripts in one HTTP call, no polling or WebSocket
  • 2Runs on Universal-3.5 Pro, the same model used in async and realtime products
  • 3~134ms p50 latency on a 2-second clip, versus 5-6 seconds via async polling
  • 4Supports clips from 80ms to 2 minutes, files up to 40MB, 18 languages
  • 5Accepts conversation_context for prior turns, custom prompts, and word_boost keyterms
  • 6Word error rate of 1.59% on short-form audio in benchmarks
  • 7Priced at $0.45/hr, same as Universal-3.5 Pro Realtime
Universal-3.5 Pro
ParameterBeforeNow
Latency (2s clip)5-6s (Async submit-and-poll)~134ms (Sync API, p50)
Word error rate (short-form)Not specified1.59%
PriceNot specified$0.45/hr
Clip length supportedNot specified80ms to 2 minutes
File size limitNot specifiedUp to 40 MB
LanguagesNot specified18 (same as Universal-3.5 Pro)

Why it matters

Many voice-AI workloads involve short, already-complete audio clips (a finished user turn, a push-to-talk message, an IVR utterance) where neither async polling nor a persistent WebSocket is a good fit. The Sync API removes that architectural mismatch, letting developers get flagship-accuracy transcripts back inside a single request fast enough to feel instant, which matters directly for turn-based conversational agents and dictation apps.

What it means for AI agents and contact centers

If your voice agent stack already handles turn detection, Sync API could replace a WebSocket session or async job for each utterance with a single stateless HTTP call per turn, simplifying infrastructure (no reconnect logic, no session affinity) while cutting latency compared to submit-and-poll. Worth benchmarking against the current STT setup for Lithuanian accuracy (confirm it's among the 18 supported languages), per-call latency under real call conditions, and cost per minute versus the existing pipeline, especially for IVR and short-utterance routing where dead air during polling currently hurts caller experience.

🧪 Worth Testing

Directly applicable to turn-based voice agents, IVR, and dictation-style features: sub-300ms latency with flagship accuracy and a simple stateless HTTP integration makes it a strong candidate to benchmark against current STT latency, accuracy and cost.

AssemblyAI Sync API· New

Sources

  • AssemblyAIOfficialPrimary source
    „Introducing the Sync API: Finished Transcripts in a Single API Call“
    Original article →
Published by source
—
Found by our system
6 Oct 2026, 20:35
Summary generated
6 Oct 2026, 20:38

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

AssemblyAI launches Sync API: full transcript in a single HTTP call, ~134ms latency · TENESYS AI NEWS