Skip to content
Friday, 9 October 2026
Tenesys AI News
Subscribe
Voice AI· Important· 🧪 Worth Testing

Mistral releases open-source real-time Arabic speech-to-text model Voxtral Mini 4B Realtime Arabic

In short: Mistral AI has released Voxtral Mini 4B Realtime Arabic, a streaming speech-to-text model fine-tuned for Arabic dialects, Modern Standard Arabic, and code-switching between languages. The 4.4-billion-parameter model processes 16 kHz audio in real time with a configurable transcription delay, achieving an 8.82% average Character Error Rate across seven Arabic benchmarks at a 480ms delay — close to the offline Voxtral Transcribe Arabic model's 7.91% CER. It was co-developed with Morocco's digital transition ministry as part of a sovereign AI partnership and is released under Apache 2.0.

Source: Mistral AI (Hugging Face)Mistral AIVoxtral-Mini-4B-Realtime-ArabicOriginal article ↗

This summary was generated automatically by AI from Mistral AI (Hugging Face)'s publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1New model: Voxtral Mini 4B Realtime Arabic, fine-tuned from Voxtral-Mini-4B-Realtime-2602
  • 2~4.4B parameters, BF16 weights, causal audio encoder/decoder architecture
  • 3Processes 16 kHz audio with configurable transcription delay for live captions/voice interfaces
  • 48.82% average CER across seven Arabic benchmarks at 480ms delay, vs 7.91% for offline Voxtral Transcribe Arabic
  • 5Trained with focus on Arabic dialect code-switching, including unintended gains for other language pairs
  • 6Released under Apache 2.0 license, supported via vLLM and Transformers >=5.2.0
Voxtral-Mini-4B-Realtime-Arabic
ParameterBeforeNow
ParametersNot specified~4.4 billion
PrecisionNot specifiedBF16
Audio inputNot specified16 kHz streaming audio
Transcription delayNot specifiedConfigurable, 480 ms tested
Accuracy (CER)Voxtral Transcribe Arabic: 7.91%8.82% average across seven Arabic benchmarks
LicenseNot specifiedApache 2.0

Why it matters

Real-time, low-latency, open-source STT for Arabic dialects with strong code-switching handling fills a gap for voice interfaces and live captioning in Arabic-speaking markets, and its accuracy is close to offline transcription quality despite running in streaming mode.

What it means for AI agents and contact centers

For companies running voice agents or contact centers in multilingual or code-switching environments, this model demonstrates a workable open-source approach to low-latency streaming STT with configurable delay — worth benchmarking against your current STT stack for latency, dialect accuracy, and self-hosting cost, especially if you plan to expand beyond major languages.

🧪 Worth Testing

Its configurable low-latency streaming architecture and strong code-switching performance make it a useful reference point for evaluating real-time STT latency/accuracy tradeoffs, even outside Arabic-language deployments.

Voxtral Mini 4B Realtime Arabic· New

Sources

  • Mistral AI (Hugging Face)OfficialPrimary source
    „mistralai/Voxtral-Mini-4B-Realtime-Arabic released on Hugging Face (automatic-speech-recognition)“
    8 Oct 2026, 14:36
    Licence: Model card; license per model · our summary (content changed)
    Original article →
Published by source
8 Oct 2026, 14:36
Found by our system
9 Oct 2026, 20:39
Summary generated
9 Oct 2026, 20:41

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

Mistral releases open-source real-time Arabic speech-to-text model Voxtral Mini 4B Realtime Arabic · TENESYS AI NEWS