How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf audarai/Audar-ASR-V1-Turbo:Q4_K_M
Use Docker
docker model run hf.co/audarai/Audar-ASR-V1-Turbo:Q4_K_M
Quick Links

Audar-ASR-V1-Turbo Β· GGUF

Audar's Arabic-first speech-recognition model β€” leaderboard-grade, dialect-aware.

From Arabic to the world.

License Task Format Params Open-AR-ASR Avg CER Emirati GitHub

🧭 Overview Β· πŸ“Š Benchmarks Β· πŸ’» GGUF Deploy Β· πŸ€— Transformers Β· πŸŽ™οΈ Streaming Β· ⚑ vLLM Β· πŸ“„ Tech Report Β· πŸ™ GitHub Β· ☁️ Audar API Β· πŸ“œ License


🧭 What it is

Audar-ASR-V1-Turbo is an Arabic-first generative speech-recognition model β€” the accuracy tier of the Audar-ASR family. It recasts transcription as audio-conditioned next-token prediction over a unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is built on a permissively-licensed open-weight audio-LLM foundation and adapted in-house β€” the contribution is the adaptation (the data curriculum and the alignment rubric), not the foundation:

  • 🧱 Large-scale bilingual pretraining β€” 300,000+ hours of labeled audio, primarily Arabic and English, spanning MSA, Gulf, Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
  • 🎯 Dialect-targeted fine-tuning β€” hardness sampling and multi-task conditioning focused on proper nouns, code-switching, and dialect-faithful orthography.
  • 🧠 KTO preference alignment β€” Kahneman-Tversky Optimization on accented dialectal Arabic, with unpaired binary-desirability labels from trained native annotators across the Gulf, Levantine, Egyptian, and Maghrebi dialects, along five axes: verbatim accuracy, diacritic correctness, code-switch handling, named-entity preservation, and output formatting.

The result is state-of-the-art dialectal Arabic ASR β€” the lowest average WER and CER of any evaluated system on the Open Universal Arabic ASR Leaderboard. It transcribes MSA and every major Arabic dialect, code-switched Arabic–English, and English, across 30 languages in total.

Built on a permissively-licensed open-weight audio-LLM foundation; the adaptation, data, and alignment are Audar's. Full method and results: Audar-ASR-V1 Technical Report.

Model summary

ModelAudar-ASR-V1-Turbo β€” Arabic-first generative ASR (accuracy tier)
TaskAutomatic speech recognition (audio β†’ text)
ApproachGenerative ASR β€” audio encoder + language-model decoder (audio-conditioned next-token prediction)
Trainingbuilt on an open-weight audio-LLM foundation; adapted via a 4-stage curriculum β€” 300k+ hrs bilingual pretraining β†’ multi-task fine-tuning β†’ dialect PEFT β†’ KTO alignment
Decoder parameters2,031,739,904 (2.03B)
Audio encoder parameters317,477,504 (0.32B)
Total parameters2,349,217,408 (2.35B, bf16)
Audio input16 kHz mono; 30 s context (longer audio is chunked/streamed)
LanguagesArabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more
RuntimeGGUF / llama.cpp β€” CPU Β· GPU Β· edge
LicenseAudarAI Community License v1.0

πŸ“Š Benchmarks

Arabic dialectal ASR is hard β€” heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, Audar-ASR-V1-Turbo ranks #1 of 37 systems with the lowest average WER (23.2 %) and the lowest average CER (9.2 %) of any model evaluated β€” and it is the single best system on SADA, MASC-clean, MGB-2 and Casablanca.

Open Universal Arabic ASR Leaderboard β€” full standings

Per-dataset WER % across all six leaderboard test sets, plus the two composite averages. Lower is better; Avg WER is the ranking metric. Audar rows show the leaderboard maintainers' independent reproduction (Aug 2026) under the leaderboard's current normalization; other rows are as previously published by the leaderboard and may shift slightly when the full board is recomputed under the updated normalization. Ours in bold.

# Model Avg WER Avg CER SADA CV-18 MASC-clean MASC-noisy MGB-2 Casablanca
1 Audar-ASR-V1-Turbo (Ours) 23.17 9.20 28.92 8.09 16.73 27.19 11.08 47.02
2 CohereLabs/cohere-transcribe-arabic-07-2026 25.87 11.80 37.47 5.82 19.60 27.07 15.54 49.71
3 omnilingual-asr/omniASR_LLM_7B 28.32 12.52 41.61 8.75 19.69 29.29 14.13 56.46
4 omnilingual-asr/omniASR_LLM_3B 29.96 13.77 46.18 9.15 19.90 30.03 14.22 60.27
5 omnilingual-asr/omniASR_LLM_1B 29.96 13.40 43.84 9.55 20.03 30.26 15.34 60.68
6 CohereLabs/cohere-transcribe-03-2026 30.67 16.37 60.11 8.17 8.66 19.01 25.33 62.71
7 Qwen/Qwen3-Omni-30B-A3B-Instruct 30.71 13.67 44.82 11.46 21.47 30.85 13.09 62.55
8 nvidia-conformer-ctc-large-arabic (lm) 32.91 13.84 44.52 8.80 23.74 34.29 17.20 68.90
9 omnilingual-asr/omniASR_LLM_300M 32.96 14.84 51.38 12.03 20.66 32.45 16.58 64.64
10 google/gemma-4-E4B-it 32.98 13.71 43.40 19.65 24.86 33.59 17.72 58.63
11 Qwen/Qwen3-ASR-1.7B 33.36 12.33 45.53 16.90 24.37 34.29 16.57 64.47
12 mistralai/Voxtral-Small-24B-2507 34.47 15.29 50.82 15.25 23.96 34.43 16.03 66.30
13 nvidia-conformer-ctc-large-arabic (greedy) 34.74 13.37 47.26 10.60 24.12 35.64 19.69 71.13
14 google/gemma-4-E2B-it 35.87 15.34 46.23 23.76 27.47 36.15 20.72 60.87
15 openai/whisper-large-v3 36.86 17.21 55.96 17.83 24.66 34.63 16.26 71.81
16 omnilingual-asr/omniASR_CTC_3B 37.78 19.79 69.85 14.19 21.48 34.60 18.96 67.58
17 omnilingual-asr/omniASR_CTC_7B 38.12 20.91 72.69 12.47 21.08 35.04 20.43 67.02
18 facebook/seamless-m4t-v2-large 38.16 17.03 62.52 21.70 25.04 33.24 20.23 66.25
19 omnilingual-asr/omniASR_CTC_1B 39.29 20.47 71.42 17.55 22.76 35.73 19.96 68.32
20 openai/whisper-large-v3-turbo 40.05 18.87 60.36 25.73 25.51 37.16 17.75 73.79
21 openai/whisper-large-v2 40.20 19.55 57.46 21.77 27.25 38.55 25.17 71.01
22 Qwen/Qwen3-ASR-0.6B 42.19 16.23 53.75 28.28 31.34 42.63 25.45 71.68
23 openai/whisper-large 42.57 20.49 63.24 26.04 28.89 40.79 24.28 72.18
24 mistralai/Voxtral-Mini-3B-2507 42.58 19.90 63.65 22.12 28.37 41.27 22.56 77.52
25 asafaya/hubert-large-arabic-transcribe 45.50 17.35 67.82 8.01 32.94 50.16 37.51 76.53
26 openai/whisper-medium 45.57 22.27 67.71 28.07 29.99 42.91 29.32 75.44
27 nvidia-Parakeet-ctc-1.1b-concat 46.54 23.88 70.70 26.34 30.49 45.95 24.94 80.80
28 omnilingual-asr/omniASR_CTC_300M 46.65 21.86 78.11 27.90 28.40 43.26 26.85 75.35
29 nvidia-Parakeet-ctc-1.1b-universal 51.96 25.19 73.58 40.01 36.16 50.03 30.68 81.30
30 microsoft/VibeVoice-ASR 52.99 28.95 69.83 44.25 32.95 52.43 25.10 93.37
31 facebook/mms-1b-all 54.54 21.45 77.48 26.52 38.82 57.33 39.16 87.95
32 openai/whisper-small 55.13 21.68 78.02 24.18 35.93 56.36 48.64 87.64
33 whitefox123/w2v-bert-2.0-arabic-4 58.13 27.62 87.34 41.79 37.82 53.28 40.66 87.88
34 jonatasgrosman/wav2vec2-large-xlsr-53-arabic 60.98 25.61 86.82 23.00 42.75 64.27 56.29 92.72
35 speechbrain/asr-wav2vec2-commonvoice-14-ar 65.74 30.93 88.54 29.17 49.10 69.57 64.37 93.68

Bold = best in column. The 37th system, our sibling edge model Audar-ASR-V1-Flash (0.78B), enters at 32.04 avg WER β€” see its card for the full row. Audar-ASR-V1-Turbo owns both composite averages and leads on SADA, MASC-clean, MGB-2 and Casablanca; the recent Cohere and OmniASR systems are the closest competitors, each strongest on a subset of the conversational and clean-read sets. Casablanca (Moroccan Darija) is the hardest set for every system.

Emirati Arabic

Set WER % CER %
Emirati (Mixat, full 1,585-clip test) 19.4 7.3

On Emirati, the real recognition error is β‰ˆ 7.3 % β€” near-parity with spontaneous English β€” while the residual up to 19.4 % WER is largely orthographic convention (near-miss spelling of the same word, e.g. Ψ§Ω†ΨͺΩˆβ†”Ψ§Ω†Ψͺوا, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.

🏁 Benchmark-parity inference (qwen-asr) β€” recommended

Our leaderboard numbers were produced with the qwen-asr package, which implements this model's I/O protocol natively β€” and were independently reproduced by the leaderboard maintainers with this exact code:

# pip install qwen-asr torch
import torch
from qwen_asr import Qwen3ASRModel

model = Qwen3ASRModel.from_pretrained(
    "audarai/Audar-ASR-V1-Turbo",
    dtype=torch.bfloat16, device_map="cuda:0",
    max_inference_batch_size=16, max_new_tokens=256,
)
results = model.transcribe(audio=["clip.wav"], language=["Arabic"])
print(results[0].text)

Protocol handling is mandatory, not optional:

  • language="Arabic" makes the package prefill language Arabic<asr_text> into the prompt, so the model never free-runs language identification.
  • The model's no-speech verdict (language None<asr_text>) is mapped to an empty transcript; without this, non-speech audio (music, silence) can produce repetition loops.
  • max_new_tokens=256 and bf16 are the exact decode settings behind our published numbers.

If you use raw transformers (below), you must strip the language <Lang><asr_text> output prefix yourself and expect degraded scores on non-speech-heavy data.

πŸ’» GGUF inference (llama.cpp)

Turbo runs on llama.cpp via the multimodal (mtmd) path β€” a quantized decoder GGUF plus a BF16 audio projector (mmproj). Build a recent llama.cpp (with Qwen3-ASR support), then:

./llama-mtmd-cli \
  -m       Audar-ASR-V1-Turbo-Q8_0.gguf \
  --mmproj mmproj-Audar-ASR-V1-Turbo.gguf \
  --audio  clip.wav \
  -sys     "فرّغ Ψ§Ω„ΩƒΩ„Ψ§Ω… Ψ§Ω„ΨΉΨ±Ψ¨ΩŠ Ψ§Ω„ΨͺΨ§Ω„ΩŠ." \
  --temp 0

⚠️ The audio projector (mmproj) must stay BF16 (its ClippableLinear is numerically sensitive). The decoder quantizes normally.

Prefer a managed endpoint? The Audar-ASR family is also available via the Audar API/SDK β€” streaming, speaker-attributed transcription, and diarization, production-hosted.

GGUF variants

File Approx. size Notes
Audar-ASR-V1-Turbo-Q4_K_M.gguf ~1.28 GB Smallest; constrained hardware
Audar-ASR-V1-Turbo-Q8_0.gguf ~2.16 GB Near-lossless (recommended)
Audar-ASR-V1-Turbo.gguf (BF16) ~4.07 GB Full precision decoder
mmproj-Audar-ASR-V1-Turbo.gguf ~0.64 GB BF16 audio encoder β€” required, keep BF16

πŸ€— Transformers (full-precision safetensors)

The full-precision bf16 weights are published at the repo root β€” the reference checkpoint the GGUF and W4A16 builds are derived from (2,349,217,408 params, safetensors). Standard πŸ€— Transformers, loaded with trust_remote_code=True (the repo ships the self-contained Qwen3-ASR code).

# pip install "transformers==4.57.6" torch librosa
import torch, librosa
from transformers import AutoProcessor, AutoModelForCausalLM

repo  = "audarai/Audar-ASR-V1-Turbo"
proc  = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo, trust_remote_code=True,
    dtype=torch.bfloat16, device_map="cuda:0",
).eval()

SYSTEM = "فرّغ Ψ§Ω„ΩƒΩ„Ψ§Ω… Ψ§Ω„ΨΉΨ±Ψ¨ΩŠ Ψ§Ω„ΨͺΨ§Ω„ΩŠ."          # "Transcribe the following Arabic speech."
audio, _ = librosa.load("clip.wav", sr=16000, mono=True)

conv = [{"role": "system", "content": SYSTEM},
        {"role": "user",   "content": [{"type": "audio"}]}]     # audio placeholder (a list, not "<audio>")
text   = proc.apply_chat_template(conv, tokenize=False, add_generation_prompt=True)
inputs = proc(text=text, audio=audio, sampling_rate=16000, return_tensors="pt").to(model.device)
inputs["input_features"] = inputs["input_features"].to(model.dtype)   # features are fp32 -> cast to bf16

out = model.generate(**inputs, max_new_tokens=440, do_sample=False)
print(proc.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0].strip())

The self-contained modeling code targets transformers==4.57.6 (the version this checkpoint was built and validated with). For version-independent, high-throughput serving, prefer vLLM β€” it implements Qwen3-ASR natively (no custom code); see below.

File (repo root) Approx. size Notes
model.safetensors ~4.7 GB Full bf16 weights (2,349,217,408 params)
config.json Β· *_audar_asr.py Β· __init__.py β€” Config + self-contained Qwen3-ASR modeling code
tokenizer files Β· preprocessor_config.json β€” Qwen3 tokenizer + Whisper-mel feature extractor

⚑ vLLM inference (GPU serving)

Turbo also runs on vLLM for high-throughput GPU serving with an OpenAI-compatible API. vLLM implements the Qwen3-ASR architecture natively (Qwen3ASRForConditionalGeneration + Qwen3ASRRealtimeGeneration) β€” no custom serving code: point vLLM at a checkpoint and it exposes /v1/chat/completions, /v1/audio/transcriptions, and a realtime /v1/realtime WebSocket.

vLLM serves quantized compressed-tensors checkpoints (not the GGUF files β€” those are for llama.cpp). A vLLM-ready 4-bit (W4A16) build is provided in the vllm-w4a16/ folder:

Build Folder Size Decoder Audio encoder / lm_head / embeddings Accuracy
W4A16 vllm-w4a16 ~2.6 GB INT4 (group-128) BF16 (kept) ~+1 pp CER vs BF16

Only the language-model decoder is quantized; the audio encoder + projector stay BF16 (the projector's ClippableLinear is numerically sensitive β€” the same rule as the GGUF mmproj), as do lm_head and the token embeddings. An FP8 build (lossless vs BF16, ~3.3 GB) can be produced with the same recipe β€” see the note at the end.

For full-precision GPU serving, point vLLM at the repo (the full bf16 root weights) instead of the 4-bit build β€” same native Qwen3-ASR support, no quantization.

1. Install (audio support required)

vLLM needs the audio extras (PyAV + librosa + soundfile) to decode audio; the stock image does not ship them:

FROM vllm/vllm-openai:v0.24.0
RUN pip install --no-cache-dir av librosa soundfile
docker build -t vllm-audio:0.24.0 .

(Or in a plain environment: pip install "vllm>=0.24" av librosa soundfile.)

2. Get the weights & serve

# download just the vLLM build
hf download audarai/Audar-ASR-V1-Turbo --include "vllm-w4a16/*" --local-dir ./turbo

docker run -d --name audar-asr --gpus '"device=0"' \
  -v $PWD/turbo/vllm-w4a16:/model:ro -p 8000:8000 \
  vllm-audio:0.24.0 \
  --model /model --served-model-name audar-asr-v1-turbo \
  --trust-remote-code --max-model-len 8192 --gpu-memory-utilization 0.4

vLLM auto-detects the compressed-tensors quantization (Marlin INT4 kernel). Weights + KV cache fit on any β‰₯12 GB GPU.

3. Transcribe

Turbo is prompt-steerable: the system message sets the task/language. For Arabic use فرّغ Ψ§Ω„ΩƒΩ„Ψ§Ω… Ψ§Ω„ΨΉΨ±Ψ¨ΩŠ Ψ§Ω„ΨͺΨ§Ω„ΩŠ.; steer other languages with the equivalent instruction. Send 16 kHz mono audio as base64 input_audio and decode greedily (temperature: 0).

import base64, requests

audio = base64.b64encode(open("clip.wav", "rb").read()).decode()   # 16 kHz mono wav
r = requests.post("http://localhost:8000/v1/chat/completions", json={
    "model": "audar-asr-v1-turbo",
    "temperature": 0,
    "max_tokens": 320,
    "messages": [
        {"role": "system", "content": "فرّغ Ψ§Ω„ΩƒΩ„Ψ§Ω… Ψ§Ω„ΨΉΨ±Ψ¨ΩŠ Ψ§Ω„ΨͺΨ§Ω„ΩŠ."},
        {"role": "user", "content": [
            {"type": "input_audio", "input_audio": {"data": audio, "format": "wav"}}
        ]},
    ],
})
print(r.json()["choices"][0]["message"]["content"])

The OpenAI-style POST /v1/audio/transcriptions (multipart file upload) endpoint is also available for Whisper-style clients.

4. Accuracy (FLEURS Arabic, greedy)

Character Error Rate vs the BF16 source β€” CER is the stable cross-precision metric for Arabic, where minor و-segmentation differences inflate WER without changing the characters:

Build AR CER Ξ” vs BF16
BF16 source 2.46 % β€”
W4A16 (this build) 3.73 % +1.27
FP8 (optional) 2.46 % +0.00 (lossless)

Leaderboard-grade full-test-set numbers are in the Benchmarks section above; 4-bit quantization keeps them within ~1 pp CER (FP8 keeps them exactly).

Notes

  • Realtime streaming: vLLM also registers Qwen3ASRRealtimeGeneration, exposing an OpenAI-Realtime-compatible /v1/realtime WebSocket; pair it with VAD/endpointing for stable incremental output.
  • Long audio: the audio encoder is a 30 s window; chunk longer inputs client-side.
  • Producing other precisions (needs the BF16 source weights): quantize the decoder Linears only via llm-compressor model_free_ptq, ignoring the audio tower, lm_head, and embeddings β€” scheme="W4A16" (4-bit) or "FP8_DYNAMIC" (lossless), ignore=["re:.*lm_head.*","re:.*embed_tokens.*","re:.*audio_tower.*"].

πŸŽ™οΈ Real-time streaming

Audar-ASR streams via LocalAgreement-2: as audio arrives the trailing window is re-decoded each hop and a word is committed only once two consecutive decodes agree on it β€” giving stable, low-latency incremental output over the GGUF runtime. Audar's production realtime engine serves the same policy over an OpenAI-Realtime-compatible WebSocket with model-based endpointing and β‰₯64 concurrent streams on a single A100-80GB.

🌍 Languages, dialects & tasks

  • Primary: Arabic β€” MSA and dialectal (Gulf/Emirati, Egyptian, Levantine, Maghrebi), plus code-switched Arabic–English; emits dialect-faithful orthography from audio alone.
  • Also: English + 28 additional languages.
  • Task: transcription (audio β†’ UTF-8 text), prompt-steerable for language and formatting.

Intended use & limitations

Intended use. Broadcast/media transcription, meeting & contact-center intelligence, voice agents, captioning, and accessibility β€” cloud or on-prem.

Limitations.

  • Maghrebi / Moroccan Darija (Casablanca) remains the hardest condition (~63 % WER) for all systems.
  • Heavily code-switched telephony and low-SNR audio degrade accuracy relative to clean MSA.
  • Long-form audio can drift on very long recordings.
  • Not evaluated for, and must not be used for, covert speaker identification.

πŸ“œ License

Released under the AudarAI Community License v1.0 β€” research and limited commercial use for qualifying Community Entities; enterprise / large-scale / MaaS use requires an AudarAI Enterprise License. See audarai.com/license/audarai-community-license-v1.0.

Citation

@misc{audar-asr-turbo-2026,
  title  = {Audar-ASR-V1: A Multilingual, Arabic-First Generative Speech Recognition Foundation Model},
  author = {AudarAI},
  year   = {2026},
  note   = {Audar-ASR-V1-Turbo},
  url    = {https://github.com/AudarAI/Audar-ASR-V1/blob/main/report/Audar-ASR-V1-Technical-Report.pdf}
}

About AudarAI

Leading Arabic-First Multilingual Audio Intelligence

AudarAI starts with Arabic β€” and expands to the world.

We are building advanced multilingual audio intelligence that helps individuals, enterprises, and governments communicate across languages, cultures, and borders. By combining Arabic-first speech technology with global multilingual AI, AudarAI transforms voice into understanding, interaction, and connection.

Our work spans speech recognition, speech understanding, voice-enabled digital assistants, human-computer interaction, and intelligent audio systems designed for real-world impact. From empowering people to access technology in their native language to helping organizations communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and connected.

Arabic-first. Multilingual by design. Human-centered at heart.

🌐 www.audarai.com Β· πŸ€— Hugging Face Β· GitHub Β· contact@audarai.com

Β© 2026 AUDARAI PTE. LTD. Β· Licensed under the AudarAI Community License v1.0

Downloads last month
5,366
Safetensors
Model size
2B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ 1 Ask for provider support

Model tree for audarai/Audar-ASR-V1-Turbo

Quantizations
1 model

Spaces using audarai/Audar-ASR-V1-Turbo 3