Qwen3.8-Omni-Flash - Tokenstead

4 min read Original article ↗

Models / Qwen3.8-Omni-Flash

Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC). Released 2026-09-14 by Alibaba’s Qwen team. Total parameters and architecture are undisclosed - the predecessor qwen3-omni-flash was a Thinker-Talker MoE, but Qwen has not published the 3.8 breakdown, so this card lists no parameter count.

  • Context and I/O: 1M tokens in (991,808 usable without thinking), up to 131K out. Thinking is on by default with adjustable reasoning effort. Function calling, implicit caching, and Responses session caching are supported.
  • Audio: ASR covers 113 languages and dialects (the same set as Qwen3.5-Omni); spatial (multichannel) audio input is supported. The realtime variant takes camera frames at 1 fps / 720p and handles about an hour of audio-visual input with speaker diarization.
  • Benchmarks vs Gemini 3.8 Flash (Alibaba’s own numbers): wins the overall-audio and most audio-visual boards - WildClawBench-MM 71.0 vs 58.9, DailyOmni 85.1 vs 84.0, SpotSoundBench 67.2 vs 39.7, MMAU 81.8 vs 76.9, AliMeeting DER 3.4 vs 17.2 - and loses the agentic and video boards: OmniGAIA 74.0 vs 78.6, Video-MME-v2 65.0 vs 71.0, AgenticVBench 36.8 vs 45.0. Mixed, not a sweep - “beats Gemini on multimodal” is true on audio, not across the board.

No open weights. The model runs only via Alibaba Cloud Model Studio (Beijing and Singapore, plus Hong Kong, Tokyo, Frankfurt, and US-Virginia endpoints). Only the companion repos (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced. The community reaction said it plainly: “No OSS :/”.

Cloud API: $0.15/1M input, $0.47/1M output, $0.016/1M cache-hit input on the Singapore endpoint ($0.113 / $0.382 / $0.014 on mainland/Global). Unlike its predecessor there is no per-modality split - audio and image tokens bill at the flat input rate, which is how Alibaba gets its “>98% cheaper per hour of audio than Qwen3.5-Omni-Plus” claim.

general reasoning vision agentic

Parameters
125.0B

Context
1000k

License
proprietary

Developer
Alibaba

Origin
🇨🇳 China

Released
Sep 2026

What people are building with Qwen3.8-Omni-Flash

Real demos from X

Announcement: first omni-modal model built around agentic capabilities - +19.5 avg agent points, video input cost down ~89% View on X →

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Score per dollar

533 pts per $/M input

general_score (80) divided by cheapest input price ($0.15/M). Higher is better value. See live pricing.

Related models

Qwen3.8-Flash-Next

180.0B 6.0B active enthusiast 24GB min RAM Aug 26, 2026

Qwen3.8-27B

27.0B enthusiast 8GB min RAM Aug 5, 2026

Qwen3.8-Max

2400.0B 95.0B active premier no local build $2.00/M in Aug 3, 2026

Qwen3.6 35B A3B

35.0B 3.0B active enthusiast 20GB min RAM $0.05/M in Apr 22, 2026

Qwen3.6 27B

27.0B enthusiast 17GB min RAM $0.30/M in Nov 1, 2025

Qwen3 14B

14.7B consumer 9GB min RAM $0.10/M in Apr 28, 2025

Qwen3 235B A22B

235.0B 22.0B active workstation 68GB min RAM $0.46/M in Apr 28, 2025

Qwen3 30B A3B

30.5B 3.0B active enthusiast 16GB min RAM $0.12/M in Apr 28, 2025

Qwen3 32B

32.8B enthusiast 20GB min RAM $0.08/M in Apr 28, 2025

Qwen3 4B

4.0B edge 3GB min RAM Apr 28, 2025

Qwen3 8B

8.2B consumer 5GB min RAM $0.12/M in Apr 28, 2025

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Or run it in the cloud

Live per-provider pricing, throughput and uptime - refreshed 18 days ago via OpenRouter. Click a column to sort.

some pricing may be stale - last verified 2026-09-14

Provider Type Input $/M Output $/M Cache $/M Tok/s Latency Uptime Value
API 0.15 0.47 0.016 - - - cheapest

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

Detailed API pricing page + JSON endpoint →

See who runs Alibaba in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...