Signum
Feed
Useful signal19 Sept 2026medium confidence

Alibaba's Qwen releases Qwen3.8-Omni-Flash, a multimodal agent model priced well below Google's Gemini 3.8 Flash

Qwen released Qwen3.8-Omni-Flash, a multimodal (audio+video) model with a 1M-token context window, available via Qwen Studio, Qwen Cloud, and API, priced at $0.15/M input and $0.47/M output tokens — undercutting Gemini 3.8 Flash's introductory $0.75/M input and $3.75/M output pricing. Qwen also released open-source Qwen-MM-Plugins (video editing, speaker recognition, PDF video notes, workflows for agents like Claude Code, Gemini CLI, Qwen Code) and Qwen-Live Harness for real-time camera/mic interaction.

CapabilityEconomicsAccess

Entities: Qwen, Alibaba, Qwen3.8-Omni-Flash, Google, Gemini 3.8 Flash, Qwen-MM-Plugins

67Useful signal
1 source
0 primary
Was this useful?
01

What happened

Alibaba's Qwen team released Qwen3.8-Omni-Flash, a multimodal model handling audio and video with a 1M-token context window, available through Qwen Studio, Qwen Cloud and API. Pricing is $0.15 per million input tokens and $0.47 per million output tokens, well below Google's Gemini 3.8 Flash introductory pricing of $0.75 input and $3.75 output. Qwen also released two supporting open-source tools: Qwen-MM-Plugins (video editing, speaker recognition, PDF video notes, agent workflow integrations) and Qwen-Live Harness for real-time camera and microphone interaction.

02

Why it matters

If the pricing holds and the model performs as claimed, this matters for anyone building agents that need to process audio or video cheaply, since the cost gap to Gemini 3.8 Flash is roughly 5x on input and 8x on output. Developers and enterprises evaluating multimodal API costs now have a concrete alternative to benchmark against, and the open-source plugin ecosystem (with hooks for Claude Code and Gemini CLI) could speed adoption among agent builders. The competitive pressure also gives Google and other frontier labs an incentive to cut Flash-tier pricing further, which is good for buyers generally.

03

What is noise

The claim that Qwen3.8-Omni-Flash "matches" Gemini 3.8 Flash on multimodal benchmarks comes from Qwen itself, with no independent benchmark data or primary source link provided in this coverage. Price-undercutting announcements in this space are routine and often followed quickly by the incumbent matching or beating the new price, so treat this as a snapshot rather than a durable competitive shift.

04

Watch next

  1. 01Independent benchmark comparisons (e.g. from LMSYS, Artificial Analysis, or third-party developers) testing Qwen3.8-Omni-Flash against Gemini 3.8 Flash on actual multimodal tasks, not vendor-reported scores
  2. 02Whether Google adjusts Gemini 3.8 Flash pricing downward in response within the next 4-8 weeks
  3. 03Developer adoption signals: GitHub stars/usage on Qwen-MM-Plugins and Qwen-Live Harness, and whether Claude Code or Gemini CLI officially integrate Qwen Code workflows

Coverage

1 story

More capability signals

Full feed →