Alibaba's Qwen team announced Qwen3.8-Omni-Flash on September 18, an AI model that understands audio and video together. The international API charges $0.15 per million input tokens, 80% below the current standard rate for Gemini 3.8 Flash. Qwen says its audio and video performance is close to Gemini's, but the benchmark table it published includes cases where the two models swap places depending on how the video is read. For tasks such as searching a long recording for the relevant scenes and getting work done, what matters is not only the unit price but also how much material must be processed to reach an answer.
Understanding audio and video, and using tools to get work done
Qwen3.8-Omni-Flash accepts audio and video in addition to text and images, and has a 1-million-token context window. It is a multimodal model that can keep long materials in a conversation's context while researching their contents and calling the tools it needs. Qwen says that, compared with the previous-generation Qwen3.5-Omni-Plus, it has strengthened multi-step use cases such as video editing and post-meeting work.
However, the standard API outputs text. Thinking of it as a model that directly generates video or music would misjudge its purpose. It is designed to understand video content, plan edits, and operate generation services and editing tools to arrive at a finished product. The API supports Chat Completions and Responses, and is offered in multiple regions, including Tokyo.
For real-time audio and video conversation, there is a separately named model, "Qwen3.8-Omni-Flash-Realtime." The official announcement provides connection examples and releases Qwen-Live Harness, an open-source tool that supports ongoing dialogue. The standard text-output API and the voice-responding Realtime version should be treated separately. The pricing below applies to the standard version and cannot be applied directly to the Realtime version.
$0.15 per million input tokens, with conditions on hourly rates
The Alibaba Cloud international pricing page and the QwenCloud model page show the same rates: $0.15 for input and $0.47 for output. Compared with Gemini 3.8 Flash, the gap is large for both input and output.
| Standard API pricing | Qwen3.8-Omni-Flash (international) | Gemini 3.8 Flash (standard paid tier) |
|---|---|---|
| Input, per 1M tokens | $0.15 | $0.75 |
| Output, per 1M tokens | $0.47 | $3.75 |
As of September 2026. The comparison excludes cache discounts and batch processing, and Google's output price includes thinking tokens. According to Google's pricing page, the Gemini rates above apply through December 31, 2026. From January 1, 2027, they are scheduled to become $1.50 for input and $7.50 for output.
Per million tokens, Qwen's input is 80% cheaper and its output about 87.5% cheaper. The calculations are 1 − 0.15 ÷ 0.75 and 1 − 0.47 ÷ 3.75, respectively. However, the same video will not necessarily produce the same token count in both models, so the price ratio cannot be treated as the cost-reduction rate for actual processing.
Qwen also published estimates converted to material length. In the dollar figures in the comparison chart in the announcement, audio input costs less than $0.01 per hour and video with audio about $0.20. For Gemini 3.8 Flash, the figures are $0.09 and $0.84, respectively. The previous-generation Qwen3.5-Omni-Plus is listed at $0.28 and $3.27; the announcement's emphasis on price cuts of over 98% for audio and over 93% for video with audio is relative to that predecessor.
These hourly figures are estimates obtained by multiplying the input cost of a 2-minute clip by 30. Video is read at 720p and 1 frame per second, Gemini's resolution setting is high, and other API settings are defaults. They do not represent the cost of examining an hour of footage in detail at a high frame rate, nor the total including answers and video production. To judge where the low price helps, you need to align how finely the material is read.
Read the whole video at once, or search for the relevant scenes
Qwen compares processing that interprets the input as is in one pass with processing that uses Qwen Code, an agent execution environment. The latter starts from a question, chooses where to look and listen, and gathers clues step by step. If the answer to a question about a long video lies only in a short segment, there is no need to keep reading the whole thing at the same level of detail.
Comparing the evaluation table Qwen published between like methods gives the following. Higher numbers are better. All results are the developer's own evaluations, not independent replications.
| Video benchmark | Qwen (one-pass) | Gemini (one-pass) | Qwen + Qwen Code | Gemini + Qwen Code |
|---|---|---|---|---|
| OmniVideoBench | 63.4 | 65.2 | 67.8 | 70.1 |
| Video-MME-v2 | 65.0 | 71.0 | 71.3 | 72.7 |
| LVOmniBench | 63.3 | 70.7 | 73.6 | 70.7 |
In the table, Qwen refers to Qwen3.8-Omni-Flash and Gemini to Gemini 3.8 Flash. On LVOmniBench, the one-pass result had Qwen 7.4 points behind Gemini, but with Qwen Code applied to both models, Qwen comes out 2.9 points ahead. The differences are calculated as 63.3 − 70.7 and 73.6 − 70.7.
On this long-video reasoning task, how the material given to the model is searched for mattered enough to change the ranking. On the other two benchmarks, however, Gemini remains ahead even with Qwen Code. Placing only Qwen's tool-assisted score next to Gemini's one-pass score would hide this difference. A comparison that aligns the processing method actually used is more useful than a summary like "on par with Gemini."
The amount read also differed. In another Qwen comparison, OmniVideoBench accuracy rose from 63.4 to 67.8 while token consumption per question fell from 145,736 to 79,117. Dividing the reduction by the original 145,736 gives a cut of about 45.71%. This result is with a setting that retains conversation context across turns.
If, on top of the low input price, the relevant scenes can be narrowed down, the burden of repeatedly examining recordings could shrink. However, this token reduction is not a reduction in processing time or in the final bill. You need to look at where costs arise, including tool execution and additional answer generation.
Speaker diarization in meetings improves, but transcription has weaknesses
One improvement Qwen highlighted is speaker diarization, distinguishing who spoke when in a multi-person meeting. On the AliMeeting evaluation, the diarization error rate (DER) fell from 88.11 for the previous generation to 3.35. Because the model can handle video and audio together, it is intended to link statements to people and produce minutes and action items.
Not every audio result improved uniformly. On the FLEURS transcription benchmark, which aggregates 60 languages including Japanese, word error rate (lower is better) was 9.3 for Qwen3.8-Omni-Flash versus 7.2 for Qwen3.5-Omni-Plus. It also trails Gemini 3.8 Flash's 7.9 on this metric. Speech translation BLEU is 31.8 versus 32.2 for the previous generation, and 33.0 for Gemini.
Distinguishing speakers and transcribing what they said accurately are different abilities. For meeting minutes, even if speaker names are correct, you will want to check that technical terms and decisions have not been mistaken. Input support for 113 languages and dialects is broad, but the aggregated FLEURS figures do not reveal quality for Japanese alone. Adopting it for Japanese meetings would still require evaluation on recordings close to the actual acoustics and speaking style.
Total video production cost is not determined by model pricing alone
The public Qwen-MM-Plugins repository shows concretely how video production works. omni-chatcut, which creates video from music and handles movie commentary and video translation, uses generation and audio-video understanding services along with editing software such as FFmpeg. For translated audio, it may also use an external dubbing service. This is why the model's input price alone cannot represent the cost of producing a finished video.
The same plugin collection includes a feature that creates PDF notes with screen captures from explainer videos, and one that records the characters, dialogue, and sounds of a long video for reuse. There is room to shift from re-reading an entire video every time to extracting and reusing the needed content. The repository also explains that many agent execution environments cannot pass audio directly to the main model and currently handle it through the API. Designs need to account for connections with existing tools.
What will decide adoption is how far costs, including output and external services, fall when the same quality of deliverable can be produced from the same recording. If that condition can be met for Japanese meeting minutes and teaching materials, regularly re-examining long audio and video becomes a realistic way of working.
