Alibaba's Qwen3.8-Omni-Flash matches Gemini 3.8 Flash on audio-visual tasks while slashing input costs by over 90%
Alibaba releases Qwen 3.8 Omni Flash
Alibaba released Qwen3.8-Omni-Flash, a native omnimodal model with a 1M-token context window that handles text, image, audio, and video. It improves average scores by over 25% versus Qwen3.5-Omni-Plus, cuts audio input prices by more than 98% and audio-visual input by over 93%, and reaches audio-visual performance close to Gemini 3.8 Flash. The model powers agentic workflows for video editing, music video creation, film commentary, and real-time conversation, supported by open-sourced Qwen-Live Harness and Qwen-MM-Plugins.
By scaling data, context, and agentic environments, Qwen3.8-Omni-Flash achieves audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash.
- mavamaarten
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.
I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.
E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
- _ache_
If the performances are comparable, and there is no evidence it's not.
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs:
https://www.alibabacloud.com/help/en/model-studio/model-pric...
- syntaxing
> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash
Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.
They also made a new harness but github link seems to 404.
- conception
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
- podocarp
Please what is flash pro ultra and all these, can they just use semver or something
- esquire_900
The blog post itself is technically unimpressive, and feels like the slop future. Amongst many things
- Scrolling in firefox is a nightmare (only uBlock origin) and the console is full or warnings and debug data.
- Videos are all flashy but just fail to communicate anything beyond what can be said in a small paragraph (and with horrible stock music).
- Figure 1, the headpiece; too small to read, can't zoom in
- tolugenius
Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.
- lxe
Looks like the harness repo is already removed?