Qwen 3.8 Omni Flash
jjcm
136 points
30 comments
September 17, 2026
Related Discussions
Found 5 related stories in 71.3ms across 7,105 title embeddings via pgvector HNSW
- Qwen3.8-Flash-Next tosh · 657 pts · August 26, 2026 · 81% similar
- Qwen 3.8 Max linzhangrun · 26 pts · July 19, 2026 · 74% similar
- Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency _ache_ · 11 pts · August 26, 2026 · 73% similar
- Qwen3.8-Flash-Next Intelligence, Performance and Price Analysis theanonymousone · 15 pts · August 27, 2026 · 73% similar
- Qwen3.8-Flash-Next Technical Report [pdf] Philpax · 30 pts · August 26, 2026 · 71% similar
Discussion Highlights (6 comments)
tolugenius
Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.
conception
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
syntaxing
> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash. They also made a new harness but github link seems to 404.
_ache_
If the performances are comparable, and there is no evidence it's not. in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47 That is a massive cost reduction. Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni
lxe
Looks like the harness repo is already removed?
andy_ppp
Why can the Chinese build models and Europe cannot? The algorithms behind this stuff are not that complicated, are they? Is it the cost of energy? The illegality and difficulty of obtaining all the data in the world? Lack of capital to start moonshot labs? Lack of optimism? The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs. I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?