2026
12 篇文章在 SageMaker HyperPod 上部署 Qwen3.8-2.4T-A95B:單節點跑 2.4T 開源模型的實戰配置
AWS 示範如何在單一 p6-b300 節點上用 NVFP4 量化與 vLLM 部署 2.4T 參數的 Qwen3.8,省下多節點成本。
閱讀文章 ↗Qwen3.8-Flash-Next 發表:新架構把啟用參數壓到 6B
2026 年 8 月 26 日,Qwen 開源 Qwen3.8-Flash-Next:125B 主模型加 51B n-gram 嵌入表,每個 token 只啟用 6B 參數,原生 262K 上下文,官方定位是極致成本效率。
閱讀文章 ↗Qwen3.8-27B 開源釋出:體積小、看得懂圖,只是預設想太多
阿里巴巴 Qwen 兌現一週前承諾,釋出 Apache 2.0 的 Qwen3.8-27B 開放權重:262K 上下文、原生視覺、單機可跑;Simon Willison 實測稱讚能力,也點出 xhigh 推理預設讓簡單任務慢上數倍。
閱讀文章 ↗Qwen3.8-Max 正式發布並開放權重:鎖定寫程式與代理協作
2026 年 8 月 3 日,阿里巴巴正式發布 Qwen3.8-Max:百萬 token 上下文、多模態輸入、輸入每百萬 token 2 美元,並成為 Max 系列首個開放權重的模型,主打寫程式與代理協作。
閱讀文章 ↗Qwen 3.8 登場:阿里 2.4 兆參數旗艦先上訂閱,權重還沒開
2026 年 7 月 19 日,阿里巴巴發布 Qwen 3.8,2.4 兆參數旗艦以 Qwen3.8-Max-Preview 形態先在訂閱方案上線,開放權重尚未釋出,被視為對 Kimi K3 的回應。
閱讀文章 ↗Qwen3.6-35B-A3B 開源釋出:3B 啟動參數的代理編碼模型
阿里巴巴 4 月 16 日以 Apache-2.0 開源 Qwen3.6-35B-A3B:35B 總參數僅啟動 3B 的 MoE,原生 262K 上下文,Terminal-Bench 2.0 達 51.5;Simon Willison 筆電實測在 SVG 任務勝過同日發布的 Claude Opus 4.7。
閱讀文章 ↗Qwen3.5-Omni 登場:原生全模態、聽 113 種語言、說 36 種
阿里巴巴 Qwen 團隊推出原生全模態模型 Qwen3.5-Omni,單一管線處理文字、影像、音訊與視訊並即時說話,256K 上下文、可聽逾 10 小時音訊,Plus 版宣稱在一般音訊任務超越 Gemini 3.1 Pro。
閱讀文章 ↗USCC 報告:中國開源 AI 的「雙循環」正在強化工業主導地位
2026 年 3 月 23 日,美中經濟與安全審查委員會發布《Two Loops》研究報告:中國全面押注開源 AI,透過數位與實體兩個互相強化的循環鞏固工業主導地位。Qwen 衍生模型已超過十萬個,而美國出口管制主要只打到數位循環。
閱讀文章 ↗Qwen3.5 小模型補齊戰線:9B 在多項基準超越 gpt-oss-120b
阿里巴巴 2 月底補上 Qwen3.5 的 0.8B–9B 四款小模型:Apache-2.0 開源、原生 262K 情境、支援影像影片輸入。9B 以 9.65B 密集參數在 MMLU-Pro、GPQA Diamond 超越 gpt-oss-120b,目標是手機與筆電上的本地多模態。
閱讀文章 ↗阿里巴巴開源 Qwen3.5:397B 參數、201 種語言的 Agent 時代模型
2026 年 2 月 16 日除夕,阿里巴巴開源 Qwen3.5:首發 Qwen3.5-Plus 為 397B-A17B 稀疏 MoE 原生多模態模型,支援 201 種語言、內建可操作手機與電腦的視覺 agent,長情境解碼吞吐量達 Qwen3-Max 的 8.6 倍。本文解析規格、開源策略與中國 Agent 競賽。
閱讀文章 ↗Qwen3-Max-Thinking 登場:阿里巴巴的兆級參數專有推理 model
阿里巴巴 Qwen 團隊於 2026 年 1 月下旬(約 25 日)發表 Qwen3-Max-Thinking:兆級參數的專有推理旗艦,主打自適應工具使用。InfoWorld 認為它讓企業的 model 選項再添一員。本文解析其定位與選型意義。
閱讀文章 ↗阿里巴巴升級通義 App:AI 助理直接幫你點餐、訂行程
2026 年 1 月 15 日,阿里巴巴把通義 App 升級為可實際辦事的個人 AI 助理:點外送、訂機票、比價都交給代理,背後串接支付寶、飛豬與高德。本文解析一億月活用戶背後的中國 Agent 大戰、變現難題與產品設計課題。
閱讀文章 ↗
2026
13 ARTICLESWhat Deploying Qwen3.8-2.4T-A95B on HyperPod Changes for Self-Hosting Frontier Models
A practical walkthrough for serving a 2.4T open-weights MoE on a single 8-GPU node with vLLM and SageMaker HyperPod.
READ POST ↗Qwen3.8-Flash-Next: 125B MoE with only 6B active parameters
Qwen open-sourced Qwen3.8-Flash-Next: a 125B main model plus a 51B n-gram embedding table, only 6B parameters active per token, 262K context, built for cost efficiency.
READ POST ↗Qwen3.8-27B: Great Open Weights That Overthink by Default
Alibaba's Qwen shipped the Apache 2.0 Qwen3.8-27B a week after promising it: 262K context, native vision, laptop-class local AI — behind a costly xhigh reasoning default.
READ POST ↗Qwen3.8-Max Goes GA: 1M Context and Open Weights
Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.
READ POST ↗Qwen 3.8: Subscription First, Open Weights Later
Alibaba announced Qwen 3.8 on July 19, 2026, putting its 2.4T-parameter flagship behind the Qwen Cloud token plan as Qwen3.8-Max-Preview, with open weights still pending.
READ POST ↗Qwen Code Update: Auto Model Fallback and Nested Subagents for Reliable AI Agents
Qwen Code's v0.19.6-0.19.8 releases center on reliability: fallback masks only capacity and rate-limit errors, nested sub-agents get depth caps and two-layer protection, plus session upgrades.
READ POST ↗Qwen3.6-35B-A3B: Open-Weight MoE Punches at Agentic Coding
Alibaba open-sources Qwen3.6-35B-A3B (Apache-2.0): a 35B MoE with 3B active params, 262K context, Terminal-Bench 2.0 at 51.5 — beating Opus 4.7 on a same-day laptop SVG test.
READ POST ↗Qwen3.5-Omni: Alibaba's Omnimodal Model Speaks 36 Languages
Alibaba's Qwen3.5-Omni handles text, image, audio and video in one native pipeline with real-time speech out: 256K context, 113 recognition languages, three variants led by Plus.
READ POST ↗USCC: China's Open-Source AI Reinforces Industrial Power
A new USCC report argues China's open-source AI strategy is self-reinforcing: open models spread globally while factory deployment feeds data back. Qwen alone has 100,000+ derivatives.
READ POST ↗Qwen3.5 Small Models: 9B Rivals gpt-oss-120b at the Edge
Alibaba adds 0.8B-9B small models to Qwen3.5 under Apache-2.0: native 262K context, image and video input. The 9B beats gpt-oss-120b on MMLU-Pro and GPQA Diamond, targeting phones and laptops.
READ POST ↗Alibaba Open-Sources Qwen3.5 for the Agentic AI Era
Alibaba open-sourced Qwen3.5 on Lunar New Year's Eve: a multimodal 397B-A17B MoE with 201 languages, 8.6x Qwen3-Max decoding throughput, and visual agents that operate phones and desktops.
READ POST ↗Qwen3-Max-Thinking Arrives: Alibaba's Trillion-Parameter Proprietary Reasoner
Alibaba Qwen team released Qwen3-Max-Thinking in late January 2026: a trillion-parameter proprietary reasoning flagship with adaptive tool use, aimed at enterprise model selection.
READ POST ↗Alibaba Upgrades Qwen App Into a Task-Doing AI Agent
Alibaba's Qwen app can now order food delivery, book travel, and hunt deals as an agent powered by Qwen3.8, wired into Alipay, Fliggy, and Amap. Inside China's consumer agent war at 100 million MAU.
READ POST ↗