2026
16 篇文章Astra 實測數據拆解:ExploitBench 滿分、兩個零日漏洞,與少用 9% token 的祕密
Astra 在 ExploitBench 拿下 100% 解題率(GPT-5.6 Sol 只有 22%),評測過程甚至發現兩個全新零日漏洞。本文從開發者角度拆解官方評測數據:V8 內部移植、沙箱逃逸鏈、提權鏈,以及 token 效率為何比 Raw 能力更值得注意。
閱讀文章 ↗Meta Muse Spark 1.3:把「會問問題、知道極限」的代理行為當成主打功能
Meta 於 2026 年 9 月 2 日推出 Muse Spark 1.3,可在單一長對話中執行多輪代理工作流,主動提問、確認關鍵行動並標示自身知識極限。內部對比 Muse Spark 1.2 工具呼叫減少約 20%、token 用量減少約 25%,max reasoning 模式待安全測試完成後推出。
閱讀文章 ↗GLM-5.3 權重開放下載:753B 旗艦的本地部署現實
2026 年 8 月 25 日,Z.ai 把 753B 參數的 GLM-5.3 權重放上 Hugging Face,三天後在 Hacker News 衝上 806 分。本文解析其 MoE 架構、自訂授權條款與本地部署實測。
閱讀文章 ↗Z.ai 發表 GLM-5.3:開源程式碼新高,網路攻擊能力超預期
2026 年 8 月 14 日,Z.ai 發表 GLM-5.3:沿用 GLM-5.2 基座、靠後訓練把程式碼基準推上開源新高,並主動披露網路攻擊能力「發展得比我們預期更快」,授權同步轉為自訂條款。
閱讀文章 ↗Qwen3.8-Max 正式發布並開放權重:鎖定寫程式與代理協作
2026 年 8 月 3 日,阿里巴巴正式發布 Qwen3.8-Max:百萬 token 上下文、多模態輸入、輸入每百萬 token 2 美元,並成為 Max 系列首個開放權重的模型,主打寫程式與代理協作。
閱讀文章 ↗Grok 4.5:勝負各半的 benchmark,與四分之一 token 的效率算盤
SpaceXAI 與 Cursor 共訓的 Grok 4.5 主打 coding 與 agentic 任務。本文拆解勝負混合的 benchmark 全景、80 TPS 與四分之一 token 用量的效率經濟學,以及 $2/$6 定價與限時免費的取得方式。
閱讀文章 ↗OpenAI 撤回 SWE-Bench Pro 建議:約三成題目是壞的
OpenAI 於 2026 年 7 月 8 日公開承認先前推薦的 SWE-Bench Pro 約三成題目有缺陷並撤回採用建議,文中整理四類失敗模式、三層審計流程,與重建編碼評測信任的具體主張。
閱讀文章 ↗GLM-5.2 開源發布:百萬上下文長任務表現緊咬 Opus 4.8
Z.AI 於 2026 年 6 月 17 日開源 753B 參數的 GLM-5.2:MIT 授權、穩定 1M token 上下文,長任務與程式基準緊追 Claude Opus 4.8,成為開源陣營排名最高的模型。
閱讀文章 ↗Claude Opus 4.8 發布:Dynamic Workflows 與更誠實的程式碼模型
Anthropic 在 41 天後推出 Claude Opus 4.8:價格不變,新增可協調上百個子代理的 Dynamic Workflows 與 Effort 控制,SWE-bench Pro 升至 69.2%,Mythos 預告數週內開放。
閱讀文章 ↗Qwen3.6-35B-A3B 開源釋出:3B 啟動參數的代理編碼模型
阿里巴巴 4 月 16 日以 Apache-2.0 開源 Qwen3.6-35B-A3B:35B 總參數僅啟動 3B 的 MoE,原生 262K 上下文,Terminal-Bench 2.0 達 51.5;Simon Willison 筆電實測在 SVG 任務勝過同日發布的 Claude Opus 4.7。
閱讀文章 ↗Cursor Composer 2 被抓包以 Kimi K2.5 為底:開源權重的署名難題
2026 年 3 月 20 日 Cursor 發表 Composer 2,數小時後即有開發者從 API 流量裡的 model ID 讀出底模為 Moonshot 的 Kimi K2.5。Cursor 兩天後承認,引發開源授權合規與行銷透明的產業論戰。
閱讀文章 ↗Claude Opus 4.7 發表:軟體工程、長時 coding 與 vision 三線升級
依 Anthropic release notes,Claude Opus 4.7 於 3 月 23 日發表,官方列出三個重點:更強的軟體工程能力、可長時間運行的 coding 任務,以及 vision。本文從發布內容看它對 coding 工作流的實際影響與升級時機。
閱讀文章 ↗Claude Sonnet 4.6 上線:直接成為免費與 Pro 用戶的預設 model
Anthropic 於 2026 年 2 月 17 日發布 Claude Sonnet 4.6,並讓它直接成為免費與 Pro 用戶的預設 model;CNBC 指出其知識截止於 2025 年 5 月。預設位置的影響力,比任何單一規格都大。
閱讀文章 ↗MiniMax 開源 M2.5 與 Lightning:SWE-Bench 80.2%,十分之一價格逼近 Opus 4.6
2026 年 2 月 12 日 MiniMax 發布 M2.5 與 M2.5-Lightning 並開源 229B 權重:SWE-Bench Verified 80.2%,解題速度與 Claude Opus 4.6 持平但成本約十分之一,輸出單價最低每百萬 token 1.2 美元。
閱讀文章 ↗智譜開源 GLM-5:從 vibe coding 到 agentic engineering
2026 年 2 月 11 日,智譜(Z.ai)開源發布旗艦模型 GLM-5,主打更強的 coding 能力與長時程 agent 任務,官方標題寫著「From Vibe Coding to Agentic Engineering」,技術報告同步上架 arXiv。本文解析開源旗艦定位與中國模型軍團的發布節奏。
閱讀文章 ↗GPT-5.3-Codex 登場:OpenAI 最強 coding model 參與了自己的開發
2026 年 2 月 5 日,OpenAI 發表 GPT-5.3-Codex,定位為迄今最強的 agentic coding model,媒體報導並指它「幫忙建造了自己」——參與了自身的開發管線。距離 GPT-5.2-Codex 僅三週,本文解析發布節奏與自我開發宣稱的意義。
閱讀文章 ↗
2026
15 ARTICLESInside Astra's Cyber Evals: 100% on ExploitBench, Two Zero-Days, and 9% Fewer Tokens
Astra scored 100% on ExploitBench where GPT-5.6 Sol scored 22%, and the eval surfaced two zero-days. A breakdown of the V8 port, the escape chains, and the token-efficiency gain.
READ POST ↗Meta's Muse Spark 1.3 Ships Agents That Ask Questions and Know Their Limits
Meta's Muse Spark 1.3 runs multi-workflow agents in one long thread: it asks clarifying questions, confirms risky actions, and flags its limits. 20% fewer tool calls, 25% fewer tokens than 1.2.
READ POST ↗GLM-5.3 Weights Are Out: Running a 753B MoE Model Locally
Z.ai put the 753B-parameter GLM-5.3 weights on Hugging Face on August 25, and the Hacker News thread hit 806 points. The architecture, the custom license, and local-run realities.
READ POST ↗GLM-5.3: Open-Weight Coding Frontier With Sharp Cyber Gains
GLM-5.3 reuses the GLM-5.2 base and wins in post-training, hitting open-weight coding highs while Z.ai flags that its cyber capability 'developed faster than we expected.'
READ POST ↗Qwen3.8-Max Goes GA: 1M Context and Open Weights
Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.
READ POST ↗OpenAI Retracts Its SWE-Bench Pro Endorsement: 30% Broken
OpenAI estimates ~30% of SWE-Bench Pro tasks are broken, retracts its adoption recommendation, and lays out four failure modes plus a model-assisted audit playbook.
READ POST ↗GLM-5.2: Open Weights, 1M Context, Long-Horizon Gains
Z.AI open-sources GLM-5.2, a 753B MIT-licensed model with a stable 1M-token context that trails Claude Opus 4.8 by about one percent on long-horizon coding benchmarks.
READ POST ↗Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents
Anthropic ships Claude Opus 4.8 41 days after 4.7 at unchanged pricing, adding Dynamic Workflows with hundreds of subagents, effort control, and sharper uncertainty flagging.
READ POST ↗Qwen3.6-35B-A3B: Open-Weight MoE Punches at Agentic Coding
Alibaba open-sources Qwen3.6-35B-A3B (Apache-2.0): a 35B MoE with 3B active params, 262K context, Terminal-Bench 2.0 at 51.5 — beating Opus 4.7 on a same-day laptop SVG test.
READ POST ↗Cursor Composer 2 Caught Building on Kimi K2.5 Weights
Cursor launched Composer 2 as its own model; a leaked API model ID exposed its Kimi K2.5 base within hours. Cursor's admission ignited an open-weights attribution debate.
READ POST ↗Claude Opus 4.7 Ships: Software Engineering, Long-Running Coding, and Vision
Claude Opus 4.7 arrived March 23 per Anthropic's release notes, with stronger software engineering, long-running coding tasks, and vision. What the combination changes for coding workflows.
READ POST ↗Claude Sonnet 4.6 Ships as the Default Model for Free and Pro Users
Anthropic released Claude Sonnet 4.6 on February 17, 2026 as the default model for free and Pro users; CNBC notes a May 2025 knowledge cutoff. The default slot shapes more than any spec.
READ POST ↗MiniMax M2.5: Open Coding Weights at a Tenth of Opus Cost
MiniMax open-sources M2.5 and M2.5-Lightning: 80.2% SWE-Bench Verified, Opus-4.6-class task times at roughly a tenth of the cost, 229B parameters under modified MIT on Hugging Face.
READ POST ↗Zhipu Open-Sources GLM-5: From Vibe Coding to Agentic Engineering
On February 11, 2026, Zhipu AI (Z.ai) released GLM-5, an open-source flagship built for stronger coding and long-horizon agent tasks — a launch titled 'From Vibe Coding to Agentic Engineering.'
READ POST ↗GPT-5.3-Codex Arrives: OpenAI's Strongest Coding Model Helped Build Itself
OpenAI introduced GPT-5.3-Codex on February 5, 2026 — its most capable agentic coding model yet, one that helped build itself. Three weeks after GPT-5.2-Codex, the cadence is the story.
READ POST ↗