2026
7 篇文章NVIDIA AVO 於 ARC-AGI-3 滿分:模型之外的代理系統才是關鍵
NVIDIA 的 AVO 代理系統在互動推理基準 ARC-AGI-3 公開集拿到 100.00 RHAE、用 6,624 步解完 183 關,本文解析其監督者與持續記憶體設計,以及「模型不是整個代理」的意義。
閱讀文章 ↗DeepSeek V4-Flash 0731:只重做後訓練,代理能力越級跳
2026 年 7 月 31 日,DeepSeek 推出 V4-Flash-0731:架構與規模不變、僅重新後訓練,代理基準大幅超越 V4-Pro-Preview,輸出速度每秒 210 個 token,權重續以 MIT 授權開放。
閱讀文章 ↗Gemini 3.5 Flash 內建電腦操作:OSWorld 78.4 追平 Sonnet
Google DeepMind 把電腦操作能力直接內建到 Gemini 3.5 Flash,OSWorld 拿下 78.4 分追平 Claude Sonnet 4.6,並配上對抗訓練與兩項企業級安全防護,透過 Gemini API 與企業代理平臺開放。
閱讀文章 ↗你的工具真的適合 Agent 嗎?Hugging Face 用完整工作過程來測
Hugging Face 的 agentic benchmark 不只看答案對錯,而是記錄 turns、tokens、錯誤率與 marker 採用率,再掃過模型與工具版本。Transformers 案例顯示:CLI 加 Skill 讓大模型省時間,卻讓 Qwen3-14B 正確率從 67% 掉到 43%。
閱讀文章 ↗視覺代理比結構化 API 貴 45 倍:Reflex 的完整實測
Reflex 讓 Claude 用兩種方式完成同一個後台任務:視覺 Computer Use 燒掉約 55 萬輸入 tokens、花 17 分鐘;結構化 API 只用 1.2 萬 tokens、20 秒完成。45 倍差距來自介面設計,不是模型好壞。
閱讀文章 ↗Zapier 推出自動化 Agent 評測集 AutomationBench:前沿模型全數不及格
Zapier 開源 AutomationBench:47 個模擬 App、600 多個真實商業任務,以最終資料狀態而非模型回答計分。最強的 GPT 6 Astra 完成率也只有 41.4%,最大失敗模式是回報成功但狀態錯誤的假自信。
閱讀文章 ↗GPT-5.2 寫下 METR 時間視野新紀錄:6.6 小時的 Agent 門檻
METR 於 2026 年 2 月 4 日將 GPT-5.2 納入時間視野追蹤:高推理模式的 50% 時間視野約 6.6 小時,刷新紀錄。本文解析指標計算方式、與前代模型對比,以及約每 4–7 個月翻倍的指數趨勢對工程團隊的意義。
閱讀文章 ↗
2026
6 ARTICLESNVIDIA AVO Hits 100% on ARC-AGI-3: The Agent Is the System
NVIDIA's AVO agent system scored 100.00 RHAE on ARC-AGI-3's public set, clearing all 183 levels in 6,624 actions — the harness, not just the model, drives autonomy.
READ POST ↗DeepSeek V4-Flash 0731: Same Architecture, Sharper Agents
V4-Flash-0731 (July 31, 2026) keeps the same architecture and only redoes post-training — agent benchmarks far exceed V4-Pro-Preview, at 210 tok/s with MIT-licensed weights.
READ POST ↗Gemini 3.5 Flash Gets Built-in Computer Use
Google built computer use directly into Gemini 3.5 Flash: 78.4 on OSWorld ties Claude Sonnet 4.6, with adversarial training and enterprise guardrails, available via the Gemini API.
READ POST ↗Reflex Benchmark: Computer Use Costs 45x More Than APIs
Reflex tested one task two ways: vision computer use burned ~551k input tokens in ~17 minutes; structured APIs took ~12k in 20 seconds. The interface sets the cost.
READ POST ↗Zapier's AutomationBench: Real Work Is Still Hard for Agents
Zapier's AutomationBench: 47 simulated apps, 600+ business tasks, scored on final data state, not replies. Top model GPT 6 Astra hits 41.4%; false confidence dominates failures.
READ POST ↗GPT-5.2 Sets a METR Time-Horizon Record: 6.6 Hours
On Feb 4, 2026, METR added GPT-5.2 to its time-horizons tracker: roughly 6.6 hours at 50% success in high-reasoning mode, a new record. What the metric measures and why the doubling trend matters.
READ POST ↗