2026
17 篇文章用時間點快照做回溯測試:Exa Snapshot 如何幫你把網頁資料變成可重現的評估環境
Exa 推出 Snapshot 功能,讓開發者用 snapshotAsOf 參數查詢過去某個時間點的網頁內容,解決模型評估中的資料洩漏與回測問題。
閱讀文章 ↗當 AI 測試環境意外連上真實網路:Anthropic 三起事件的檢討
Anthropic 在回顧網路安全評估時,發現 Claude 模型因環境設定錯誤而意外存取真實系統。本文整理事件經過、原因與後續改進,並探討對 AI 評估與產品建構者的啟示。
閱讀文章 ↗Anthropic 撥款五百萬美元,資助 AI 對幸福感影響的獨立評估
Anthropic 推出五百萬美元資助計劃,支持獨立研究 AI 對用戶幸福感的影響,並公開評估指引,涵蓋多輪對話、臨床專家參與、評分驗證、申請時程與成果開源要求。
閱讀文章 ↗NVIDIA AVO 於 ARC-AGI-3 滿分:模型之外的代理系統才是關鍵
NVIDIA 的 AVO 代理系統在互動推理基準 ARC-AGI-3 公開集拿到 100.00 RHAE、用 6,624 步解完 183 關,本文解析其監督者與持續記憶體設計,以及「模型不是整個代理」的意義。
閱讀文章 ↗DARPA 與美國空軍讓 F-16 全 AI 控制飛行:VENOM 計畫的自主里程碑
2026 年 8 月 4 日,DARPA 與美國空軍在 Eglin 空軍基地完成 F-16 全 AI 控制飛行,安全飛行員採 human-on-the-loop。VENOM 測試機隊同時支援 AIR 計畫,在真實飛行中評測多個 AI agents。
閱讀文章 ↗第三方資安測試中,模型為何越界?OpenAI 揭露兩起評估事件
OpenAI 公布 UK AISI 與 Irregular 在第三方資安評估中發生的模型越界事件,分析測試環境設定與模型能力交互下的風險,並提出強化評估環境的方向。
閱讀文章 ↗兩個 API 設定讓 GPT-5.6 Sol 在 ARC-AGI-3 分數三倍跳:評估背後的隱藏變數
OpenAI 發現,保留推理與壓縮上下文這兩個 API 設定,能讓 GPT-5.6 Sol 在 ARC-AGI-3 基準測試的成績從 13.3% 提升到 38.3%,同時減少 6 倍輸出 token。這提醒我們,基準測試衡量的不只是模型能力,還包括測試框架的設計選擇。
閱讀文章 ↗DharmaOCR 對上更新模型:OCR 專用訓練為何仍有價值
Dharma-AI 用巴西葡萄牙文基準比較 DharmaOCR、Mistral OCR4 與 Unlimited-OCR:兩階段訓練、逐 token 漂移機制、Chico Buarque 誤轉案例,以及 0.925 對 0.798 的廠商自評分數背後,產品團隊該看的四個訊號。
閱讀文章 ↗Shippy 的生產經驗:可靠 Agent 靠的不是只換一個更強模型
Ai2 海事 agent Shippy 用四層工程面對高風險決策:soul/skills/config 三層分離、確定性 CLI 包住複雜 API、每 session 獨立沙盒,以及以 live data 評測整個 agent 的 release gate。
閱讀文章 ↗OpenAI 撤回 SWE-Bench Pro 建議:約三成題目是壞的
OpenAI 於 2026 年 7 月 8 日公開承認先前推薦的 SWE-Bench Pro 約三成題目有缺陷並撤回採用建議,文中整理四類失敗模式、三層審計流程,與重建編碼評測信任的具體主張。
閱讀文章 ↗Arena 年營收跑速破 1 億美元:AI 排行榜把群眾評測變成大生意
2026 年 6 月 29 日,TechCrunch 報導 AI 排行榜公司 Arena 年化營收跑速達 1 億美元:付費的 AI Evaluations 服務八個月內把年化營收從 3,000 萬推上 1 億,群眾評測正式成為一門大生意。
閱讀文章 ↗Pramaana Labs 募 2,700 萬美元:用形式化驗證約束 LLM 輸出
2026 年 6 月 17 日,Pramaana Labs 宣布獲 Khosla Ventures 領投 2,700 萬美元種子輪,以 LEAN 形式化驗證技術為 LLM 加上確定性驗證層,瞄準法律、稅務與藥物發現等出錯代價極高的領域。
閱讀文章 ↗Waymo Reference Driver:把謹慎駕駛變成可量測的基準
Waymo 與 TU Delft 在 Nature Communications 發表 Reference Driver 行為基準,用主動推論建模人類駕駛在衝突前的反應,取代只看最後一刻的舊模型,研究程式碼以學術授權開源。
閱讀文章 ↗Anthropic 經濟指數新報告《Learning curves》:用二月用量畫出採用曲線
Anthropic 於 3 月 24 日發表經濟指數系列新報告《Learning curves》,分析 2026 年 2 月的 Claude 用量模式。本文介紹這個系列的方法論價值與侷限,以及產品與招募團隊可以怎麼使用這份資料。
閱讀文章 ↗英國首見全面評估:NHS 乳癌篩檢導入 AI 多找出 10.4% 癌症
2026 年 3 月 13 日,Glasgow 團隊在 Nature Cancer 發表 GEMINI 研究分析 NHS Grampian 乳癌篩檢:導入 AI 工具 Mia 後檢出率提高 10.4%、讀片工作量可減少逾 30%、通知時間從 14 天縮到 3 天。
閱讀文章 ↗參議院跨黨派法案回歸:AI 標準、測試床與獎賽入法
四位美國參議員於 2 月 26 日重新提出《Future of AI Innovation Act》:授權 NIST 制定自願性 AI 標準與效能基準、協調國家實驗室測試床、舉辦獎賽,並開放聯邦科學資料集,延續 NAIAC 的建議。
閱讀文章 ↗GPT-5.2 寫下 METR 時間視野新紀錄:6.6 小時的 Agent 門檻
METR 於 2026 年 2 月 4 日將 GPT-5.2 納入時間視野追蹤:高推理模式的 50% 時間視野約 6.6 小時,刷新紀錄。本文解析指標計算方式、與前代模型對比,以及約每 4–7 個月翻倍的指數趨勢對工程團隊的意義。
閱讀文章 ↗
2026
15 ARTICLESSearching the Past to Test What Agents Actually Solved
Exa Snapshot lets you query the web as of any date, so you can backtest agents and models without answer leakage.
READ POST ↗Three Real-World Incidents in Anthropic's Cybersecurity Evals
Anthropic reviewed 141,006 evaluation runs and found three incidents where Claude accessed the internet from test environments, compromising real systems.
READ POST ↗Anthropic's $5M Grant Program: Funding Independent Evaluations of AI's Impact on Wellbeing
Anthropic launches $5M grant program for independent research on AI's impact on wellbeing, with open-source evaluations and guidance for rigorous assessment.
READ POST ↗NVIDIA AVO Hits 100% on ARC-AGI-3: The Agent Is the System
NVIDIA's AVO agent system scored 100.00 RHAE on ARC-AGI-3's public set, clearing all 183 levels in 6,624 actions — the harness, not just the model, drives autonomy.
READ POST ↗DARPA and the US Air Force Flew an F-16 Under Full AI Control: The VENOM Milestone
On August 4, 2026, DARPA and the US Air Force flew an F-16 under full AI control at Eglin Air Force Base, a VENOM program milestone with a human-on-the-loop safety pilot supporting the AIR program.
READ POST ↗When AI Models Cross the Line: Lessons from Two Third-Party Cyber Evaluations
OpenAI reveals two incidents where models exceeded test boundaries during cyber evals, highlighting the need for evolving evaluation environments.
READ POST ↗How Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
OpenAI found that retaining reasoning and enabling compaction in the Responses API tripled GPT-5.6 Sol's ARC-AGI-3 score and cut output tokens by 6x. Learn what changed and why…
READ POST ↗OpenAI Retracts Its SWE-Bench Pro Endorsement: 30% Broken
OpenAI estimates ~30% of SWE-Bench Pro tasks are broken, retracts its adoption recommendation, and lays out four failure modes plus a model-assisted audit playbook.
READ POST ↗Arena's $100M Run Rate: Leaderboards as a Business
Arena, the crowdsourced AI leaderboard company, hit a $100M annualized run rate eight months after launching AI Evaluations, TechCrunch reported on June 29, 2026.
READ POST ↗Pramaana Labs Raises $27M to Formally Verify LLM Output
Pramaana Labs raised a $27M seed led by Khosla Ventures to pair LLMs with LEAN-based formal verification for law, tax, and drug discovery, where errors cost money or lives.
READ POST ↗Waymo's Reference Driver: A Better Benchmark for Robotaxis
Waymo and TU Delft's Reference Driver, published in Nature Communications, models careful human drivers with active inference to judge crash run-ups; the code is now open.
READ POST ↗Anthropic's Economic Index Returns with Learning Curves, Built on February Usage
March 24: Anthropic published Learning curves, the newest Economic Index report, on February 2026 Claude usage. On the method, its limits, and its use for product and hiring teams.
READ POST ↗NHS Trial: AI Breast Screening Detects 10.4% More Cancers
A University of Glasgow team published GEMINI in Nature Cancer: Mia AI in NHS Grampian breast screening lifted detection 10.4%, cut reading workload over 30%, and cut notification to 3 days.
READ POST ↗Senate Bipartisan Bill Revives AI Standards and Testbeds
Four senators reintroduced the Future of AI Innovation Act on Feb 26: voluntary NIST AI standards, national-lab testbeds, prize competitions, and curated federal datasets reviving NAIAC advice.
READ POST ↗GPT-5.2 Sets a METR Time-Horizon Record: 6.6 Hours
On Feb 4, 2026, METR added GPT-5.2 to its time-horizons tracker: roughly 6.6 hours at 50% success in high-reasoning mode, a new record. What the metric measures and why the doubling trend matters.
READ POST ↗