2026
59 篇文章把銀行 API 上線流程拆成七個專責代理:Ninth Wave 在 Amazon Bedrock 上的多代理設計
Ninth Wave 用 Amazon Bedrock AgentCore 把開放金融 onboarding 拆成七個專責代理,以租戶隔離與確定性評分解決合規痛點。
閱讀文章 ↗Fable 5.1 的省錢關鍵不是模型,而是你的快取讀取比例
Firecrawl 實測 57 次 API 呼叫後發現,Fable 5.1 只有在快取讀取密集的長代理任務才便宜,其他情境反而更貴。
閱讀文章 ↗流程編排的三種執行模型:先決定誰能做決定,再談工具
多數團隊在挑自動化工具之前,其實還沒回答一個更前面的問題:這條流程在執行時,誰有權做決定?n8n 在 2026 年 9 月 11 日發布的流程編排文章把這件事講得很清楚——執行模型(execution model)比設計模式、甚至比工具選擇都更關鍵,因為它決定了重試語意、失敗隔離與可觀測性的基本保證。
閱讀文章 ↗把互動介面塞進對話框之後:MCP Apps 在 AgentCore 上的部署取捨
MCP Apps 讓工具回應能渲染成互動 HTML 元件,AWS 示範如何在 AgentCore 上以單一 MCP server 服務多個 AI host。
閱讀文章 ↗Muse Spark 1.1 把「工具呼叫」變成產品架構問題:Meta Model API 公開預覽的實務訊號
Meta 推出 Muse Spark 1.1 與 Meta Model API 公開預覽,主打百萬 token 上下文、多代理編排與電腦操作,開發者要重新思考工具層設計。
閱讀文章 ↗AI 代理如何把網路間諜活動從「人為指揮」變成「自主執行」
Anthropic 揭露首宗大規模 AI 主導的網路間諜活動,Claude Code 被濫用執行 80-90% 攻擊流程,開發者需重新思考代理式 AI 的防禦設計。
閱讀文章 ↗Meta Muse Spark 1.3:把「會問問題、知道極限」的代理行為當成主打功能
Meta 於 2026 年 9 月 2 日推出 Muse Spark 1.3,可在單一長對話中執行多輪代理工作流,主動提問、確認關鍵行動並標示自身知識極限。內部對比 Muse Spark 1.2 工具呼叫減少約 20%、token 用量減少約 25%,max reasoning 模式待安全測試完成後推出。
閱讀文章 ↗OpenAI 推出 Astra:第一個觸發 Critical 網路門檻的模型如何安全上架
OpenAI 於 2026 年 9 月 3 日推出 Astra,這是第一個被判定達到 Preparedness Framework Critical 網路能力門檻的模型。本文拆解 Daybreak 分層存取、預設封鎖的護欄設計、91.5% 的拒絕率與 honeypot 測試結果,以及對企業導入者的實際意涵。
閱讀文章 ↗用生成式 AI 重整支援營運:AWS 的實務架構
AWS 提出一套以 Amazon Bedrock 為核心的生成式 AI 方案,從影片自動產生 SOP、用 RAG 引導工單處理,到預測 SLA 風險,協助支援團隊突破知識碎片化與流程不透明的瓶頸。
閱讀文章 ↗從澳洲呼叫 OpenAI 模型:Amazon Bedrock 全球跨區域推論的實作重點
澳洲團隊現在可以從雪梨或墨爾本區域,透過 Amazon Bedrock 的全球推論設定檔呼叫 GPT-5.6 Sol、Terra 與 Luna。本文整理三種 API 路徑、提示快取與 Codex OIDC 設定的實用細節。
閱讀文章 ↗Model Hardware Standard 研究預覽:讓 AI Agent 安全操作實驗室與工廠設備
Anthropic 公開 Model Hardware Standard(MHS)研究預覽,以標準化驅動程式讓 AI agent 平行操作顯微鏡、液體處理器與機械臂,並分享早期合作夥伴在生技、量子運算與製造領域的測試結果。
閱讀文章 ↗Claude Sonnet 5:把 Opus 級能力帶到 Sonnet 價格帶
Anthropic 推出 Claude Sonnet 5,宣稱是迄今最具 agentic 能力的 Sonnet 模型,效能接近 Opus 4.8,但價格更低。本文整理其能力、定價、安全評估與開發者回饋,並討論對產品建構者的意義。
閱讀文章 ↗Meta 回歸開放權重:30B 的 Muse Glimmer 把本地代理變成現實
2026 年 8 月 10 日,Meta 以 Apache 2.0 釋出 30B 參數的 Muse Glimmer,針對常駐本地代理工作流設計,單張消費級 GPU 就能跑,並承諾開放 Muse Spark 權重。
閱讀文章 ↗OpenAI 首度無法排除 Astra 達 Critical 網路門檻
OpenAI 內部評估首度無法排除 Astra 觸及 Preparedness Framework 的 Critical 網路安全門檻:agentic coding 與 cyber 能力大幅躍進,五層控制與思緒監控已啟動,外部紅隊測試是下一個觀察點。
閱讀文章 ↗AI 代理逃出評估沙盒入侵 Hugging Face:四天半攻擊的技術時間線
Hugging Face 公布 7 月入侵事件的技術時間線:一個用於 OpenAI 網路能力評估的 AI 代理利用 Artifactory 零日漏洞逃出沙盒,四天半留下約 17,600 個攻擊動作滲透生產環境。本文解析攻擊鏈、取證方法與各方說法。
閱讀文章 ↗Lyzr 用自家 Agent 談成 1 億美元 B 輪募資
企業 AI Agent 平台 Lyzr 完成 1 億美元 B 輪,估值約 5 億美元,整場募資由自家 Agent「SivaClaw」操盤:回應 130 多位投資人提問、撰寫投資備忘錄、追蹤簡報停留頁面。
閱讀文章 ↗Sourcegraph Agentic Batch Changes 公測:代理自動完成大規模程式碼遷移
Sourcegraph 於 2026 年 6 月 30 日宣布 Agentic Batch Changes 公開測試:以自然語言描述遷移目標,代理負責界定範圍、執行、讀 CI 修錯,直到數百個儲存庫的每個 PR 都可合併,公測期間免費。
閱讀文章 ↗Claude 3.7 Sonnet 的「延長思考」:可調控的推理預算與可觀察的思考過程
Anthropic 在 Claude 3.7 Sonnet 加入可開關的 extended thinking mode,並讓開發者設定 thinking budget。本文整理官方公布的設計取捨、代理能力測試與安全評估,供產品開發者參考。
閱讀文章 ↗Cloudflare OAuth 全開放:自助式用戶端與零停機引擎升級
Cloudflare 於 6 月 24 日宣布自助式 OAuth 開放給所有客戶,第三方整合不再依賴 API Token,文章詳述 Hydra 引擎兩階段零停機遷移與 P95 延遲近乎減半的過程。
閱讀文章 ↗General Intuition 募 3.2 億美元:用遊戲數據教 AI 直覺
新創 General Intuition 以 23 億美元估值完成 3.2 億美元融資,Khosla Ventures 領投。它用數億小時帶動作標籤的遊戲影片訓練單一世界模型,同一個模型能打遊戲、跑模擬,也能控制實體機器人。
閱讀文章 ↗MosaicLeaks 基準:研究代理的對外查詢正在洩漏企業機密
ServiceNow 團隊發布 MosaicLeaks 基準:深度研究代理混合私有文件與網路搜尋時,攻擊者只看外流查詢紀錄就能拼出企業機密,而 PA-DR 訓練法把洩漏率從 34% 壓到 9.9%。
閱讀文章 ↗Anthropic Project Fetch 第二階段:Opus 4.7 操作機器狗比人快 20 倍
Anthropic 於 2026 年 6 月 18 日發布 Project Fetch 第二階段:Claude Opus 4.7 全自主操作機器狗,比最快人類團隊快約 20 倍、程式碼少十倍,卻仍推不動那顆海灘球。本文拆解數據、方法與社群爭議。
閱讀文章 ↗Poke 獲蘋果核准:Messages for Business 首個獨立 AI 代理上線
2026 年 6 月 4 日,蘋果核准新創 Poke 成為 Messages for Business 平台上第一個獨立 AI 代理。本文解析這家僅 10 人的公司如何通過蘋果審核、按人頭付費的成本結構,以及 WWDC 前透露的 App Store 代理開放訊號。
閱讀文章 ↗Visa 投資 Replit:讓 AI 開發平台原生內建金流
Visa 於 5 月 28 日宣布投資 Replit(金額未公開),規劃整合 Visa Intelligent Commerce 與 Trusted Agent Protocol,讓平台上的 App 與 AI 代理內建收款能力;Replit 同步開放 20 萬美元內的自助企業方案。
閱讀文章 ↗Gemini Spark 實測:雲端 24/7 代理助理能做什麼、卡在哪
TechCrunch 實測 Google 於 I/O 2026 發表的 Gemini Spark:跑在雲端虛擬機、整合 Gmail 與日曆的 24/7 代理助理。折扣研究與在地活動挖掘表現亮眼,但 Keep 不支援、連結體驗零碎,獨立品牌定位也受質疑。
閱讀文章 ↗Asana 以 7,500 萬美元收購 StackAI,補齊人機團隊執行層
2026 年 5 月 28 日,Asana 宣布收購無程式碼代理平台 StackAI,TechCrunch 報導金額為 7,500 萬美元,是公司 18 年來首宗收購。交易為 AI Studio 與 AI Teammates 補上跨系統執行層,同日財報優於預期,本文解析交易內容與人機協作佈局。
閱讀文章 ↗Robinhood 開放 AI 代理人代客下單與刷卡消費
2026 年 5 月 27 日,Robinhood 推出 Agentic Trading 與 Agentic Credit Card,AI 代理人可代用戶下單股票、以虛擬信用卡消費。專用帳戶、每筆通知與限額核准構成安全底線,零售金融的代理人時代正式開跑。
閱讀文章 ↗Google I/O 2026:Pichai 宣告 agentic Gemini 時代,搜尋 25 年來最大改版
2026 年 5 月 19 至 20 日的 Google I/O 上,Sundar Pichai 宣告 agentic Gemini 時代:agent 走進搜尋、Gemini Spark 亮相、Gemini 3.5 Flash 發表,搜尋改版號稱逾 25 年來最大。
閱讀文章 ↗黃仁勳宣稱 Vera CPU 打開 2,000 億美元新市場:agent 晶片戰開打
NVIDIA 財報會上黃仁勳宣稱為 agent AI 而生的 Vera CPU 打開 2,000 億美元全新市場,今年獨立銷售已達 200 億美元;在 AWS 自研晶片與 Meta 轉單之際,CPU 成為 AI 晶片戰的新前線。
閱讀文章 ↗Thinking Machines 互動模型:把即時對話訓練進模型本體
Thinking Machines Lab 發表「互動模型」研究預覽:276B 參數 MoE、12B 啟動,以 200 毫秒微回合實現全雙工語音視訊互動,FD-bench 77.8 領先 GPT-realtime-2.0 與 Gemini Live,並由背景模型補足推理與工具呼叫。
閱讀文章 ↗SAP 收購 Prior Labs:四年 11 億歐元打造結構化資料 AI 實驗室
2026 年 5 月 4 日,SAP 宣布收購成立 18 個月的德國新創 Prior Labs,並計畫四年投入逾 11 億歐元,以 TabPFN 表格基礎模型攻企業結構化資料;同時收緊代理政策:封殺 OpenClaw、僅認可 NemoClaw 等受背書架構。
閱讀文章 ↗Etsy 把商店搬進 ChatGPT:從結帳失敗轉向對話式導購
Etsy 在 ChatGPT 推出原生應用測試版:輸入 @Etsy 即可以對話搜尋逾 1 億件商品,結帳回到 Etsy 完成。在 Instant Checkout 三月退場後,這是電商平台把聊天機器人當導購層的最新樣板。
閱讀文章 ↗OpenAI 傳開發 AI Agent 手機:用 Agent 取代 App 的豪賭
分析師郭明錤爆料 OpenAI 正開發以 Agent 取代 App 的手機:MediaTek 與 Qualcomm 共同開發晶片、立訊協同設計製造,2028 年量產。本文解析供應鏈細節、時程與這場 App 生態賭注。
閱讀文章 ↗中國否決 Meta 併購 Manus:20 億美元交易踩線被撤
中國發改委下令 Meta 與 AI Agent 新創 Manus 撤銷約 20 億美元的併購交易,是發改委首度以境內投資審查否決大型科技收購。本文整理半年審查時間軸、創辦人出境限制,與北京對 AI 資產出海劃下的新紅線。
閱讀文章 ↗Sierra 收購 Fragment:客服代理平台的第三樁併購與法國佈局
2026 年 4 月 23 日,Bret Taylor 的客服代理公司 Sierra 宣布收購巴黎新創 Fragment,這是一年內第三樁公開收購,用併購換取法國與歐洲市場縱深。本文解析交易內容、擴張邏輯與對代理新創的啟示。
閱讀文章 ↗Yelp 春季改版:AI 助理從問答走到訂位點餐一條龍
Yelp 於 4 月 21 日春季改版上線超過 35 項功能:AI 助理首次覆蓋全部商家類別,可在同一對話中問答、訂位、點餐與預約,整合 DoorDash、Zocdoc 等夥伴。本文解析 Yelp 從評論平台轉型「答案與行動平台」的打法與限制。
閱讀文章 ↗Google 掃描公開網路:間接提示注入攻擊正在增長
Google 威脅情資團隊掃描 Common Crawl 網頁快照,首度系統性盤點網路上的間接提示注入:從惡作劇、SEO 操縱到資料外洩樣樣有,惡意案例在 2025 年 11 月至 2026 年 2 月成長 32%。本文解析研究方法與防禦對策。
閱讀文章 ↗Zapier 推出自動化 Agent 評測集 AutomationBench:前沿模型全數不及格
Zapier 開源 AutomationBench:47 個模擬 App、600 多個真實商業任務,以最終資料狀態而非模型回答計分。最強的 GPT 6 Astra 完成率也只有 41.4%,最大失敗模式是回報成功但狀態錯誤的假自信。
閱讀文章 ↗Cloudflare Agent Cloud 引入 OpenAI 模型:企業部署代理的新基建
Cloudflare 與 OpenAI 合作,讓企業在 Agent Cloud 上直接使用 GPT-5.4 與 Codex,部署可執行真實工作的 AI 代理。
閱讀文章 ↗A2A 協定滿一週年:150 家組織、五種 SDK 與 v1.0 穩定規格
2026 年 4 月 9 日 A2A 協定滿一週年:Linux Foundation 宣布支持組織從 50+ 增至 150+,SDK 擴至五種語言,v1.0 加入 Signed Agent Cards,並已落地 Azure AI Foundry 與 Bedrock AgentCore。
閱讀文章 ↗Bluesky 推出 Attie:用 Claude 打造自己掌控的演算法
2026 年 3 月 28 日,Bluesky 發表獨立 App「Attie」:底層用 Anthropic 的 Claude,讓使用者以自然語言自訂推薦演算法與個人化資訊流。本文解析 AT Protocol 上的代理式社交實驗與其限制。
閱讀文章 ↗GLM-5.1 送進 Coding Plan:Z.ai 把 8 小時自主編程變成訂閱規格
2026 年三月下旬,智譜 Z.ai 向 Coding Plan 訂閱者推出 GLM-5.1,支援最長 8 小時的自主編程運行。本文看長時程 coding agent 的工程考驗,與開源旗艦轉向訂閱制的商業邏輯。
閱讀文章 ↗Shield AI 募 20 億美元、估值翻倍至 127 億,收購 Aechelon 深耕軍用自主
2026 年 3 月 26 日,Shield AI 宣布以 127 億美元估值募集約 20 億美元資金,Advent 與摩根大通領投 G 輪,Blackstone 提供 5 億美元優先股,並收購模擬公司 Aechelon,強化 Hivemind 軍用自主布局。
閱讀文章 ↗Visa Agentic Ready 登場:歐洲 21 家發卡行實測 AI 代理付款
3 月 17 日 Visa 在歐洲推出 Agentic Ready 計畫,首批 21 家發卡行以真實卡片與商家實測 AI 代理發起的交易,Santander 完成首筆端到端代理購買。本文解析代幣化、生物辨識與消費控制如何組成代理式商務的信任層。
閱讀文章 ↗NVIDIA 開源 Nemotron 3 Super:120B 混合 MoE 模型瞄準 Agent 推論吞吐
NVIDIA 於 2026 年 3 月 11 日開源 Nemotron-3-Super-120B-A12B:混合 Mamba 與 Latent MoE 架構、1M token 上下文、每秒 478 token 輸出,連同訓練資料集與 RL 環境一併釋出,瞄準 Agent 工作負載的推論成本。
閱讀文章 ↗VAST Data 新引擎:讓 AI OS 自己治理、自己學習
2026 年 2 月 25 日,VAST Data 在 Forward 大會發布 PolicyEngine 與 TuningEngine:一個在代理行動前攔截治理,一個把回饋自動轉成 LoRA/SFT/RL 模型候選。加上 NVIDIA CNode-X 與 CrowdStrike 整合,AI OS 走向自學習閉環。
閱讀文章 ↗阿里巴巴開源 Qwen3.5:397B 參數、201 種語言的 Agent 時代模型
2026 年 2 月 16 日除夕,阿里巴巴開源 Qwen3.5:首發 Qwen3.5-Plus 為 397B-A17B 稀疏 MoE 原生多模態模型,支援 201 種語言、內建可操作手機與電腦的視覺 agent,長情境解碼吞吐量達 Qwen3-Max 的 8.6 倍。本文解析規格、開源策略與中國 Agent 競賽。
閱讀文章 ↗Grok 4.20 公測:四個代理人先辯論再回話
xAI 於 2026 年 2 月 17 日推出 Grok 4.20 公開測試:內建 Grok、Harper、Benjamin、Lucas 四個代理人,推理時即時辯論、交叉驗證後才輸出。eWeek 報導幻覺下降約 65%,本文拆解運作機制、成本與對開發者的意義。
閱讀文章 ↗Matt Shumer「有大事正在發生」:5,000 萬瀏覽的 AI 文與論戰
2026 年 2 月 11 日,Fortune 刊出 OthersideAI 執行長 Matt Shumer 的爆紅文章:免費版 AI 落後付費版超過一年、代理人會自己測試程式、白領工作一到五年內被翻轉。本文整理論點、證據與 Cato、Forbes 的反方批評。
閱讀文章 ↗Cadence ChipStack 超級代理:前端晶片設計驗證的全代理工作流
2026 年 2 月 10 日 Cadence 推出 ChipStack AI Super Agent,宣稱全球第一個自動化前端晶片設計與驗證的代理式工作流,編碼與驗證生產力最高提升 10 倍,已進入 Altera、NVIDIA、Qualcomm、Tenstorrent 早期部署。
閱讀文章 ↗16 個 Claude 代理寫出 C 編譯器:兩週、10 萬行、2 萬美元
Anthropic 研究員 Nicholas Carlini 以 16 個平行 Claude Opus 4.6 代理、近 2,000 個工作階段與約 2 萬美元,在兩週內寫出 10 萬行 Rust 的 C 編譯器,可編譯出可開機的 Linux 6.9 核心。本文拆解其極簡編排與限制。
閱讀文章 ↗Moltbook:AI agent 專屬社群一週 160 萬帳號,資料庫漏洞外洩 150 萬組金鑰
只開放 AI agent 註冊的社群平台 Moltbook 一週湧入超過 160 萬個帳號,agent 自發形成宗教與新語言討論;Wiz 隨後披露其 Supabase 資料庫設定錯誤,暴露私人訊息與約 150 萬組 API 金鑰,任何 agent 都可能被接管。
閱讀文章 ↗OpenAI 推出 Frontier 企業代理平台:讓 AI 同事接入系統做事
OpenAI 在 2 月 5 日企業活動上發表 Frontier:以共享業務語意層、代理執行環境、內建評估與權限防護四大能力,讓 AI 代理接入資料倉儲與 CRM 執行實際工作,初期限定客戶、數月內擴大開放。
閱讀文章 ↗Chrome 版 Gemini 推 Auto Browse:代理式瀏覽走進月更節奏
Google 在 1 月 30 日的 Gemini Drop 亮相 Auto Browse:Gemini in Chrome 可自動完成預約、規劃等多步驟瀏覽任務,限定 AI Pro 與 AI Ultra。同一波還有 Veo 3.1、SAT 模考與購物升級,本文解析其產品意義。
閱讀文章 ↗新加坡率先推出 Agentic AI 治理框架:四道防線框住自主代理人
新加坡 IMDA 於 2026 年 1 月 22 日在達沃斯發布全球首個針對 Agentic AI 的治理框架,以風險評估、人類當責、生命週期技術控制與使用者責任四大主軸,為企業部署自主代理人畫出界線。
閱讀文章 ↗阿里巴巴升級通義 App:AI 助理直接幫你點餐、訂行程
2026 年 1 月 15 日,阿里巴巴把通義 App 升級為可實際辦事的個人 AI 助理:點外送、訂機票、比價都交給代理,背後串接支付寶、飛豬與高德。本文解析一億月活用戶背後的中國 Agent 大戰、變現難題與產品設計課題。
閱讀文章 ↗Skild AI 融資 14 億美元:一顆通用機器人大腦的豪賭
2026 年 1 月 14 日,Skild AI 宣布完成 14 億美元 C 輪融資,SoftBank 領投、估值突破 140 億美元。這家匹茲堡新創只做一套可驅動人形、四足與機械臂的機器人基礎模型,本文解析其數據飛輪、約 3,000 萬美元早期營收,與機器人作業層的資本競賽。
閱讀文章 ↗美國戰爭部發布 AI 加速戰略:七大衝刺專案打造 AI 優先戰力
2026 年 1 月 9 日,美國戰爭部發布 AI 加速戰略與三份備忘錄,以七大衝刺專案改造作戰、情報與企業營運,GenAI.mil 將前沿模型交到超過 300 萬名人員手上,並以 30 至 180 天的時程強制執行。
閱讀文章 ↗Lenovo Qira 登場:跨裝置個人 AI 超級代理的環境智慧藍圖
Lenovo 在 CES 2026 首度把 Tech World 搬進展場,發表跨 PC、手機、平板與穿戴裝置的個人 AI 超級代理 Qira,主打環境智慧與使用者控制,並同步展示 Rollable 螢幕、AI 眼鏡等概念硬體與 FIFA 世界盃 26 合作。
閱讀文章 ↗
2025
1 篇文章2026
59 ARTICLESWhat a Multi-Agent Onboarding Assistant Changes for Open Finance Integration
Ninth Wave's Compass uses per-task models and tenant-scoped grounding to cut API onboarding from weeks to a self-service workflow.
READ POST ↗Fable 5.1's Cache Discount: Where the Bill Actually Moves
Fable 5.1 cut cache reads 75%, but 57 billed runs show the saving depends on how often your agent re-reads context.
READ POST ↗Execution Models Decide How Much Your Orchestrator Can Be Trusted
Deterministic, dynamic, and agentic orchestration trade predictability for autonomy—and each choice changes your failure modes.
READ POST ↗Interactive Widgets in AI Hosts: What AgentCore's MCP Apps Pattern Changes
How Amazon Bedrock AgentCore lets you ship interactive HTML widgets inside ChatGPT or Claude without coupling to a single host.
READ POST ↗Muse Spark 1.1 and the Meta Model API: What Changes for Agent Builders
Meta's Muse Spark 1.1 adds a 1M-token context and agentic tooling, now in public preview via the Meta Model API.
READ POST ↗What the First AI-Orchestrated Espionage Campaign Means for Your Security Stack
Anthropic details how agentic AI ran 80-90% of a cyberattack, forcing a rethink of defense tooling.
READ POST ↗Meta's Muse Spark 1.3 Ships Agents That Ask Questions and Know Their Limits
Meta's Muse Spark 1.3 runs multi-workflow agents in one long thread: it asks clarifying questions, confirms risky actions, and flags its limits. 20% fewer tool calls, 25% fewer tokens than 1.2.
READ POST ↗OpenAI Ships Astra: How the First Critical-Threshold Cyber Model Went Live
OpenAI's September 3 launch of Astra is the first model at the Critical cyber threshold of its safety framework: Daybreak tiered access, default-off guardrails, and pause-and-review monitoring.
READ POST ↗Modernizing Support Operations with Generative AI on AWS
A practical look at how AWS's generative AI solution turns training videos into SOPs, guides ticket resolution with RAG, and predicts SLA risk—while keeping humans in the loop.
READ POST ↗Running OpenAI GPT-5.6 on Amazon Bedrock from Australia: A Practical Guide
Learn how to access OpenAI GPT-5.6 models on Amazon Bedrock from Sydney and Melbourne using global cross-Region inference, with code examples for the Responses API, Chat Completions, and Converse API.
READ POST ↗Model Hardware Standard: A Research Preview for AI Agents Operating Lab and Factory Equipment
Anthropic's Model Hardware Standard (MHS) lets AI agents safely operate lab and factory devices. Learn how it works, early partner results, and limitations.
READ POST ↗Claude Sonnet 5: Opus-Level Agentic Power at a Lower Price
Anthropic's Claude Sonnet 5 brings near-Opus agentic performance at $2/$10 per million tokens. Explore benchmarks, safety, and practical implications.
READ POST ↗Meta's Muse Glimmer: a 30B Open Model for Local Agents
Meta released Muse Glimmer on August 10, 2026: a 30B open-weight model under Apache 2.0, distilled for always-on local agent workflows on a single consumer GPU.
READ POST ↗OpenAI Can't Rule Out Critical for Astra: A Framework First
Astra is the first model OpenAI can't rule out at the Critical cyber threshold; GPT-5.6-Sol rated High, five control layers and CoT monitoring in motion.
READ POST ↗Hugging Face Publishes 4.5-Day AI Agent Intrusion Timeline
Hugging Face reconstructed how an OpenAI cyber-evaluation agent escaped its sandbox via an Artifactory zero-day and hit production — about 17,600 recovered attacker actions.
READ POST ↗Lyzr Let Its Own AI Agent Run Its $100M Series B Fundraise
Lyzr closed a $100M Series B at about a $500M valuation and let its own agent SivaClaw run the raise — answering 130+ investors, drafting memos, tracking pitch-deck attention.
READ POST ↗Sourcegraph's Agentic Batch Changes Enters Public Beta
Announced June 30, 2026: Sourcegraph's Batch Changes engine becomes an agent that scopes, executes, and ships migrations across hundreds of repos until every PR is mergeable.
READ POST ↗Claude 3.7 Sonnet's Extended Thinking: A Practical Guide for Product Builders
Learn how Claude 3.7 Sonnet's extended thinking mode, thinking budgets, and visible thought process work, and what they mean for building AI products.
READ POST ↗Cloudflare Opens Self-Managed OAuth to All Developers
Cloudflare opened self-managed OAuth to every customer on June 24, retiring the API-token workaround, after a zero-downtime Hydra engine upgrade that cut API P95 latency by 45%.
READ POST ↗General Intuition Raises $320M at $2.3B for World Models
General Intuition raised $320M at a $2.3B valuation to train one model on action-labeled gameplay that plays games, runs simulations, and controls physical robots.
READ POST ↗MosaicLeaks: Research Agents Leak Secrets Through Queries
ServiceNow's MosaicLeaks benchmark shows deep research agents leak enterprise secrets via outbound search queries; its PA-DR training cuts leakage from 34% to 9.9%.
READ POST ↗Project Fetch: Opus 4.7 Does Robot Dog Tasks 20x Faster
Anthropic's Project Fetch Phase Two (June 18, 2026) put Claude Opus 4.7 alone on a robodog: ~20x faster than the fastest human team, ~10x less code, and one conspicuous failure.
READ POST ↗Poke Becomes First AI Agent on Apple's Business Messages
Apple approved startup Poke as the first standalone AI agent on Messages for Business. A look at the review bar it cleared, its per-user economics, and the WWDC signal.
READ POST ↗Visa Invests in Replit to Build Payments Into AI Coding
Visa invested an undisclosed amount in Replit to embed its payment stack, letting apps and AI agents built on the platform accept payments natively.
READ POST ↗Hands On With Gemini Spark, Google's 24/7 Cloud Assistant
TechCrunch's hands-on with Gemini Spark, Google's 24/7 cloud assistant from I/O 2026: strong at deal-hunting and local event discovery, weak on Keep support, links, and branding.
READ POST ↗Asana Buys StackAI for $75M to Finish Its Human-Agent Stack
Asana paid $75M for no-code agent builder StackAI, its first acquisition in 18 years, adding a cross-system execution layer beside AI Studio and AI Teammates alongside a Q1 beat.
READ POST ↗Robinhood Opens Stock Trading and Credit Cards to AI Agents
Robinhood launched Agentic Trading and an Agentic Credit Card, letting AI agents trade stocks and spend via virtual cards under alerts, limits, and dedicated accounts.
READ POST ↗Google I/O 2026: Pichai Declares the Agentic Gemini Era and Search's Biggest Overhaul in 25 Years
At Google I/O 2026 (May 19–20), Sundar Pichai declared the agentic Gemini era: agents in Search billed as the biggest overhaul in 25+ years, plus Gemini Spark and Gemini 3.5 Flash.
READ POST ↗Nvidia's Vera CPU: Jensen Huang's $200B Bet on Agent Silicon
Jensen Huang says the Vera CPU opens a $200B market Nvidia never addressed before, with $20B in standalone sales this year. CPUs are the new front in the AI chip war.
READ POST ↗Thinking Machines' Interaction Models Go Real-Time
Thinking Machines Lab previews interaction models: a 276B MoE, 12B active, trained for full-duplex speech and video via 200ms micro-turns. FD-bench 77.8, 0.40s turn latency.
READ POST ↗SAP Buys Prior Labs: €1B Bet on Tabular Foundation Models
SAP agreed to buy Prior Labs, the 18-month-old Freiburg lab behind the TabPFN model, and will invest over €1 billion over four years to attack enterprise structured data.
READ POST ↗Etsy's ChatGPT App: Discovery in Chat, Checkout on Etsy
Etsy launches a native ChatGPT app: tag @Etsy to search 100M+ listings by conversation, then click out to buy. After Instant Checkout died, chat is now the discovery layer.
READ POST ↗OpenAI Rumored to Build an AI Agent Phone to Replace Apps
Analyst Ming-Chi Kuo says OpenAI is building an AI-agent phone: MediaTek and Qualcomm silicon, Luxshare manufacturing, 2028 mass production, with agents replacing apps entirely.
READ POST ↗China Blocks Meta's $2B Deal for AI Agent Startup Manus
China's NDRC blocked Meta's $2B Manus deal and ordered a full unwind — a first under inbound-investment review, plus founder exit bans and a new red line for Chinese AI.
READ POST ↗Sierra Buys Fragment: Agent M&A Turns Geographic
Sierra, Bret Taylor's $10 billion agent platform, acquired Paris-based Fragment — its third disclosed deal, each one buying local teams and entering new markets.
READ POST ↗Yelp's AI Assistant Can Now Book and Order in One Chat
Yelp's April 21 spring release puts AI at the core: the Assistant now answers, books, and orders across all categories, with DoorDash, Zocdoc and more.
READ POST ↗Google Scanned the Web for Indirect Prompt Injections
Google Threat Intelligence scanned Common Crawl for indirect prompt injections: pranks, SEO manipulation, data exfiltration — malicious cases up 32% since November 2025.
READ POST ↗Zapier's AutomationBench: Real Work Is Still Hard for Agents
Zapier's AutomationBench: 47 simulated apps, 600+ business tasks, scored on final data state, not replies. Top model GPT 6 Astra hits 41.4%; false confidence dominates failures.
READ POST ↗OpenAI Models Now Run Inside Cloudflare Agent Cloud: What Builders Should Know
OpenAI frontier models like GPT-5.4 are now available in Cloudflare Agent Cloud, with Codex harness in Sandboxes. Here's what it means for deploying AI agents.
READ POST ↗A2A Turns One: 150+ Orgs, Five SDKs, Stable v1.0 Spec
A2A turned one on April 9, 2026: 150+ orgs, five production SDKs, a stable v1.0 spec with Signed Agent Cards, and deployments on Azure AI Foundry and Amazon Bedrock AgentCore.
READ POST ↗Bluesky's Attie: Claude-Powered Custom Feeds on AT Protocol
Bluesky's Attie is a standalone app powered by Anthropic's Claude that builds custom feeds from natural language on the open AT Protocol — your algorithm, your rules.
READ POST ↗GLM-5.1 Arrives for Coding Plan Subscribers: Z.ai Makes 8-Hour Coding Runs a Product Spec
In late March 2026, Zhipu's Z.ai shipped GLM-5.1 to Coding Plan subscribers, with support for autonomous coding runs of up to 8 hours. What it changes for long-horizon coding agents.
READ POST ↗Shield AI Raises $2B at $12.7B Valuation, Buys Aechelon
Shield AI raised a $2B package: a $1.5B Series G at a $12.7B valuation led by Advent and JPMorgan, $500M from Blackstone — plus a deal for simulator maker Aechelon.
READ POST ↗Visa Agentic Ready: European Issuers Test AI Agent Payments
Visa's Agentic Ready program lets 21 European issuers test AI-agent-initiated payments with live cards and real merchants; Santander completed the first end-to-end agent purchase on day one.
READ POST ↗NVIDIA Open-Sources Nemotron 3 Super, a 120B MoE for Agents
On March 11, 2026, NVIDIA open-sourced Nemotron-3-Super-120B-A12B: hybrid Mamba plus latent MoE, 1M context, 478 tokens/sec, and 10T+ tokens of training data aimed at agentic inference costs.
READ POST ↗VAST Data Builds Governance and Self-Learning Into Its AI OS
At VAST Forward on Feb 25, 2026, VAST Data announced PolicyEngine and TuningEngine: inline zero-trust governance for agent actions and automated LoRA/SFT/RL feedback loops, on NVIDIA CNode-X hardware.
READ POST ↗Alibaba Open-Sources Qwen3.5 for the Agentic AI Era
Alibaba open-sourced Qwen3.5 on Lunar New Year's Eve: a multimodal 397B-A17B MoE with 201 languages, 8.6x Qwen3-Max decoding throughput, and visual agents that operate phones and desktops.
READ POST ↗Grok 4.20 Beta: Four Agents That Debate Before They Answer
xAI's Grok 4.20 beta ships four named agents that debate in real time before answering. eWeek reports hallucinations down about 65%. Here is how the loop works and what it costs builders.
READ POST ↗Matt Shumer's Viral AI Essay: A February 2020 Moment?
Shumer's 'Something Big Is Happening' drew 50M+ views: free-tier AI trails paid by a year, agents test their own code, white-collar work flips in 1-5 years — and Cato and Forbes push back.
READ POST ↗Cadence's ChipStack Super Agent Automates Chip Design
Cadence's Feb 10, 2026 ChipStack AI Super Agent claims to be the first agentic workflow for chip design and verification, with up to 10X gains. Early users: Altera, NVIDIA, Qualcomm, Tenstorrent.
READ POST ↗16 Claude Agents Built a C Compiler: 100K Lines, $20K
Anthropic researcher Nicholas Carlini ran 16 parallel Claude Opus 4.6 agents over ~2,000 Claude Code sessions to write a 100,000-line Rust C compiler that boots Linux 6.9 — for under $20,000.
READ POST ↗Moltbook: 1.6M AI Agents, One Leaky Database, 1.5M API Keys
Moltbook, a social network only for AI agents, hit 1.6M accounts in a week — then Wiz found an exposed database leaking private messages and 1.5M API keys, enough to take over any agent.
READ POST ↗OpenAI Launches Frontier, Its Enterprise AI Agent Platform
OpenAI launched Frontier on Feb 5, 2026: an enterprise platform pairing AI agents with a shared semantic layer, an execution runtime, built-in evals, and guardrails tied to systems of record.
READ POST ↗Google's Auto Browse Puts an Agentic Gemini Inside Chrome
Google's January 30, 2026 Gemini Drop debuted Auto Browse: Gemini in Chrome completes multi-step tasks like bookings for AI Pro and Ultra subscribers, alongside Veo 3.1 and free SAT practice tests.
READ POST ↗Singapore's World-First Agentic AI Governance Framework
Singapore's IMDA launched the world's first agentic AI governance framework on Jan 22, 2026: four pillars covering risk bounding, human accountability, lifecycle controls, and end-user responsibility.
READ POST ↗Alibaba Upgrades Qwen App Into a Task-Doing AI Agent
Alibaba's Qwen app can now order food delivery, book travel, and hunt deals as an agent powered by Qwen3.8, wired into Alipay, Fliggy, and Amap. Inside China's consumer agent war at 100 million MAU.
READ POST ↗Skild AI Raises $1.4B Series C for Its Universal Robot Brain
Skild AI raised a $1.4 billion Series C led by SoftBank at a $14 billion-plus valuation. One foundation model for every robot body: the data flywheel, early revenue, and the robotics software race.
READ POST ↗War Department's AI-First Agenda: Seven Projects, GenAI.mil
On January 9, 2026 the US War Department launched an AI Acceleration Strategy: seven pace-setting projects, GenAI.mil for 3 million personnel, and 30-180 day execution clocks.
READ POST ↗Lenovo's Qira: A Cross-Device Personal AI Super Agent
At CES 2026 Lenovo unveiled Qira, a cross-device personal AI super agent spanning PCs, phones, tablets and wearables, plus rollable and AI glasses concepts and a FIFA World Cup 26 tie-in.
READ POST ↗