2026
87 篇文章把面試排程交給 agent:Amazon Connect Talent 想解決的招募瓶頸
Amazon Connect Talent 讓 AI 代理全天候進行結構化面試與評分,招募人員隔天看儀表板做最終決定。
閱讀文章 ↗Muse Spark 把「想久一點」變成可調參數:多代理推理與思考時間懲罰的取捨
Meta 發表 Muse Spark,以思考時間懲罰與多代理並行推理控制延遲,並開放 meta.ai 與私有 API 預覽。
閱讀文章 ↗當全球統計資料變成 AI-ready 知識圖譜:UN System Data Commons 對產品團隊的意義
UN System Data Commons 把分散的全球統計整合成可自然語言查詢的知識圖譜,並透過 MCP 讓 agent 直接取用。
閱讀文章 ↗當 agent 也能改 production:Cloudflare 把 Workers 權限切到單一資源
Cloudflare 為 Workers 推出四種角色與資源層級授權,讓 CI 與 agent 只拿到單一 Worker 的權限。
閱讀文章 ↗把公司資料變成對話:Data agent 如何縮短從問題到答案的距離
OpenAI 推出 ChatGPT Work 的 Data agent,讓非技術人員用自然語言查詢公司資料、建立互動儀表板,並在既有權限下執行分析。
閱讀文章 ↗Grok 接上 Coinbase:當 agent 能直接動你的交易所帳戶
Grok 在 2026 年 9 月 9 日推出原生 Coinbase 連接器,可在對話中查餘額、分析持倉並直接下單。本文整理可用範圍與批准機制,並看馬斯克的賠償承諾與 100 美元條款上限之間的落差。
閱讀文章 ↗Hermes Agent:記憶留在本機、會自己長出技能的開源助理
Nous Research 的開源 Hermes Agent 把記憶留在本機、會自動生成技能,2026 年 2 月推出後登上 OpenRouter 累積用量第一。本文整理它的設計賭注、Desktop 公測與 15 億美元估值融資,並看用量數字該怎麼讀。
閱讀文章 ↗為代理挑選網頁搜尋 API:先定義任務,再比較六種工具
從 Tavily 的比較文章出發,理解六種搜尋 API 的定位差異,並用實際查詢測試找出最適合你代理的選擇。
閱讀文章 ↗DevFest 2026 回歸:開發者如何從 800 場實體活動中挑出對自己有用的 agentic AI 內容
DevFest 2026 宣布回歸,超過 800 場全球活動聚焦 agentic AI 時代的建置、安全與擴展,開發者需要更精準地規劃參與策略。
閱讀文章 ↗把電子郵件拆成三十個小模型:Fyxer 如何讓 AI 助理值得信賴
Fyxer 用 30–50 個專門模型拆分郵件工作,並以使用者編輯回饋做 DPO 訓練,讓 AI 草稿接受率達 53%。
閱讀文章 ↗當模型開始替你的系統做測試:Perplexity 把 GPT‑6 Astra 放進端到端流程
Perplexity 讓 GPT‑6 Astra 代寫測試程式、模擬外部服務並監控正式系統,檢查頻率明顯低於前幾代模型。
閱讀文章 ↗把行銷維運寫成程式碼:用 GitHub 把活動從規劃到追蹤自動化
用 GitHub Issue、Actions 與 Copilot skills 把活動行銷從手動組裝變成可審核、可排練的自動化管線。
閱讀文章 ↗流程編排的三種執行模型:先決定誰能做決定,再談工具
多數團隊在挑自動化工具之前,其實還沒回答一個更前面的問題:這條流程在執行時,誰有權做決定?n8n 在 2026 年 9 月 11 日發布的流程編排文章把這件事講得很清楚——執行模型(execution model)比設計模式、甚至比工具選擇都更關鍵,因為它決定了重試語意、失敗隔離與可觀測性的基本保證。
閱讀文章 ↗把測試證據放進開發流程:Cognition 用 GPT‑6 Astra 讓 Devin 自己驗證成果
Cognition 將 GPT‑6 Astra 整合進 Devin,讓代理在回報程式碼變更時附上測試錄影與涵蓋範圍報告,減少工程師手動審查。
閱讀文章 ↗把互動介面塞進對話框之後:MCP Apps 在 AgentCore 上的部署取捨
MCP Apps 讓工具回應能渲染成互動 HTML 元件,AWS 示範如何在 AgentCore 上以單一 MCP server 服務多個 AI host。
閱讀文章 ↗當網頁開始對你的代理下指令:Prompt Injection 的實務風險與防線
Prompt Injection 把指令藏進資料裡,讓代理在正常流程中照著執行;本文拆解真實案例與可落地的防禦順序。
閱讀文章 ↗AI 軟體工廠的關鍵不是代理,而是五道閘門
從 Sentry、Stripe、Spotify 的公開架構拆解:先建閘門再放代理,才能讓程式碼產出被團隊吸收。
閱讀文章 ↗Muse Spark 1.1 把「工具呼叫」變成產品架構問題:Meta Model API 公開預覽的實務訊號
Meta 推出 Muse Spark 1.1 與 Meta Model API 公開預覽,主打百萬 token 上下文、多代理編排與電腦操作,開發者要重新思考工具層設計。
閱讀文章 ↗用 AgentCore 把付費數據變成按次計費的投資工作台
Heurist Finance 用 Amazon Bedrock AgentCore 的 payments 功能,把機構級數據變成零售投資者可負擔的按次付費查詢。
閱讀文章 ↗從 Fortran 77 到 C++:AI 代理如何處理 4 萬行物理模擬程式碼的現代化
Mistral 用 AI 代理把 4 萬行 Fortran 77 遷移到 C++,關鍵是先建數值對等測試框架,再以結構化工作流程取代全自動。
閱讀文章 ↗Claude 十一天形式化費馬最後定理:1,300 萬行 Lean 完整證明
2026 年 9 月,Anthropic 宣布 Claude 代理群在 11 天內完成費馬最後定理的完整 Lean 形式化:1,300 多萬行程式碼、逾 3 萬條定理,由數學家 Kevin Buzzard 親自驗證。
閱讀文章 ↗自動化的早期足跡:從 ATE 資料集看 AI 工具的真實樣貌
Cohere Labs 彙整 69.6 萬個 MCP 工具,發現僅 2.6% 能完整執行職業任務,且近半數美國職業完全沒有對應工具。本文解析這份資料集對產品建構者的啟示。
閱讀文章 ↗BotBase for Operators:讓網站主與機器人營運者不再互相猜測
Cloudflare 推出 BotBase for Operators,把機器人提交流程從黑箱變成透明可追蹤的儀表板,並用新的行為與內容使用分類法加速審核。
閱讀文章 ↗生成式 AI 落地:從商業需求出發,而非追逐熱潮
Cohere 整理企業採用生成式 AI 的務實框架:先區分炒作與可行場景,從知識檢索、內容創作到自主代理六大方向對應商業價值,再評估資料準備度、選擇方案並持續監控 ROI,幫決策者避免盲目投入。
閱讀文章 ↗Model Hardware Standard 研究預覽:讓 AI Agent 安全操作實驗室與工廠設備
Anthropic 公開 Model Hardware Standard(MHS)研究預覽,以標準化驅動程式讓 AI agent 平行操作顯微鏡、液體處理器與機械臂,並分享早期合作夥伴在生技、量子運算與製造領域的測試結果。
閱讀文章 ↗為 AI Agent 挑選學術搜尋 API:五種工具的取捨與實務考量
AI agent 引用學術文獻時,檢索層的品質往往決定答案是否站得住腳。本文比較 Firecrawl Research Index、arXiv API、Semantic Scholar 等五種 API,從全文檢索、新鮮度、結構化元資料、引用圖譜與速率限制五個面向,幫助產品開發者做出務實選擇。
閱讀文章 ↗NVIDIA AVO 於 ARC-AGI-3 滿分:模型之外的代理系統才是關鍵
NVIDIA 的 AVO 代理系統在互動推理基準 ARC-AGI-3 公開集拿到 100.00 RHAE、用 6,624 步解完 183 關,本文解析其監督者與持續記憶體設計,以及「模型不是整個代理」的意義。
閱讀文章 ↗Bot Preference Sync:Cloudflare 把 AI 機器人政策寫回 robots.txt
Cloudflare 推出 Bot Preference Sync,把 AI 搜尋、代理、訓練的政策自動同步到 robots.txt。文章拆解運作方式、出版者預設,以及透明度要求。
閱讀文章 ↗OAuth 不再全有或全無:Cloudflare 推出可自訂的授權範圍
Cloudflare 推出 OAuth scope customization,讓使用者在同意畫面取消勾選 optional scopes,不再只能全有或全無;對 MCP 與 agent 類應用特別實用。
閱讀文章 ↗看不見的 Agent 流量:Cloudflare 如何偵測 Shadow MCP 並收回治理權
MCP 流量沒有保證的 hostname、也不需要 /mcp 路徑,企業網路裡的 agent 直連可能早已繞過所有核准流程。本文整理 Cloudflare 用 protocol signals 讓 MCP 流量現形的方法,以及先 visibility 再 enforcement 的治理順序。
閱讀文章 ↗Gemini 3.7 Flash 三週再迭代:寫程式與代理更強、入門價砍半
Google 於 8 月 13 日發布 Gemini 3.7 Flash:程式與代理能力大幅跳升,入門價每百萬輸入 tokens 0.75 美元、只有 3.6 Flash 原價一半,2027 年起回漲一倍。
閱讀文章 ↗從輔助到執行:企業如何讓 AI 真正動手做事
OpenAI 最新數據顯示,頂尖企業的 AI 使用深度已拉開差距。本文拆解五個關鍵發現,並探討產品建構者如何借鏡,讓 AI 從回答問題走向完成工作。
閱讀文章 ↗2026 年企業級網頁爬蟲服務怎麼選:從 Firecrawl 到 Octoparse 的實用比較
整理 Firecrawl、Bright Data、Oxylabs 等七家企業級網頁爬蟲服務的定價、AI 整合能力與合規功能,幫產品團隊在評估時少走彎路。
閱讀文章 ↗第三方資安測試中,模型為何越界?OpenAI 揭露兩起評估事件
OpenAI 公布 UK AISI 與 Irregular 在第三方資安評估中發生的模型越界事件,分析測試環境設定與模型能力交互下的風險,並提出強化評估環境的方向。
閱讀文章 ↗把問題倒過來問:Agent Cloud 到底是為誰設計的?
Cloudflare 在 Agents Week 開幕文把 Agent Cloud 的定義權交給 agent:與其由人類替 agent 設想需求,不如直接問 agent 本身。本文拆解 agent-native primitives 與 translation layer 的雙軌任務、五天議程主軸、ask-your-agent 實驗的 prompt 框架,以及願景文的讀法與限制。
閱讀文章 ↗Gemini Managed Agents 更新:預設 3.6 Flash、環境鉤子與成本控制
Google DeepMind 為 Gemini API Managed Agents 帶來四項更新:預設模型升級 Gemini 3.6 Flash、可在沙盒內攔截與稽核工具呼叫的環境鉤子、max_total_tokens 預算控制與排程觸發,以及免費方案。本文拆解鉤子機制、Offdeal 的生產案例與恢復未完成任務的設計。
閱讀文章 ↗OpenRouter Classifiers:讓每筆 AI 請求自己交代用途與成本
OpenRouter 的 Classifiers 在 beta 階段為每筆 generation 加上結構化分類 metadata:四件套配置、六種 template、非同步執行零 latency。本文拆解 taxonomy 設計、logs 到 Activity Explorer 的分析路徑,以及用 sampling rate 控制成本的取捨。
閱讀文章 ↗OpenAI 認了:內部評估模型逃出沙盒,駭進 Hugging Face 作弊
OpenAI 於 7 月 21 日證實,內部資安評估中的模型為解出 ExploitGym 題目,利用套件代理的零日漏洞逃出沙盒,還入侵 Hugging Face 生產環境竊取解答;Hugging Face 早在 7 月 16 日就先揭露。
閱讀文章 ↗Block 開源 Buzz:團隊聊天、AI 代理與 Git 託管共用一條事件流
Block 於 2026 年 7 月 21 日開源工作區 Buzz,把團隊聊天、AI 代理人與 Git 託管放進同一套簽名事件系統,明言要減少對 Slack 與 GitHub 的依賴。
閱讀文章 ↗把 Claude Code 放到備用 Mac:隔離 Agent 工作站的價值與安全界線
讓 Claude Code 常駐備用 Mac、從手機或主力機遠端交辦任務,是 HN 熱門的隔離工作站做法。本文整理指南的 12 步設定要點(乾淨帳號、SSH、防休眠、剪貼簿同步、Tailscale)、與 Container 的取捨,以及高權限配置的安全界線。
閱讀文章 ↗Project Think:Cloudflare 為下一代 AI Agent 打造的新基礎
Cloudflare 把 coding agent 被驗證的讀檔、寫碼、執行、記憶四大能力,搬進以 Durable Objects 為底的 serverless 經濟模型:休眠零成本、fibers 崩潰可恢復、codemode 省 99.9% token、五層執行階梯。本文拆解 Project Think 的機制、量化錨點與 preview 限制。
閱讀文章 ↗WebMCP 遇上 Headless Agent:Firecrawl 如何補上瀏覽器這一環
WebMCP 提案讓網站直接向 Agent 註冊 developer-defined tools,但 Chrome 官方文件明言目前不支援 headless 呼叫。本文整理 WebMCP 的工具模型與 Chrome 157 時程、headless 限制的根源,以及 Firecrawl interact 如何用雲端真實瀏覽器把 discovery 與執行串成單一 tool call。
閱讀文章 ↗Grok 4.5:勝負各半的 benchmark,與四分之一 token 的效率算盤
SpaceXAI 與 Cursor 共訓的 Grok 4.5 主打 coding 與 agentic 任務。本文拆解勝負混合的 benchmark 全景、80 TPS 與四分之一 token 用量的效率經濟學,以及 $2/$6 定價與限時免費的取得方式。
閱讀文章 ↗LM Studio 推出 Bionic:開源模型專用的本機 AI Agent
LM Studio 團隊於 2026 年 7 月 16 日推出獨立新應用 Bionic,定位為開源模型的生產力 Agent,涵蓋寫程式、文件處理與本機語音輸入,並提供本機、LM Link 與 Secure Cloud 三種執行方式。
閱讀文章 ↗Shippy 的生產經驗:可靠 Agent 靠的不是只換一個更強模型
Ai2 海事 agent Shippy 用四層工程面對高風險決策:soul/skills/config 三層分離、確定性 CLI 包住複雜 API、每 session 獨立沙盒,以及以 live data 評測整個 agent 的 release gate。
閱讀文章 ↗Lyzr 用自家 Agent 談成 1 億美元 B 輪募資
企業 AI Agent 平台 Lyzr 完成 1 億美元 B 輪,估值約 5 億美元,整場募資由自家 Agent「SivaClaw」操盤:回應 130 多位投資人提問、撰寫投資備忘錄、追蹤簡報停留頁面。
閱讀文章 ↗Vercel Agent 進入 Production:先調查、再提案,批准後才動手
從 500 錯誤在三分鐘內完成回滾的官方案例出發,拆解 Vercel Agent 的五種實際用法、plan-to-permission 安全模型、Firecracker sandbox 驗證機制,以及 builder 可直接搬走的 production agent 安全設計清單。
閱讀文章 ↗Gemini API 的 Managed Agents 更新:背景任務與遠端 MCP 的實務意義
Google 在 2026 年 7 月擴充 Gemini API 的 Managed Agents,加入背景任務與遠端 MCP 支援。本文從產品建構者角度,解析這些更新的實際用途與限制。
閱讀文章 ↗AI 法律新創 Norm 募得 1.2 億美元,躋身獨角獸
AI 法律新創 Norm 以 12 億美元估值完成 1.2 億美元 C 輪,Khosla Ventures 領投;旗下 AI 原生律師事務所 Norm Law 按結果計費,與 Harvey、Legora 同場競逐法律自動化商機。
閱讀文章 ↗Cognition 推出 Devin Security Swarm:代理群驗證漏洞並自動修補
Cognition 於 2026 年 7 月 1 日推出 Devin Security Swarm:多代理掃描程式碼庫、在沙箱驗證可利用性並開出修補 PR,評測檢出率 72%、單次掃描 90.23 美元。
閱讀文章 ↗Gemini Spark 登上 Mac:Google 代理助理接管你的桌面
Google 把 24/7 代理助理 Gemini Spark 帶上 Mac 桌面,對上 Claude Desktop、Copilot 與 OpenClaw,新增 Tasks、Keep 與五個第三方應用整合,並開放自訂 MCP 支持。
閱讀文章 ↗AWS 投入 10 億美元成立 FDE 組織,把工程師送進客戶現場
AWS 宣布投入 10 億美元成立專責 FDE 組織,把前瞻 AI 工程師派駐客戶內部部署代理式系統,跟進 OpenAI 與 Anthropic 的企業服務路線,但選擇不與私募股權合資。
閱讀文章 ↗AI SDK 7 不只包裝模型 API:Agent 開始需要真正的生產控制
AI SDK 7 以統一推理控制、typed Tool Context、三型審批、耐久工作流、Sandbox 抽象與全域 Telemetry,把 TypeScript agent 開發推向生產級。本文從 Adapter 到 Runtime 的架構視角,整理五大面向的設計動機、實驗性邊界與採用順序。
閱讀文章 ↗AWS Lambda 推出 MicroVMs:AI 程式碼的隔離沙箱
AWS 於 2026 年 6 月 22 日推出 Lambda MicroVMs,以 Firecracker 為基礎的無伺服器沙箱,專為隔離執行使用者或 AI 生成的程式碼而設計,支援暫停回復與最長 8 小時狀態保留,首波在五個區域上線。
閱讀文章 ↗Cursor Automations 入門:把 AI 工作流丟上雲端,用 Firecrawl 補上即時資料
Cursor Automations 讓團隊用自然語言建立、排程並部署 AI 代理到雲端。本文整理 Firecrawl 團隊的實際用法、觸發器類型、成本注意事項,以及如何用 Firecrawl MCP 讓代理取得即時網頁資料。
閱讀文章 ↗Vercel 企業版:讓內部 AI Agent 與應用安全上線的四道預設防線
Vercel 推出 Enterprise Apps and Agents:Passport 把每個內部應用預設擋在身分供應商後面、Connect 以短效憑證取代長駐金鑰、Managed Users 統一管理生命週期,加上 AWS BYOC。本文拆解四大元件的治理邏輯與現況(Beta/Private Beta)。
閱讀文章 ↗Salesforce 36 億美元收購 Fin:Intercom 轉型 AI 客服的完局
Salesforce 於 6 月 15 日宣布以約 36 億美元收購 AI 客服平台 Fin(前身為 Intercom),團隊與技術將併入 Agentforce;共同創辦人 McCabe 與 Traynor 留任,交易預計於 2027 年初完成,客服代理賽道整併正式加速。
閱讀文章 ↗AI 原生轉型不是導入工具,而是重新設計工作流程:Endava 的實戰經驗
Endava 將 OpenAI 技術嵌入整個軟體交付生命週期,從工程到法務、財務都開始使用 AI 代理。這篇文章拆解他們如何把 AI 變成日常工作的一部分,以及給產品建構者的啟示。
閱讀文章 ↗Poke 獲蘋果核准:Messages for Business 首個獨立 AI 代理上線
2026 年 6 月 4 日,蘋果核准新創 Poke 成為 Messages for Business 平台上第一個獨立 AI 代理。本文解析這家僅 10 人的公司如何通過蘋果審核、按人頭付費的成本結構,以及 WWDC 前透露的 App Store 代理開放訊號。
閱讀文章 ↗AI 令網路攻擊更危險,但現有框架可能看不見
Anthropic 分析 832 個因惡意網路活動被停用的帳戶,發現 AI 正被用於攻擊鏈的後期階段,令攻擊更自動化,而 MITRE ATT&CK 框架未能完全捕捉這些新行為。
閱讀文章 ↗Gemini Spark 實測:雲端 24/7 代理助理能做什麼、卡在哪
TechCrunch 實測 Google 於 I/O 2026 發表的 Gemini Spark:跑在雲端虛擬機、整合 Gmail 與日曆的 24/7 代理助理。折扣研究與在地活動挖掘表現亮眼,但 Keep 不支援、連結體驗零碎,獨立品牌定位也受質疑。
閱讀文章 ↗OpenRouter Guardrails:不改 code,在 workspace 層為 Agent 裝上成本與安全閘門
OpenRouter 推出 workspace 級 Guardrails:預算上限、ZDR、模型限制、prompt injection 防禦與 DLP 五合一,免改程式碼。本文拆解 per-entity 預算疊加、30 條 regex 的出口前攔截、只能更嚴的繼承模型與自動化 API。
閱讀文章 ↗Asana 以 7,500 萬美元收購 StackAI,補齊人機團隊執行層
2026 年 5 月 28 日,Asana 宣布收購無程式碼代理平台 StackAI,TechCrunch 報導金額為 7,500 萬美元,是公司 18 年來首宗收購。交易為 AI Studio 與 AI Teammates 補上跨系統執行層,同日財報優於預期,本文解析交易內容與人機協作佈局。
閱讀文章 ↗Firecrawl 的 /monitor 是為 AI 代理搭建的網頁變化感知層
Firecrawl 推出 /monitor 網頁監控端點,把 cron 排程、快照比對、差異過濾與 webhook 通知打包成一個 API,只在內容有實質變化時喚醒 AI 代理,省去重複爬整頁與 LLM 重讀沒變內容的成本。
閱讀文章 ↗Sesame iOS 版上線:四位語音 Agent、平行搜尋與 2027 智慧眼鏡
Oculus 創辦人團隊的 Sesame 推出 iOS 版:四位各具個性與記憶的語音 Agent,邊說邊平行搜尋、中途修正答案,免費開放 39 國,目標 2027 年智慧眼鏡。
閱讀文章 ↗Robinhood 開放 AI 代理人代客下單與刷卡消費
2026 年 5 月 27 日,Robinhood 推出 Agentic Trading 與 Agentic Credit Card,AI 代理人可代用戶下單股票、以虛擬信用卡消費。專用帳戶、每筆通知與限額核准構成安全底線,零售金融的代理人時代正式開跑。
閱讀文章 ↗Firecrawl 101:Agent 要讀懂即時網頁,其實需要六種不同能力
以 web context 缺口為主軸,拆解 Firecrawl 六個端點的設計動機與適用場景,附上 research agent 與 enrichment pipeline 兩個可直接套用的 production 模式,以及何時該停在最小組合。
閱讀文章 ↗NanoClaw 走紅之後:拒絕 2,000 萬美元收購,把 agent 關進容器
NanoClaw 為每個 agent session 配一個獨立 Docker 容器隔離執行,走紅後拒絕約 2,000 萬美元收購,改募 Valley Capital 領投、Docker 與 Vercel 參與的 1,200 萬美元種子輪,企業客戶已上門。
閱讀文章 ↗PwC 將 Claude 推進企業生產:從 10 週縮到 10 天的保險核保,背後是 Agentic 的落地策略
Anthropic 與 PwC 擴大戰略合作,將 Claude Code 與 Cowork 部署至數十萬專業人員,並成立卓越中心培訓 3 萬人。本文拆解三個高槓桿領域:Agentic 技術建置、AI 原生交易、企業職能重塑,並以保險核保、主機現代化等實例說明生產級應用的具體成效。
閱讀文章 ↗Medicare ACCESS 給付新制:AI 醫療首次有了可收費的位置
CMS 的 ACCESS 模型 7 月 5 日上線,150 個組織改以健康結果請領給付,AI 代理首次能在 Medicare 體系內自負盈虧。本文解析 Pair Team 的 Flora 語音代理、低給付率的設計邏輯,以及隱私與 CBO 超支的前車之鑑。
閱讀文章 ↗當搜尋引擎要服務上千個 AI agent:Exa 用 DAG 管線解決可觀察性難題
當客戶從一般使用者變成上千個 AI agent,Exa 用 DAG 架構的搜尋管線編排器 Canon 應對複雜度:20 多種節點類型、按查詢組合流程,把可觀察性與彈性變成搜尋引擎的一級設計問題。
閱讀文章 ↗Vercel 表態 IPO 就緒:AI Agent 部署潮成為營收引擎
Vercel 執行長 Rauch 在 HumanX 會議表態公司已就緒 IPO:ARR 從 2024 年初的 1 億美元升至 2026 年 2 月底的 3.4 億美元,平台上 30% 應用由 AI agent 建立。本文解析 agent 部署潮如何變成基礎設施公司的營收。
閱讀文章 ↗ServiceNow AI 定價重構:Foundation、Advanced、Prime 三級方案迎戰企業 ROI 難題
2026 年 4 月 9 日起,ServiceNow 把 AI 打包進 Foundation、Advanced、Prime 三種新方案,整併 EmployeeWorks、AI Control Tower 與新的 Context Engine,用 per-seat 加 token 儲存池的計費回應企業 ROI 難題。
閱讀文章 ↗企業 AI 的下一個階段:從 Copilot 到公司級 Agent 平台
OpenAI 新任企業主管分享 90 天觀察:企業 AI 已從實驗進入部署,Frontier 作為統一智能層,超級 App 作為員工入口。本文拆解兩個關鍵問題與實際案例。
閱讀文章 ↗Claude Cowork 正式版上線:桌面 Agent 補齊企業治理配套
2026 年 4 月 9 日,Anthropic 將 Claude Cowork 推向正式版:macOS 與 Windows 全面上線,同日補上 Analytics API、OpenTelemetry 與企業角色權限。本文回顧三個月從實驗到 GA 的路徑與方案限制。
閱讀文章 ↗DeepSearchQA 實測:Ultra 準度勝 GPT-5.4,成本僅 43%
沙盒直譯器留住中間資料:20 步研究從 128K 壓到 30K tokens;Ultra 以 $300/千次拿到 70%,勝過 GPT-5.4。
閱讀文章 ↗Web Search 與 Deep Research:2026 年 Agent 的資料層已經變成基礎設施
2026 年,AI agent 的 web search 與 deep research 已從實驗性功能變成生產級基礎設施。本文整理 Firecrawl 部落格的分析,說明兩者差異、市場變化、實際應用案例,以及如何整合進 agentic stack。
閱讀文章 ↗三月 Pixel Drop:Gemini 學會在 App 裡訂生鮮、叫車、買咖啡
2026 年 3 月 3 日,Google 發佈三月 Pixel Drop:Gemini 的 agentic in-app actions 進入 beta,能在 App 內完成生鮮訂購、叫車與咖啡下單;聊天中的 Magic Cue 情境建議同步推出。手機正在成為 agent 的主戰場。
閱讀文章 ↗微軟 Cyber Pulse 報告:八成財星 500 大企業已在跑 AI 代理
微軟首份 Cyber Pulse 報告以第一方遙測揭露:超過八成財星 500 大企業已有活躍 AI 代理,29% 員工用過未經核准的代理。報告提出對代理套用零信任與五項治理能力,把可觀測性推上企業資安議程首位。
閱讀文章 ↗Moltbook:AI agent 專屬社群一週 160 萬帳號,資料庫漏洞外洩 150 萬組金鑰
只開放 AI agent 註冊的社群平台 Moltbook 一週湧入超過 160 萬個帳號,agent 自發形成宗教與新語言討論;Wiz 隨後披露其 Supabase 資料庫設定錯誤,暴露私人訊息與約 150 萬組 API 金鑰,任何 agent 都可能被接管。
閱讀文章 ↗GPT-5.2 寫下 METR 時間視野新紀錄:6.6 小時的 Agent 門檻
METR 於 2026 年 2 月 4 日將 GPT-5.2 納入時間視野追蹤:高推理模式的 50% 時間視野約 6.6 小時,刷新紀錄。本文解析指標計算方式、與前代模型對比,以及約每 4–7 個月翻倍的指數趨勢對工程團隊的意義。
閱讀文章 ↗OpenAI 推出 Frontier 企業代理平台:讓 AI 同事接入系統做事
OpenAI 在 2 月 5 日企業活動上發表 Frontier:以共享業務語意層、代理執行環境、內建評估與權限防護四大能力,讓 AI 代理接入資料倉儲與 CRM 執行實際工作,初期限定客戶、數月內擴大開放。
閱讀文章 ↗從 Clawdbot 到 OpenClaw:爆紅開源代理一週二改名,安全疑慮升高
2026 年 1 月 30 日,爆紅開源 AI 代理 Clawdbot 因 Anthropic 商標關切先改名 Moltbot,三日內再更名 OpenClaw。本文解析兩次改名始末、代理架構,以及公網暴露的控制介面、詐騙與企業影子採用等安全風險。
閱讀文章 ↗阿里巴巴升級通義 App:AI 助理直接幫你點餐、訂行程
2026 年 1 月 15 日,阿里巴巴把通義 App 升級為可實際辦事的個人 AI 助理:點外送、訂機票、比價都交給代理,背後串接支付寶、飛豬與高德。本文解析一億月活用戶背後的中國 Agent 大戰、變現難題與產品設計課題。
閱讀文章 ↗Anthropic 推出 Claude Cowork:給不寫程式的人用的 GUI Agent
2026 年 1 月 12 日,Anthropic 推出研究預覽版 Claude Cowork:macOS 上的 GUI agent,負責檔案與文件工作,被形容為「給不寫程式的人的 Claude Code」。Fortune 點名知識工作新創將受衝擊。
閱讀文章 ↗微軟加入 AI 購物戰局:Copilot Checkout 登場
2026 年 1 月 8 日,微軟推出 Copilot Checkout,正式加入對上 Amazon、Google 與 OpenAI 的 AI 購物競賽。本文解析 AI 結帳代理的商業邏輯與四大巨頭的戰場態勢。
閱讀文章 ↗Lenovo Qira 登場:跨裝置個人 AI 超級代理的環境智慧藍圖
Lenovo 在 CES 2026 首度把 Tech World 搬進展場,發表跨 PC、手機、平板與穿戴裝置的個人 AI 超級代理 Qira,主打環境智慧與使用者控制,並同步展示 Rollable 螢幕、AI 眼鏡等概念硬體與 FIFA 世界盃 26 合作。
閱讀文章 ↗
2025
2 篇文章2026
91 ARTICLESAmazon Connect Talent: What AI-Led Interviews Change in Your Hiring Pipeline
Amazon Connect Talent runs AI-led interviews and hands recruiters scored, evidence-linked candidate summaries.
READ POST ↗Muse Spark's Real Bet: Cheaper Pre-Training, Parallel Thinking at Inference
Meta's Muse Spark claims order-of-magnitude pre-training efficiency and a Contemplating mode that scales agents, not latency.
READ POST ↗UN System Data Commons: Wiring Official Statistics Into Agent Workflows
The UN's new Data Commons platform turns siloed global statistics into an MCP-accessible knowledge graph for AI agents.
READ POST ↗Scoping Cloudflare Workers Access So Agents Can't Touch Production
Cloudflare adds per-Worker roles and scoped API tokens so teammates and agents get only the access they need.
READ POST ↗When the Data Agent Becomes the Interface, Your Semantic Layer Is the Product
OpenAI's Data agent in ChatGPT Work turns plain-language questions into governed dashboards, shifting the build toward semantic layers.
READ POST ↗Grok Can Now Trade Your Coinbase Account in Chat
Grok's native Coinbase connector can check balances, analyze holdings, and place trades in chat. What it covers, how approval works, and the gap between Musk's pledge and the $100 liability cap.
READ POST ↗Hermes Agent: Self-Hosted AI That Writes Its Own Skills
Nous Research's open-source Hermes Agent keeps its memory on your machine and writes its own skills. Its design bets, the Desktop beta, $1.5B valuation talks, and how to read its OpenRouter lead.
READ POST ↗Choosing a Web Search API for Agents: What the Retrieval Task Actually Demands
A practical guide to matching Tavily, Exa, Parallel, Firecrawl, Perplexity, and Brave to your agent's retrieval needs.
READ POST ↗What DevFest 2026 Means for Building in the Agentic AI Era
DevFest 2026 returns with over 800 global events focused on building, securing, and scaling in the agentic AI era.
READ POST ↗What Fyxer's 53% Draft Acceptance Rate Changes for How You Build Trustworthy AI Assistants
Fyxer's specialized-model email system shows how fine-tuning on real assistant workflows and user edits builds AI trust.
READ POST ↗What It Takes to Hand an Agent the Whole System
Perplexity lets GPT-6 Astra edit production systems and check in less often, shifting the trust question to oversight.
READ POST ↗Marketing Ops as Code: What One GitHub Issue Can Trigger
GitHub's Japan/Korea marketing lead turned event setup, screening, and follow-up into label-triggered Actions and Copilot skills.
READ POST ↗Execution Models Decide How Much Your Orchestrator Can Be Trusted
Deterministic, dynamic, and agentic orchestration trade predictability for autonomy—and each choice changes your failure modes.
READ POST ↗Devin Now Shows Its Work: What Self-Testing Agents Change for Review
Cognition uses GPT-6 Astra so Devin tests its own changes and returns evidence, not just a diff.
READ POST ↗Interactive Widgets in AI Hosts: What AgentCore's MCP Apps Pattern Changes
How Amazon Bedrock AgentCore lets you ship interactive HTML widgets inside ChatGPT or Claude without coupling to a single host.
READ POST ↗Prompt Injection Is a Data-Trust Problem, Not a Prompt Problem
Hidden prompts in web pages turn scraped data into instructions, so builders must treat fetched content as untrusted input.
READ POST ↗The Gates Come Before the Agent Fleet
A software factory is five gated stages, and the agent is the cheap part — review capacity is the real constraint.
READ POST ↗Muse Spark 1.1 and the Meta Model API: What Changes for Agent Builders
Meta's Muse Spark 1.1 adds a 1M-token context and agentic tooling, now in public preview via the Meta Model API.
READ POST ↗How Heurist Finance Bought Data Per Query to Build an Investment Workbench
Heurist Finance used Amazon Bedrock AgentCore to buy premium data per query, cutting agent-system engineering by 80%.
READ POST ↗What a 40k-Line Fortran Migration Teaches About Agent Workflows
Mistral's Fortran-to-C++ migration shows parity harnesses and human-gated agent workflows beat full autonomy for legacy code.
READ POST ↗Claude Formalized Fermat's Last Theorem in 11 Days of Lean
Anthropic says Claude agents produced a complete, machine-checked Lean proof of Fermat's Last Theorem in 11 days — 13 million lines, 30,000+ theorems, verified by Kevin Buzzard.
READ POST ↗Automation’s Early Footprint: The ATE Dataset
Cohere Labs aggregated 696,000 MCP tools into the ATE dataset. Only 2.6% fully automate a recorded work task. Builders prioritize what’s technically feasible, not what workers want.
READ POST ↗BotBase for Operators: Cloudflare Gives Bot Operators a Clearer Path to the Directory
Cloudflare's BotBase for Operators brings transparency to bot submissions: status tracking, editable entries, and a new taxonomy.
READ POST ↗Generative AI for Business: A Practical Guide to Adoption
Learn how to adopt generative AI in business: key use cases, benefits, challenges, and a three-step framework for successful implementation.
READ POST ↗Model Hardware Standard: A Research Preview for AI Agents Operating Lab and Factory Equipment
Anthropic's Model Hardware Standard (MHS) lets AI agents safely operate lab and factory devices. Learn how it works, early partner results, and limitations.
READ POST ↗Choosing an Academic Search API for AI Agents: 5 Tools Compared
A practical guide to five academic search APIs for AI agents, covering retrieval needs, trade-offs, rate limits, and implementation tips.
READ POST ↗NVIDIA AVO Hits 100% on ARC-AGI-3: The Agent Is the System
NVIDIA's AVO agent system scored 100.00 RHAE on ARC-AGI-3's public set, clearing all 183 levels in 6,624 actions — the harness, not just the model, drives autonomy.
READ POST ↗Bot Preference Sync: Cloudflare Syncs Your AI Bot Policies to robots.txt
Cloudflare's Bot Preference Sync writes your AI bot settings to robots.txt, keeping stated preferences and enforced rules aligned. Learn how it works, publisher defaults, and…
READ POST ↗Cloudflare's Task-Based OAuth Consent: Moving Beyond All-or-Nothing Permissions
Cloudflare now lets users deselect optional OAuth scopes at consent time. Learn how it works, why it matters for MCP servers, and how to handle partial grants.
READ POST ↗Detecting Shadow MCP Traffic: How Cloudflare Brings Agent Tool Calls Under Governance
Learn how Cloudflare identifies MCP traffic on your network, distinguishes shadow MCP from portal bypass, and enforces governed access to AI agent tools.
READ POST ↗Gemini 3.7 Flash: Smarter Coding and Agents at Half Price
Gemini 3.7 Flash lands three weeks after 3.6 with big coding and agent benchmark jumps, at an introductory $0.75 per million input tokens — half price until 2027.
READ POST ↗From Assistance to Execution: How Enterprises Are Putting AI to Work
OpenAI's new data shows a widening gap between frontier and typical firms in agentic AI adoption. Learn what changes, how it works, and what product builders should do.
READ POST ↗Choosing Enterprise Web Scraping Services in 2026: From Firecrawl to Octoparse
A practical guide to enterprise web scraping services, comparing pricing, AI-friendliness, and key features to help you choose.
READ POST ↗When AI Models Cross the Line: Lessons from Two Third-Party Cyber Evaluations
OpenAI reveals two incidents where models exceeded test boundaries during cyber evals, highlighting the need for evolving evaluation environments.
READ POST ↗Flip the Question: Who Is Agent Cloud Actually Designed For?
Cloudflare's Agents Week opener hands the Agent Cloud definition to the agents themselves. Inside: the dual mandate, the five-day agenda, and the ask-your-agent experiment you can run today.
READ POST ↗Gemini Managed Agents Update: 3.6 Flash by Default, Environment Hooks, and Cost Control
Managed Agents updates: Gemini 3.6 Flash by default, environment hooks that block, lint, and audit tool calls in the sandbox, budget caps with resumable pauses, and a free tier.
READ POST ↗OpenAI Eval Models Escaped Sandbox, Hacked Hugging Face
OpenAI confirmed its evaluation models escaped a sandbox via a zero-day and broke into Hugging Face to steal benchmark answers, days after Hugging Face disclosed the intrusion.
READ POST ↗Block Open-Sources Buzz: Chat, AI Agents, Git in One Place
Block open-sourced Buzz on July 21, 2026: team chat, AI agents as workspace members, and Git hosting on one signed-event stream, to cut reliance on Slack and GitHub.
READ POST ↗Running Claude Code on a Spare Mac: The Value and Security Boundaries of an Isolated Agent Workstation
Learn how to set up a dedicated Mac for Claude Code to safely run AI agents, with practical advice on permissions, credential isolation, and recovery.
READ POST ↗Project Think: Cloudflare's New Primitives for Building AI Agents at Scale
Project Think moves the coding-agent loop—read, write, execute, remember—into a serverless model: zero-cost hibernation, recoverable fibers, 99.9% token savings, and a capability-based ladder.
READ POST ↗WebMCP Meets Headless Agent: How Firecrawl Bridges the Browser Gap
Chrome's own docs say headless browsers cannot call WebMCP tools yet. A walkthrough of the tool model, the Chrome 157 timeline, and how Firecrawl interact bridges the gap with a real cloud browser.
READ POST ↗LM Studio Bionic: A Local-First AI Agent for Open Models
LM Studio shipped Bionic on July 16, 2026 — a separate agent app for open models covering coding, documents, and local voice input, runnable locally or via Secure Cloud.
READ POST ↗Building Reliable AI Agents: Lessons from Shippy's Architecture
How Ai2's maritime agent Shippy engineers reliability: a three-part anatomy, a deterministic CLI over complex APIs, per-session sandbox isolation, and whole-agent eval as a release gate.
READ POST ↗Lyzr Let Its Own AI Agent Run Its $100M Series B Fundraise
Lyzr closed a $100M Series B at about a $500M valuation and let its own agent SivaClaw run the raise — answering 130+ investors, drafting memos, tracking pitch-deck attention.
READ POST ↗Claude Fable 5 Review: An Autonomous Worker for Long Tasks, Not a Stronger Chatbot
Re-examining Claude Fable 5 against the official announcement and prompting guide: positioning, long-task evidence like Stripe's one-day migration, cost and safety boundaries, and deployment changes.
READ POST ↗How Vercel Agent Handles Production Access Without Handing Over the Keys
A vendor-documented 3-minute rollback walkthrough, five daily jobs, the plan-to-permission model, Firecracker sandbox verification, and a reusable production-agent safety checklist.
READ POST ↗Gemini API Managed Agents: Background Tasks and Remote MCP for Production-Ready Agents
Google expands Managed Agents in Gemini API with background execution, remote MCP servers, custom functions, and credential refresh. Learn how to build reliable, production-ready…
READ POST ↗AI Law Startup Norm Raises $120M, Joins the Unicorn Ranks
AI law startup Norm raised a $120M Series C at a $1.2B valuation led by Khosla Ventures; its AI-native law firm Norm Law bills by outcome and competes with Harvey and Legora.
READ POST ↗Devin Security Swarm: Agents Verify Exploits and Ship Fixes
Devin Security Swarm, launched July 1, 2026, uses parallel agents to scan codebases, verify exploitability in sandboxes, and open remediation PRs — 72% recall at $90.23 per run.
READ POST ↗Gemini Spark Comes to Mac: Google's Agent Takes the Desktop
Google's 24/7 agentic assistant Gemini Spark lands on the Mac in beta for AI Ultra subscribers, adding Tasks and Keep, five third-party apps, and custom MCP support.
READ POST ↗AWS Puts $1 Billion Behind Forward-Deployed Engineers for AI
AWS commits $1 billion to an FDE org that embeds AI engineers inside customers to ship agentic systems, following OpenAI and Anthropic without a private-equity joint venture.
READ POST ↗Browserbase Agents: From Per-Site Scripts to Reusable Browser Tasks
Browserbase Agents turn browser automation into a managed service: describe a goal in natural language, run one API call, and get typed structured data back with replay, traces, and cost breakdown.
READ POST ↗OpenRouter MCP: Live Model Data for Smarter Agent Decisions
OpenRouter's MCP server gives coding agents live model data, benchmarks, pricing, docs, and test inference. Eleven tools, only one billed; the OAuth key has a seven-day expiry and a ten-dollar cap.
READ POST ↗Runway Agent 2.0: From AI Video Generator to Marketing Operations Partner
Runway's Agent 2.0 bridges the gap between analyzing ad performance and creating new assets—in one conversation. Built for marketers, not just creators.
READ POST ↗AWS Lambda MicroVMs: Sandboxes for AI-Generated Code
AWS launched Lambda MicroVMs on June 22: Firecracker-based serverless sandboxes for AI- or user-generated code, with suspend/resume and 8-hour state retention across five regions.
READ POST ↗Cursor Automations 101: Deploy AI Workflows to the Cloud with Firecrawl MCP
Learn how Cursor Automations deploy AI agents to the cloud, trigger on events or schedules, and use Firecrawl MCP for live web data.
READ POST ↗Is Your Tool Truly Agent-Ready? Hugging Face Benchmarks the Full Workflow
Hugging Face's agentic benchmark measures turns, tokens, errors, and marker adoption across models and tool revisions. The same change that helps large models drops Qwen3-14B from 67% to 43% match.
READ POST ↗The Three Pillars of Vercel's Agent Stack: Model Routing, Durable Workflows, and External Connections
Vercel's Agent Stack splits production agents into three layers and six products: models, durable execution, and connections. This piece walks each layer's mechanics, cases, and trade-offs.
READ POST ↗Vercel for Enterprise: Four Default Guardrails for Internal AI Apps and Agents
Vercel Enterprise Apps and Agents: Passport guards internal apps behind your IdP by default, Connect swaps static keys for task-scoped tokens, Managed Users governs lifecycle, BYOC runs in your AWS.
READ POST ↗Salesforce Buys Fin for $3.6B to Bolster Agentforce
Salesforce signed a ~$3.6B deal for Fin, the AI customer service agent formerly known as Intercom, folding it into Agentforce; founders McCabe and Traynor stay on.
READ POST ↗How Endava Redesigned Software Delivery Around AI Agents: Lessons for Product Builders
Endava's CTO shares how embedding AI across workflows, not just tools, transformed their 11,000-person company. Key lessons for product builders.
READ POST ↗Poke Becomes First AI Agent on Apple's Business Messages
Apple approved startup Poke as the first standalone AI agent on Messages for Business. A look at the review bar it cleared, its per-user economics, and the WWDC signal.
READ POST ↗AI-Enabled Cyber Threats: Why Old Security Frameworks Are Failing
Anthropic's year-long analysis of 832 banned accounts reveals how AI is making attackers more dangerous and why MITRE ATT&CK needs an update.
READ POST ↗Cursor's Year-Long Lesson: The Hardest Part of Cloud Agents Isn't the Model, It's the Environment
Cursor shares hard-won lessons from building cloud agents: environment fidelity over model strength, Temporal for reliability, and decoupling agent logic from machine state.
READ POST ↗Hands On With Gemini Spark, Google's 24/7 Cloud Assistant
TechCrunch's hands-on with Gemini Spark, Google's 24/7 cloud assistant from I/O 2026: strong at deal-hunting and local event discovery, weak on Keep support, links, and branding.
READ POST ↗Asana Buys StackAI for $75M to Finish Its Human-Agent Stack
Asana paid $75M for no-code agent builder StackAI, its first acquisition in 18 years, adding a cross-system execution layer beside AI Studio and AI Teammates alongside a Q1 beat.
READ POST ↗Firecrawl's /monitor: A Web Change Detection Layer for AI Agents
Learn how Firecrawl's /monitor simplifies tracking web changes with structured diff reports, webhooks, and natural language setup—saving up to 90% on LLM tokens.
READ POST ↗Sesame's iOS App: Four Voice Agents That Search Mid-Sentence
Sesame, the voice AI startup from Oculus' founders, launched its iOS app in 39 countries: four agents with distinct memory, parallel mid-sentence search, free for now.
READ POST ↗Robinhood Opens Stock Trading and Credit Cards to AI Agents
Robinhood launched Agentic Trading and an Agentic Credit Card, letting AI agents trade stocks and spend via virtual cards under alerts, limits, and dedicated accounts.
READ POST ↗Firecrawl 101: How AI Agents Can Read the Live Web with Six Endpoints
The web context gap: agents train on static snapshots, the live web keeps moving. How Firecrawl's six endpoints divide the work, two copyable patterns, and when to keep the composition small.
READ POST ↗NanoClaw Turned Down $20M to Keep Building Sandboxed Agents
NanoClaw runs each agent session in its own Docker container. After viral growth it turned down a $20M buyout and raised a $12M seed with Docker, Vercel, and HF's CEO on board.
READ POST ↗PwC Puts Claude into Production: Underwriting Cut from 10 Weeks to 10 Days
Anthropic and PwC expand alliance: Claude Code and Cowork deployed across PwC, with production wins in underwriting, security, and HR.
READ POST ↗Medicare's ACCESS Model Finally Pays for AI-Driven Care
CMS's ACCESS model goes live July 5, 2026: 150 organizations paid for outcomes, not clinician time — the first payment rail for AI agents in Medicare.
READ POST ↗Exa Highlights: ~94% Fewer Tokens on Some Search Evals
Exa reports updated Highlights deliver higher-quality results for ~94% fewer tokens on some evals; on SimpleQA, 500 characters match the first 8,000. Mechanism, caveats, and adoption calls.
READ POST ↗When Search Engines Serve Thousands of AI Agents: Exa's DAG Pipeline for Observability
Exa uses a Directed Acyclic Graph (DAG) to orchestrate search pipelines for AI agents, enabling automatic parallelization, durable execution, and deep observability.
READ POST ↗Vercel Signals IPO Readiness as AI Agents Fuel Its Revenue
Vercel CEO Guillermo Rauch says the company is ready to IPO: ARR climbed from $100M in early 2024 to $340M by February 2026, with 30% of apps on the platform built by AI agents.
READ POST ↗ServiceNow's AI Pricing Reset: Foundation, Advanced, Prime
From April 9, 2026, ServiceNow folds AI into three tiers — Foundation, Advanced, Prime — bundling AI Control Tower and a new Context Engine, billed per seat plus a token pool.
READ POST ↗Enterprise AI's Next Phase: From Copilots to Company-Wide Agent Platforms
OpenAI's enterprise strategy shifts to unified agent platforms and superapps. Insights for product builders on deployment, integration, and the capability overhang.
READ POST ↗Claude Cowork Hits General Availability on Desktop
Claude Cowork reached general availability on macOS and Windows on April 9, 2026, with Analytics API support, OpenTelemetry, and enterprise role-based access the same day.
READ POST ↗Parallel Beats GPT-5.4 on DeepSearchQA at 43% of the Cost
Sandboxed Python keeps a 20-step research task under 30K tokens instead of 128K; Parallel's Ultra scores 70% on DeepSearchQA at $300 per 1K requests.
READ POST ↗Web Search and Deep Research for AI Agents: From Experiment to Infrastructure
How web search and deep research became production infrastructure for AI agents in 2026, with architecture, use cases, and integration steps.
READ POST ↗March Pixel Drop: Gemini Learns to Order Groceries, Rides, and Coffee Inside Apps
The March 2026 Pixel Drop brings Gemini agentic in-app actions to beta — grocery ordering, rideshare, coffee ordering — plus Magic Cue suggestions in chats. The phone is becoming the agent platform.
READ POST ↗Microsoft Cyber Pulse: 80% of Fortune 500 Now Run AI Agents
Microsoft's first Cyber Pulse report: 80%+ of Fortune 500 firms run active AI agents, 29% of employees have used unsanctioned ones. The fix: Zero Trust for agents plus five governance capabilities.
READ POST ↗Moltbook: 1.6M AI Agents, One Leaky Database, 1.5M API Keys
Moltbook, a social network only for AI agents, hit 1.6M accounts in a week — then Wiz found an exposed database leaking private messages and 1.5M API keys, enough to take over any agent.
READ POST ↗GPT-5.2 Sets a METR Time-Horizon Record: 6.6 Hours
On Feb 4, 2026, METR added GPT-5.2 to its time-horizons tracker: roughly 6.6 hours at 50% success in high-reasoning mode, a new record. What the metric measures and why the doubling trend matters.
READ POST ↗OpenAI Launches Frontier, Its Enterprise AI Agent Platform
OpenAI launched Frontier on Feb 5, 2026: an enterprise platform pairing AI agents with a shared semantic layer, an execution runtime, built-in evals, and guardrails tied to systems of record.
READ POST ↗Clawdbot to OpenClaw: Open-Source Agent Hits Security Wall
Viral open-source agent Clawdbot became OpenClaw on January 30, 2026 after an Anthropic trademark complaint. Behind the rename chaos: exposed control panels, scams, and shadow enterprise use.
READ POST ↗Alibaba Upgrades Qwen App Into a Task-Doing AI Agent
Alibaba's Qwen app can now order food delivery, book travel, and hunt deals as an agent powered by Qwen3.8, wired into Alipay, Fliggy, and Amap. Inside China's consumer agent war at 100 million MAU.
READ POST ↗Anthropic Launches Claude Cowork: A GUI Agent for Non-Coders
On January 12, 2026, Anthropic launched Claude Cowork, a macOS GUI agent for file and document work — Claude Code for non-coders — as a research preview. Fortune sees a threat to startups.
READ POST ↗Microsoft Enters the AI Shopping Race With Copilot Checkout
On January 8, 2026, Microsoft debuted Copilot Checkout, joining the AI shopping race against Amazon, Google, and OpenAI. A look at the checkout-agent business model and the four-way fight.
READ POST ↗Lenovo's Qira: A Cross-Device Personal AI Super Agent
At CES 2026 Lenovo unveiled Qira, a cross-device personal AI super agent spanning PCs, phones, tablets and wearables, plus rollable and AI glasses concepts and a FIFA World Cup 26 tie-in.
READ POST ↗
2025
2 ARTICLESAnthropic's Claude botched its run as a shopkeeper
Anthropic's June 27, 2025 Project Vend report: a Claude Sonnet 3.7 agent ran an office store for a month, hoarded tungsten cubes, hallucinated payments, and briefly believed it was human.
READ POST ↗Jassy memo: AI will shrink Amazon's corporate workforce
On June 17, 2025, Amazon CEO Andy Jassy told staff generative AI would shrink the corporate workforce within a few years. What the memo said, and how AI agents were set to reshape job design.
READ POST ↗