2026
4 篇文章MosaicLeaks 基準:研究代理的對外查詢正在洩漏企業機密
ServiceNow 團隊發布 MosaicLeaks 基準:深度研究代理混合私有文件與網路搜尋時,攻擊者只看外流查詢紀錄就能拼出企業機密,而 PA-DR 訓練法把洩漏率從 34% 壓到 9.9%。
閱讀文章 ↗搜尋工具會改變 RL 結果:Exa 的實驗提醒我們別只盯著模型
Exa 用兩個只有搜尋後端不同的 RL agent 做對照:Exa 訓練的 4B 在六個 benchmark 全數領先 SERP 版,還常超過未訓練的 235B,並以 0.58B tokens 追平 1.89B 的表現。本文拆解更密 reward signal 的機制、47 萬次搜尋呼叫的穩定性數據,以及這份廠商研究的限制。
閱讀文章 ↗OpenAI 的「哥布林」之謎:獎勵機制如何悄悄塑造模型行為
OpenAI 公開調查 GPT-5.1 以來模型頻繁提及「哥布林」等生物的現象,發現源於「Nerdy」個性訓練的獎勵訊號,並因回饋迴圈擴散。本文拆解根因、傳播機制與教訓,對產品開發者與 AI 學習者深具啟發。
閱讀文章 ↗David Silver 創辦 Ineffable Intelligence:11 億美元豪賭「超級學習者」
2026 年 4 月,DeepMind 強化學習主管 David Silver 創辦的 Ineffable Intelligence 以 51 億美元估值募得 11 億美元,目標是打造純靠強化學習試錯、不依賴人類資料的「超級學習者」。本文解析資金、技術賭注與倫敦 AI 樞紐的崛起。
閱讀文章 ↗
2026
4 ARTICLESMosaicLeaks: Research Agents Leak Secrets Through Queries
ServiceNow's MosaicLeaks benchmark shows deep research agents leak enterprise secrets via outbound search queries; its PA-DR training cuts leakage from 34% to 9.9%.
READ POST ↗Search Quality Shapes RL Outcomes: Why Your Agent's Backend Matters More Than You Think
Exa's controlled experiment shows that changing the search backend during RL training significantly impacts agent performance and efficiency.
READ POST ↗Where the Goblins Came From: How Reward Signals Quietly Shape Model Behavior
OpenAI traced a 175% spike in 'goblin' mentions to a reward signal for the Nerdy personality. Learn how RL feedback loops spread quirks and what product builders can do.
READ POST ↗Ineffable Intelligence: $1.1B to Learn Without Human Data
David Silver's lab Ineffable Intelligence raised $1.1B at a $5.1B valuation to build a 'superlearner' trained by reinforcement learning alone — no human data. Inside the round.
READ POST ↗