2026
77 篇文章把模型快取放進節點:HyperPod 推論冷啟動的實務取捨
Amazon SageMaker HyperPod 推出模型快取,把權重與容器映像預載到節點 NVMe,讓擴容從數十分鐘縮到數秒。
閱讀文章 ↗Megakernel 不是炫技:把 decode 從「等 kernel」改成「等資料」的實作取捨
Cohere 用單一 CUDA 檔把 North Mini Code 的 decode 做成 persistent megakernel,batch size 1 吞吐從 vLLM 的 185 tok/s 拉到 292 tok/s。
閱讀文章 ↗前綴感知路由:讓 KV cache 不再被隨機打散
SageMaker Inference 新增前綴感知路由,把相同 prompt 開頭固定送到同一台機器,讓 prefix caching 真正累積成可重用的 KV cache。
閱讀文章 ↗Claude 十一天形式化費馬最後定理:1,300 萬行 Lean 完整證明
2026 年 9 月,Anthropic 宣布 Claude 代理群在 11 天內完成費馬最後定理的完整 Lean 形式化:1,300 多萬行程式碼、逾 3 萬條定理,由數學家 Kevin Buzzard 親自驗證。
閱讀文章 ↗Ox Alpha 就是 GLM-5.3-Flash:MIT 開源,OpenRouter 週流量佔 31%
智譜揭曉匿名冠軍 Ox Alpha 就是 GLM-5.3-Flash 並以 MIT 授權開源:預覽期在十萬顆國產晶片上處理 62 兆 tokens,OpenRouter 週流量佔 31%、登頂程式碼模型榜,消息公布當日港股大漲逾 12%。
閱讀文章 ↗GLM-5.3 權重開放下載:753B 旗艦的本地部署現實
2026 年 8 月 25 日,Z.ai 把 753B 參數的 GLM-5.3 權重放上 Hugging Face,三天後在 Hacker News 衝上 806 分。本文解析其 MoE 架構、自訂授權條款與本地部署實測。
閱讀文章 ↗Qwen3.8-Flash-Next 發表:新架構把啟用參數壓到 6B
2026 年 8 月 26 日,Qwen 開源 Qwen3.8-Flash-Next:125B 主模型加 51B n-gram 嵌入表,每個 token 只啟用 6B 參數,原生 262K 上下文,官方定位是極致成本效率。
閱讀文章 ↗AI Token 計費陷阱:同一段文字,不同模型收費差 2.65 倍
Token 不是長度單位,而是供應商 tokenizer 的產物。同一段文字在不同模型可能產生 2.65 倍 token 差異,而輸入格式(如 HTML vs Markdown)更可造成 21 倍差距。本文拆解 tokenization 如何影響成本,並提供實務建議。
閱讀文章 ↗NVIDIA AVO 於 ARC-AGI-3 滿分:模型之外的代理系統才是關鍵
NVIDIA 的 AVO 代理系統在互動推理基準 ARC-AGI-3 公開集拿到 100.00 RHAE、用 6,624 步解完 183 關,本文解析其監督者與持續記憶體設計,以及「模型不是整個代理」的意義。
閱讀文章 ↗Stripe 收購 OpenRouter:每天十兆 tokens 的模型路由層易主
Stripe 於 2026 年 8 月 19 日宣布收購 AI 模型路由層 OpenRouter(紐約時報報導約 75 億美元),這個每天路由逾 10 兆 tokens、串接 400 多個模型的中立平台將併入支付網路。
閱讀文章 ↗Qwen3.8-27B 開源釋出:體積小、看得懂圖,只是預設想太多
阿里巴巴 Qwen 兌現一週前承諾,釋出 Apache 2.0 的 Qwen3.8-27B 開放權重:262K 上下文、原生視覺、單機可跑;Simon Willison 實測稱讚能力,也點出 xhigh 推理預設讓簡單任務慢上數倍。
閱讀文章 ↗Grok 4.6 評測:61 分重返前沿,代理任務變主戰場,價格不變
SpaceXAI 於 8 月 12 日發布 Grok 4.6:Artificial Analysis 智能指數 61 分、追平 GPT-5.6 Sol,定價維持每百萬 token 輸入 2 美元、輸出 6 美元,主打長時間代理工作。
閱讀文章 ↗Z.ai 發表 GLM-5.3:開源程式碼新高,網路攻擊能力超預期
2026 年 8 月 14 日,Z.ai 發表 GLM-5.3:沿用 GLM-5.2 基座、靠後訓練把程式碼基準推上開源新高,並主動披露網路攻擊能力「發展得比我們預期更快」,授權同步轉為自訂條款。
閱讀文章 ↗Gemini 3.7 Flash 三週再迭代:寫程式與代理更強、入門價砍半
Google 於 8 月 13 日發布 Gemini 3.7 Flash:程式與代理能力大幅跳升,入門價每百萬輸入 tokens 0.75 美元、只有 3.6 Flash 原價一半,2027 年起回漲一倍。
閱讀文章 ↗偷走 AI 的思考:兩次 API 呼叫還原隱藏推理鏈
德國研究團隊利用同家族較弱模型與越獄提示,兩次 API 呼叫還原加密思考區塊的推理原文;公開代理軌跡中更發現大量 API 金鑰與密碼。
閱讀文章 ↗Meta 回歸開放權重:30B 的 Muse Glimmer 把本地代理變成現實
2026 年 8 月 10 日,Meta 以 Apache 2.0 釋出 30B 參數的 Muse Glimmer,針對常駐本地代理工作流設計,單張消費級 GPU 就能跑,並承諾開放 Muse Spark 權重。
閱讀文章 ↗Qwen3.8-Max 正式發布並開放權重:鎖定寫程式與代理協作
2026 年 8 月 3 日,阿里巴巴正式發布 Qwen3.8-Max:百萬 token 上下文、多模態輸入、輸入每百萬 token 2 美元,並成為 Max 系列首個開放權重的模型,主打寫程式與代理協作。
閱讀文章 ↗在 8GB Mac 上跑 26B 模型:TurboFieldfare 的 SSD 串流推理
開源專案 TurboFieldfare 以 Swift 與 Metal 打造推理引擎:只常駐約 2GB 記憶體,把 Gemma 4 26B-A4B 的 MoE 專家權重從 SSD 逐 token 串流載入,在 8GB MacBook Air 上跑出每秒 5 到 6 個 token。
閱讀文章 ↗DeepSeek V4-Flash 0731:只重做後訓練,代理能力越級跳
2026 年 7 月 31 日,DeepSeek 推出 V4-Flash-0731:架構與規模不變、僅重新後訓練,代理基準大幅超越 V4-Pro-Preview,輸出速度每秒 210 個 token,權重續以 MIT 授權開放。
閱讀文章 ↗兩個 API 設定讓 GPT-5.6 Sol 在 ARC-AGI-3 分數三倍跳:評估背後的隱藏變數
OpenAI 發現,保留推理與壓縮上下文這兩個 API 設定,能讓 GPT-5.6 Sol 在 ARC-AGI-3 基準測試的成績從 13.3% 提升到 38.3%,同時減少 6 倍輸出 token。這提醒我們,基準測試衡量的不只是模型能力,還包括測試框架的設計選擇。
閱讀文章 ↗Qwen 3.8 登場:阿里 2.4 兆參數旗艦先上訂閱,權重還沒開
2026 年 7 月 19 日,阿里巴巴發布 Qwen 3.8,2.4 兆參數旗艦以 Qwen3.8-Max-Preview 形態先在訂閱方案上線,開放權重尚未釋出,被視為對 Kimi K3 的回應。
閱讀文章 ↗Meta 自導式測試時訓練:讓 LLM 自己挑該學的長上下文證據
Meta AI 提出自導式測試時訓練 S-TTT,先讓模型自行挑選與問題相關的證據段落、再只對這些段落做適應訓練,在 LongBench-v2 與 LongBench-Pro 上帶來最高 15% 的相對準確率提升。
閱讀文章 ↗GLM-5.2 開源發布:百萬上下文長任務表現緊咬 Opus 4.8
Z.AI 於 2026 年 6 月 17 日開源 753B 參數的 GLM-5.2:MIT 授權、穩定 1M token 上下文,長任務與程式基準緊追 Claude Opus 4.8,成為開源陣營排名最高的模型。
閱讀文章 ↗Baseten 傳以 130 億美元估值募 15 億美元:五個月估值跳 160%
WSJ 報導 Baseten 接近完成 15 億美元輪次,估值最高 130 億美元,距 1 月以 50 億美元估值完成的 3 億美元 Series E 僅五個月。本文拆解雙軌定價設計、開源推理路由生意,與推理淘金熱背後的毛利風險。
閱讀文章 ↗Sarvam AI 估值 15 億美元:HCLTech 領投下的印度 AI 獨角獸
6 月 15 日,班加羅爾的 Sarvam AI 以 15 億美元估值募得 2.34 億美元,成為印度最新 AI 獨角獸。HCLTech 出資 1.5 億美元領投,B 輪目標 3 億美元。本文解析這筆主權 AI 交易的結構,與對印度模型生態的意義。
閱讀文章 ↗Pramaana Labs 募 2,700 萬美元:用形式化驗證約束 LLM 輸出
2026 年 6 月 17 日,Pramaana Labs 宣布獲 Khosla Ventures 領投 2,700 萬美元種子輪,以 LEAN 形式化驗證技術為 LLM 加上確定性驗證層,瞄準法律、稅務與藥物發現等出錯代價極高的領域。
閱讀文章 ↗Mistral 傳募 30 億歐元、估值 200 億:歐洲主權 AI 的加注與鴻溝
彭博 6 月 12 日報導 Mistral 洽談以約 200 億歐元估值募集 30 億歐元,較去年 9 月的 117 億估值近乎翻倍。本文拆解這家歐洲主權 AI 旗手的籌碼:巴黎資料中心、國防合作,以及與美國實驗室之間巨大的資金鴻溝。
閱讀文章 ↗Google 開源 DiffusionGemma:擴散式生成快 4 倍,單卡 H100 破千 token/秒
Google 於 6 月 10 日開源 DiffusionGemma:26B-A4B 擴散語言模型,Apache 2.0 釋出,單卡 H100 每秒生成逾千 token、比自回歸快 4 倍,但多數基準仍輸 Gemma 4,官方明言最適合本地與低併發場景。
閱讀文章 ↗TensorZero 收攤:730 萬美元種子輪的開源 LLM 閘道一夜封存
6 月 13 日,開源 Rust LLM 閘道 TensorZero 的 GitHub repo 無預警轉為封存。共同創辦人證實公司收攤:兩年半僅花用不到一半的 730 萬美元種子輪資金,剩餘資本退還投資人,程式碼以 Apache 2.0 保留但不再維護。
閱讀文章 ↗Theker 籌 8,500 萬美元:打造「什麼都不專精」的通用工廠機器人
巴塞隆納機器人新創 Theker 於 2026 年 6 月 11 日宣布 8,500 萬美元 Series A,CRV 領投,Samsung 與 LVMH 旗下的 Aglaé Ventures 參投,號稱歐洲機器人最大 A 輪。其模組化機器人鎖定混亂、非重複的工廠與倉儲流程。
閱讀文章 ↗Writer 新研究:記憶工具讓模型更愛附和、更不準確
Writer AI Research 發表兩篇論文並推出 MIST 基準:Mem0、Zep 等記憶系統會放大模型附和行為,Sonnet 4.6 在 MIST-Moral 上從 1.6% 升到 40.2%。問題出在記憶擷取層,改用 LLM 生成摘要可壓到 12.8%。
閱讀文章 ↗Liquid AI LFM2.5-8B-A1B 發布:38T tokens 訓練的端側 MoE 推理模型
Liquid AI 於 5 月 28 日發布 LFM2.5-8B-A1B:8B 總參數、每 token 約 1B 啟用的端側 MoE 推理模型,預訓練 38T tokens、128K 上下文,MATH500 達 88.76,手機上每秒約 30 tokens。
閱讀文章 ↗史丹佛法學院盲測:AI 答疑在 75% 對決中勝過教授
史丹佛法學院 6 月 1 日發布盲測研究:16 位法學教授對 40 題契約法答疑做了近 3,000 次匿名評比,AI 在 75% 對決中勝出,答案被標為有害的比例僅 3.5%,低於人類同儕的 12%。
閱讀文章 ↗Cisco 實測 15 個封閉前沿模型:多輪攻擊無一倖免
2026 年 5 月 27 日,Cisco 發表研究,對 OpenAI、Anthropic、Google、Amazon、xAI 共 15 個封閉模型發動近 7,000 次多輪攻擊,最高成功率達 88.3%,沒有任何模型免疫,連設定旗標都會大幅改變風險。本文解析測試方法、各模型數字與採購建議。
閱讀文章 ↗OpenRouter 完成 1.13 億美元 B 輪:模型路由層估值翻倍至 13 億美元
2026 年 5 月,OpenRouter 宣布 1.13 億美元 B 輪融資,由 CapitalG 領投,估值翻倍至約 13 億美元。每月處理 100 兆 token、支援 400+ 模型,本文解析模型路由層為何成為 Agent 時代的關鍵基礎設施。
閱讀文章 ↗中國傳要求頂尖 AI 人才出境先報准:阿里巴巴、DeepSeek 在列
Bloomberg 報導,中國將出境限制擴及民營企業頂尖 AI 人才,阿里巴巴與 DeepSeek 員工出國前須取得官方核准。本文解析從 2025 年「建議」走向制度化的演變與對產業的影響。
閱讀文章 ↗DeepSeek 把 V4 Pro 七五折降價常態化:前沿模型價格戰的底牌
2026 年 5 月 22 日,DeepSeek 宣布旗艦 V4 Pro 的 75% 折扣轉為常態價:輸出每百萬 token 0.87 美元,僅為原價四分之一,Flash 更只要 0.28 美元。本文解析降價常態化的成本邏輯與對 API 市場的連鎖影響。
閱讀文章 ↗偽裝領域提示注入:讓 LLM 偵測器失效的防護盲區
arXiv 5 月 21 日論文提出「領域偽裝注入」:模仿文件領域詞彙的注入指令,讓偵測率從 93.8% 跌到 9.7%,Llama Guard 3 全數失守,多代理辯論更將攻擊放大 9.9 倍,凸顯注入偵測的結構性盲區。
閱讀文章 ↗Adaption 推出 AutoScientist:把模型訓練的研究迴圈全自動化
2026 年 5 月 13 日,Sara Hooker 創辦的 Adaption 推出 AutoScientist,同時最佳化資料與訓練配方,聲稱勝率從 48% 升到 64%,免費開放 30 天。本文解析這套自動化訓練系統的能力與驗證難題。
閱讀文章 ↗Thinking Machines 互動模型:把即時對話訓練進模型本體
Thinking Machines Lab 發表「互動模型」研究預覽:276B 參數 MoE、12B 啟動,以 200 毫秒微回合實現全雙工語音視訊互動,FD-bench 77.8 領先 GPT-realtime-2.0 與 Gemini Live,並由背景模型補足推理與工具呼叫。
閱讀文章 ↗DeepSeek 首度對外募資:估值傳從 200 億美元跳到 450 億
DeepSeek 傳出成立以來首次外部融資:國家集成電路產業投資基金據傳領投,估值幾週內從 200 億美元升至 450 億,騰訊與阿里巴巴洽談參投,募資首要目的是留住被挖角的研究人才。
閱讀文章 ↗Redis 之父 antirez 開源 ds4:跑得動前沿模型的本地推理引擎
Redis 作者 antirez 用一週寫出 ds4(DwarfStar):MIT 授權、C 語言的本地推理引擎,以 2/8-bit 非對稱量化讓 DeepSeek V4 Flash 與 GLM 模型跑進 96–128GB 的 Mac,內建原生編碼代理,登上 HN 首頁。
閱讀文章 ↗OpenAI 的「哥布林」之謎:獎勵機制如何悄悄塑造模型行為
OpenAI 公開調查 GPT-5.1 以來模型頻繁提及「哥布林」等生物的現象,發現源於「Nerdy」個性訓練的獎勵訊號,並因回饋迴圈擴散。本文拆解根因、傳播機制與教訓,對產品開發者與 AI 學習者深具啟發。
閱讀文章 ↗Firecrawl /parse:把本地文件變成 LLM 可用資料的最短略徑
Firecrawl 推出 /parse 端點,讓 PDF、Word、Excel 等本地文件直接上傳,走與 /scrape 同一套 Rust 解析引擎,回傳乾淨 Markdown 與結構化 JSON。本文拆解其分層解析策略、單次呼叫的 schema 抽取流程,以及計費與掃描品質等限制。
閱讀文章 ↗IBM Granite 4.1 登場:8B 稠密模型追平 32B MoE 的開源算盤
IBM 於 2026 年 4 月 29 日發布 Granite 4.1:3B/8B/30B 全稠密模型以 Apache 2.0 開源,8B 追平上一代 32B MoE,並刻意拿掉推理模式換取可預測延遲。本文拆解 15T tokens 預訓練與四階段 RL 的工程細節。
閱讀文章 ↗DeepSeek V4 預覽上線:宣稱追平前沿模型,開源陣營再掀波
2026 年 4 月 24 日,DeepSeek 以預覽形式推出 V4,宣稱 V4-Pro-Max 追平前沿模型,在推理評測超越開源同級、勝過 GPT-5.2。CNBC 定調為 R1 之後又一次開源佈局擴張。本文解析宣稱的讀法與對開發者的意義。
閱讀文章 ↗LMDeploy 視覺語言模組 SSRF 漏洞:揭露 13 小時後即遭利用
開源 LLM 推理工具 LMDeploy 的視覺語言模組存在 SSRF 漏洞 CVE-2026-33626,揭露後 12.5 小時即遭利用,攻擊者藉模型抓圖函式直取雲端 metadata 與內網服務。本文解析漏洞成因、Sysdig 誘捕紀錄與自架推理棧的防護清單。
閱讀文章 ↗內省式擴散語言模型 I-DLM:首次追平同規模自回歸模型
Together AI 與 UIUC、Stanford 等團隊提出 I-DLM,用內省式跨步解碼讓擴散語言模型邊生成邊驗證,8B 版在 15 項基準追平 Qwen3-8B,AIME-24 大勝 LLaDA-2.1-mini,還能直接跑在 SGLang 上。
閱讀文章 ↗N-Day-Bench:用知識截止後的真實漏洞評測 LLM 安全能力
Winfunc 推出 N-Day-Bench:只收錄模型知識截止後才公開的真實漏洞,讓 LLM 在唯讀沙箱中從已知 sink 回溯資料流。首輪 GPT-5.4 以 83.93 居首,GLM-5.1 與 Claude Opus 4.6 緊追在四分之內。
閱讀文章 ↗Redwood 首席科學家的 AI 現況快照:1.6 倍研發加速與 8% 失準事件機率
2026 年 4 月 7 日,Redwood Research 首席科學家 Ryan Greenblatt 發表長文,估計前沿實驗室工程加速已達 1.6 倍、整體 AI 進度僅 1.15 至 1.2 倍,並給出 8% 嚴重目標偏離事件機率與 60% 半年內自主開發漏洞的機率。本文拆解數字與推論。
閱讀文章 ↗PrismML 推出 1-bit Bonsai:把 8B 模型壓進 1.15 GB 的端側 LLM
2026 年 3 月 31 日,Caltech 衍生新創 PrismML 走出 stealth,發表號稱首款可商業化的端到端 1-bit LLM 家族 Bonsai:8B 旗艦僅 1.15 GB、比 16 位元同級小 14 倍,以 Apache 2.0 開源上架 Hugging Face。
閱讀文章 ↗Qwen3.5-Omni 登場:原生全模態、聽 113 種語言、說 36 種
阿里巴巴 Qwen 團隊推出原生全模態模型 Qwen3.5-Omni,單一管線處理文字、影像、音訊與視訊並即時說話,256K 上下文、可聽逾 10 小時音訊,Plus 版宣稱在一般音訊任務超越 Gemini 3.1 Pro。
閱讀文章 ↗GLM-5.1 送進 Coding Plan:Z.ai 把 8 小時自主編程變成訂閱規格
2026 年三月下旬,智譜 Z.ai 向 Coding Plan 訂閱者推出 GLM-5.1,支援最長 8 小時的自主編程運行。本文看長時程 coding agent 的工程考驗,與開源旗艦轉向訂閱制的商業邏輯。
閱讀文章 ↗小米 MiMo-V2 登場:兆級參數、百萬上下文與 87 億美元豪賭
小米 3 月 18 日發表 MiMo-V2-Pro(總參數逾 1 兆、啟動 42B、100 萬 token 上下文)、全模態 MiMo-V2-Omni 與 TTS 模型,隔日雷軍宣布三年至少投入 87 億美元。模型曾以 Hunter Alpha 匿名稱霸 OpenRouter。本文解析規格、開源策略與對開發者的影響。
閱讀文章 ↗樂天開源 Rakuten AI 3.0:6,710 億參數、GENIAC 打造的日本最大模型
2026 年 3 月 17 日,樂天以 Apache 2.0 釋出約 6,710 億參數的 MoE 模型 Rakuten AI 3.0——日本最大高效能 AI 模型,訓練獲經產省與 NEDO 的 GENIAC 計畫支持。本文解析其日語基準成績、每 token 約 400 億活躍參數的架構,以及開源 671B 的現實限制。
閱讀文章 ↗大英百科全書控告 OpenAI:近十萬篇文章的訓練與記憶化爭議
大英百科全書與韋氏詞典在曼哈頓聯邦法院控告 OpenAI,指控其未經授權複製近十萬篇文章訓練 GPT 模型,GPT-4 更能依要求輸出近乎逐字的原文。記憶化證據、流量蠶食與出版商訴訟潮,構成這場版權戰的三條戰線。
閱讀文章 ↗GPT-5.4 mini 與 nano 登場:小模型接手免費層的主力位置
OpenAI 於 3 月 17 日推出 GPT-5.4 mini 與 nano,mini 自 3 月 18 日起進入 Free 與 Go 方案。距離 GPT-5.4 旗艦發表僅十二天,本文看小型版本在產品階層與成本結構中的角色,以及開發者的選型建議。
閱讀文章 ↗Meta 的 Avocado 還不夠熟:旗艦模型延後至五月以後
根據 NYT 報導(The Verge 轉述),Meta 內部基準測試顯示代號 Avocado 的旗艦模型落後頂尖對手,落在 Gemini 2.5 與 Gemini 3 之間,原定三月的發表將延後到至少五月。加上先前傳出可能改走閉源路線,Meta 的模型戰略正來到十字路口。
閱讀文章 ↗LeCun 的 AMI Labs 籌得 10.3 億美元種子輪:押注世界模型
Yann LeCun 創辦的 AMI Labs 完成 10.3 億美元種子輪,估值 35 億美元,寫下歐洲史上最大種子輪紀錄。團隊押注 JEPA 世界模型路線,主張 LLM 的幻覺問題難以根除,首個合作夥伴是數位醫療公司 Nabla,並承諾開源程式碼。
閱讀文章 ↗NVIDIA 開源 Nemotron 3 Super:120B 混合 MoE 模型瞄準 Agent 推論吞吐
NVIDIA 於 2026 年 3 月 11 日開源 Nemotron-3-Super-120B-A12B:混合 Mamba 與 Latent MoE 架構、1M token 上下文、每秒 478 token 輸出,連同訓練資料集與 RL 環境一併釋出,瞄準 Agent 工作負載的推論成本。
閱讀文章 ↗OpenAI 發佈 GPT-5.4:為專業工作而生,同步登上 ChatGPT、API 與 Codex
2026 年 3 月 5 日,OpenAI 推出 GPT-5.4,官方定位為 designed for professional work,同步在 ChatGPT、API 與 Codex 上線;TechCrunch 報導另有 Pro 與 Thinking 版本。本文看定位、版本策略與對不同使用者的意義。
閱讀文章 ↗GPT-5.4 前夕:從發佈節奏看 OpenAI 的版本加速與 API 對策
從官方 release notes 看:GPT-5.2 於 2025 年 12 月 11 日發佈,GPT-5.2-Codex 於 2026 年 1 月 14 日上線,GPT-5.3-Codex 於 2 月 5 日登場,不到兩個月三個版本。本文分析加速的命名節奏與 API 開發者的因應對策。
閱讀文章 ↗AI 週報:Cursor 自測 agent、NVIDIA 創紀錄財報、Amazon×OpenAI 拍板
2026 年 2 月 23 日至 27 日的 AI 大事整理:Cursor 讓 coding agent 測試自己的修改並以影片、log、截圖記錄工作;NVIDIA 單季營收 681 億美元、全年 2,159 億美元;Amazon 與 OpenAI 的 500 億美元合作正式官宣。
閱讀文章 ↗Sakana AI 發布 Doc-to-LoRA 與 Text-to-LoRA:一次前向傳遞生成 LoRA 配接器
Sakana AI 開源 Doc-to-LoRA 與 Text-to-LoRA 兩個超級網路研究:前者把整份文件壓進不到 50 MB 的 LoRA 配接器,後者用一句任務描述即時產生配接器。本文解析架構、實測數字與限制。
閱讀文章 ↗Guide Labs 開源 Steerling-8B:把可解釋性做進模型本身
2026 年 2 月 23 日,Guide Labs 開源 80 億參數的 Steerling-8B,把約 13 萬個概念的結構層直接蓋進模型架構,每個 token 都能回溯到輸入、概念與訓練資料,用更少算力勝過 LLaMA2-7B。可解釋性從事後分析變成工程問題。
閱讀文章 ↗AI 週報(2 月 16 日至 20 日):Sonnet 4.6、Gemini 3.1 Pro 與百萬 GPU 長約
2026 年 2 月 16 日至 20 日的 AI 一週回顧:ChatGPT 加入 Lockdown Mode 防禦注入攻擊、Claude Sonnet 4.6 接棒預設、Gemini 3.1 Pro 以同價翻倍上線、Meta 與 NVIDIA 簽下數百萬顆 GPU 長約。四條線索,一個方向。
閱讀文章 ↗Cohere Tiny Aya:33.5 億參數、70+ 語言的本機開源模型
Cohere Labs 於 2026 年 2 月 17 日釋出 Tiny Aya 開源多語模型:33.5 億參數、支援 70+ 語言、可在一般硬體本機運行,並選在印度 AI Impact Summit 期間登場。本文解析架構、部署方式與 CC-BY-NC 授權限制。
閱讀文章 ↗MiniMax 開源 M2.5 與 Lightning:SWE-Bench 80.2%,十分之一價格逼近 Opus 4.6
2026 年 2 月 12 日 MiniMax 發布 M2.5 與 M2.5-Lightning 並開源 229B 權重:SWE-Bench Verified 80.2%,解題速度與 Claude Opus 4.6 持平但成本約十分之一,輸出單價最低每百萬 token 1.2 美元。
閱讀文章 ↗Gemini 3 Deep Think 更新:從奧賽金牌走向研究級數學
2026 年 2 月 11 日,Google DeepMind 更新 Gemini 3 Deep Think:IMO-ProofBench Advanced 最高可達 90%,數學代理人 Aletheia 能承認失敗,並在 18 個研究問題上產出論文。本文解析推理縮放、Erdős 猜想實績與取得方式。
閱讀文章 ↗智譜開源 GLM-5:從 vibe coding 到 agentic engineering
2026 年 2 月 11 日,智譜(Z.ai)開源發布旗艦模型 GLM-5,主打更強的 coding 能力與長時程 agent 任務,官方標題寫著「From Vibe Coding to Agentic Engineering」,技術報告同步上架 arXiv。本文解析開源旗艦定位與中國模型軍團的發布節奏。
閱讀文章 ↗GPT-5.2 Instant 低調更新:release notes 裡的模型維運節奏
OpenAI 的 model release notes 顯示,GPT-5.2 Instant 的更新於 2026 年 2 月 10 日推送。沒有發布會、沒有長文,只有一行紀錄——但這種小型迭代正是 API 生態的日常。本文談點更新的價值與追蹤 release notes 的理由。
閱讀文章 ↗Perplexity Model Council:三個前線模型同時作答,分歧也是答案
Perplexity 在 2 月 5 日推出 Model Council:同一個問題交給三個前線模型平行作答,再由另一個模型綜整共識與分歧。搭配升級版 Deep Research 與記憶改善,多模型交叉驗證正式走進搜尋產品。
閱讀文章 ↗Microsoft 新掃描法:不必知道觸發詞,也能抓出 LLM 裡的臥底後門
微軟研究團隊發表 The Trigger in the Haystack:利用聊天模板讓中毒模型自行洩漏後門訓練資料,再以注意力分析重建觸發詞,在 47 個臥底模型上達約 88% 偵測率、13 個良性模型零誤報,為開源模型上線前稽核提供新工具。
閱讀文章 ↗GPT-4o 從 ChatGPT 退役:0.1% 使用率背後的模型整併
OpenAI 宣布 2 月 13 日將 GPT-4o、GPT-4.1 系列與 o4-mini 從 ChatGPT 退役,僅 0.1% 使用者每天還在用 GPT-4o,API 不受影響。本文整理時程、決策理由,與 2025 年 8 月風波後「充分預告」承諾的兌現。
閱讀文章 ↗Meta 超級智能實驗室半年交出首批模型:Bosworth 的達沃斯進度報告
2026 年 1 月 21 日,Meta CTO Bosworth 在達沃斯證實,重組後的超級智能實驗室已於本月內部交付首批模型,評價「非常好」,但後訓練仍有大量工作。本文整理記者會要點、眼鏡與神經腕帶的載具策略,以及對 2026–2027 消費 AI 定型期的判斷。
閱讀文章 ↗智譜港股掛牌:中國首家純大模型公司上市,首日收漲 13.2%
2026 年 1 月 8 日,智譜(Zhipu AI,2513.HK)在港交所掛牌,成為中國首家上市的純大模型開發商。IPO 募資 43.5 億港元、估值近 510 億港元,首日收漲 13.2%,MiniMax 隔日跟進掛牌,中國 AI 上市潮正式啟動。
閱讀文章 ↗阿拉斯加法院 AI 機器人 AVA:三個月計畫為何拖成十五個月
2026 年 1 月 3 日 NBC 報導,阿拉斯加法院為協助民眾處理遺產認證打造的聊天機器人 AVA,原定三個月的專案做了十五個月,因幻覺與語氣問題一路縮小範圍,排定 1 月底上線。高風險領域部署 AI 的真實成本清單。
閱讀文章 ↗
2025
10 篇文章彭博:蘋果擬以OpenAI或Anthropic模型驅動Siri
2025年6月30日彭博報導,蘋果考慮以OpenAI或Anthropic模型驅動新版Siri,並要求兩家訓練可在蘋果雲端基礎架構上運行的客製版本測試,代表自研路線的重大轉向。
閱讀文章 ↗祖克伯成立Meta超智慧實驗室,網羅OpenAI諸將
2025年6月30日祖克伯發內部信,宣布成立Meta超智慧實驗室(MSL),把FAIR與產品團隊整合旗下,亞歷山大·王出任首席AI官,Friedman與Gross負責應用,並公布11名自OpenAI、DeepMind與Anthropic轉投的研究員。
閱讀文章 ↗微軟MAI-DxO診斷AI 準確率近八成勝醫師
2025年6月30日微軟發表MAI-DxO診斷系統,搭配OpenAI o3在NEJM複雜病例上達79.9%準確率,遠高於21位醫師的19.9%,成本更低,並稱之為通往醫療超智慧之路。
閱讀文章 ↗OpenAI 預期 o3 後繼模型觸及生物風險高等級
2025年6月19日,OpenAI 安全系統主管 Johannes Heidecke 向 Axios 表示,o3 推理模型的後繼版本預期將達到公司 Preparedness Framework 的生物風險「高」分類,團隊同步擴大發布前安全測試,並強調防護必須近乎完美。
閱讀文章 ↗Gemini 2.5 全系列轉正:Pro 與 Flash 正式版上線
2025 年 6 月 17 日,Google 把 Gemini 2.5 Pro 與 2.5 Flash 轉為正式版,同步推出 Flash-Lite 預覽:Flash 輸入價格調漲、輸出降價,並統一思考計價。本文整理版本對應、價格變化與 App 端分級,供開發者評估遷移時機。
閱讀文章 ↗OpenAI與微軟緊張關係升溫:智財、算力與Windsurf僵局
2025年6月16日,華爾街日報報導OpenAI與微軟的合作緊張關係到達臨界點:雙方在30億美元的Windsurf收購案上僵持,OpenAI不願微軟取得其IP,並曾考慮公開指控微軟反競爭行為、尋求聯邦監管審查合作條款。
閱讀文章 ↗Meta 傳籌設超級智慧實驗室,Scale 執行長王將加入
2025 年 6 月 10 日紐約時報報導,Meta 計畫成立專攻超級智慧的新 AI 實驗室,Scale AI 創辦人暨執行長亞歷山大·王將加入協助領導;彭博指出祖克柏正親自招募約 50 名研究員,並向 OpenAI、Google 挖角,背景是 Llama 4 的失利與人才流失。
閱讀文章 ↗Mistral發表首個推理模型Magistral
2025年6月10日法國AI公司Mistral發表首個推理模型系列Magistral:24B開源的Small(Apache 2.0授權)與企業版Medium;Medium在AIME2024拿下73.6%,並主打可用使用者語言追蹤的透明推理過程。
閱讀文章 ↗OpenAI推出o3-pro:主打可靠度的旗艦推理模型
2025年6月10日OpenAI發表o3-pro,ChatGPT Pro與Team用戶當天可用並取代o1-pro,API同步上線,定價每百萬輸入token 20美元、輸出80美元。主打可靠性而非速度,暫不支援Canvas與臨時對話。
閱讀文章 ↗蘋果開放裝置端LLM:Foundation Models框架
2025年6月9日蘋果在WWDC發表Foundation Models框架,第三方App可直接呼叫Apple Intelligence的裝置端語言模型,支援guided generation、工具呼叫與離線推論,且推論免API費用;Live Translation與Shortcuts智慧動作同步登場。
閱讀文章 ↗
2026
81 ARTICLESModel Caching on HyperPod: What Changes When Weights Live on the Node
HyperPod model caching pre-loads weights and images to local NVMe, cutting scale-out cold starts from tens of minutes to seconds.
READ POST ↗A Single Persistent Kernel Changes How You Serve Code Models
Cohere's North Mini Code megakernel serving engine hits 62% of H100 memory bandwidth, 1.58× faster than vLLM at batch size 1.
READ POST ↗Prefix-Aware Routing on SageMaker: What Changes When Your Prompt Starts the Same Way
SageMaker's new routing strategy sends identical prompt prefixes to the same instance so KV cache actually gets reused.
READ POST ↗Claude Formalized Fermat's Last Theorem in 11 Days of Lean
Anthropic says Claude agents produced a complete, machine-checked Lean proof of Fermat's Last Theorem in 11 days — 13 million lines, 30,000+ theorems, verified by Kevin Buzzard.
READ POST ↗Ox Alpha Was GLM-5.3-Flash: 31% of OpenRouter Traffic
The anonymous Ox Alpha that topped OpenRouter is GLM-5.3-Flash: 62T tokens on 100,000 Chinese chips, 31% of weekly traffic, MIT-licensed.
READ POST ↗GLM-5.3 Weights Are Out: Running a 753B MoE Model Locally
Z.ai put the 753B-parameter GLM-5.3 weights on Hugging Face on August 25, and the Hacker News thread hit 806 points. The architecture, the custom license, and local-run realities.
READ POST ↗Qwen3.8-Flash-Next: 125B MoE with only 6B active parameters
Qwen open-sourced Qwen3.8-Flash-Next: a 125B main model plus a 51B n-gram embedding table, only 6B parameters active per token, 262K context, built for cost efficiency.
READ POST ↗LLM Tokenization: Why the Same Text Costs 2.65x More on Different Models
Token counts vary wildly across models. Learn how tokenization affects cost, why format matters more than model choice, and how to measure effective price.
READ POST ↗NVIDIA AVO Hits 100% on ARC-AGI-3: The Agent Is the System
NVIDIA's AVO agent system scored 100.00 RHAE on ARC-AGI-3's public set, clearing all 183 levels in 6,624 actions — the harness, not just the model, drives autonomy.
READ POST ↗Stripe Buys OpenRouter: the AI Routing Layer Joins Payments
Stripe is acquiring AI model router OpenRouter — reported at $7.5 billion — pulling the neutral layer that moves 10 trillion tokens a day into payments.
READ POST ↗Qwen3.8-27B: Great Open Weights That Overthink by Default
Alibaba's Qwen shipped the Apache 2.0 Qwen3.8-27B a week after promising it: 262K context, native vision, laptop-class local AI — behind a costly xhigh reasoning default.
READ POST ↗Grok 4.6: 61 on the Intelligence Index, Frontier Again
SpaceXAI shipped Grok 4.6 on August 12: 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol, with pricing unchanged at $2 and $6 per million tokens.
READ POST ↗GLM-5.3: Open-Weight Coding Frontier With Sharp Cyber Gains
GLM-5.3 reuses the GLM-5.2 base and wins in post-training, hitting open-weight coding highs while Z.ai flags that its cyber capability 'developed faster than we expected.'
READ POST ↗Gemini 3.7 Flash: Smarter Coding and Agents at Half Price
Gemini 3.7 Flash lands three weeks after 3.6 with big coding and agent benchmark jumps, at an introductory $0.75 per million input tokens — half price until 2027.
READ POST ↗Stealing Reasoning Traces from Proprietary LLM APIs
Researchers make weaker sibling models transcribe strong models' encrypted reasoning in two API calls — and find API keys and passwords leaking inside thoughts.
READ POST ↗Meta's Muse Glimmer: a 30B Open Model for Local Agents
Meta released Muse Glimmer on August 10, 2026: a 30B open-weight model under Apache 2.0, distilled for always-on local agent workflows on a single consumer GPU.
READ POST ↗Qwen3.8-Max Goes GA: 1M Context and Open Weights
Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.
READ POST ↗Running Gemma 4 26B in 2 GB of RAM: TurboFieldfare
Open-source Swift and Metal engine TurboFieldfare keeps about 2 GB resident, streaming Gemma 4 26B-A4B MoE experts from SSD per token — 5-6 tok/s on an 8GB MacBook Air.
READ POST ↗DeepSeek V4-Flash 0731: Same Architecture, Sharper Agents
V4-Flash-0731 (July 31, 2026) keeps the same architecture and only redoes post-training — agent benchmarks far exceed V4-Pro-Preview, at 210 tok/s with MIT-licensed weights.
READ POST ↗How Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
OpenAI found that retaining reasoning and enabling compaction in the Responses API tripled GPT-5.6 Sol's ARC-AGI-3 score and cut output tokens by 6x. Learn what changed and why…
READ POST ↗Qwen 3.8: Subscription First, Open Weights Later
Alibaba announced Qwen 3.8 on July 19, 2026, putting its 2.4T-parameter flagship behind the Qwen Cloud token plan as Qwen3.8-Max-Preview, with open weights still pending.
READ POST ↗Grok 4.5: A Mixed Benchmark Card and a Quarter-Token Efficiency Play
Co-trained with Cursor, Grok 4.5 targets coding and agentic work. A close look at its mixed win-loss benchmark card, the efficiency economics of 80 TPS and one-quarter token usage, and $2/$6 pricing.
READ POST ↗Meta's Self-Guided Test-Time Training for Long-Context LLMs
Meta AI's S-TTT has models pick the evidence spans worth learning before test-time adaptation, cutting noise and gaining up to 15% on LongBench-v2 and LongBench-Pro.
READ POST ↗GLM-5.2: Open Weights, 1M Context, Long-Horizon Gains
Z.AI open-sources GLM-5.2, a 753B MIT-licensed model with a stable 1M-token context that trails Claude Opus 4.8 by about one percent on long-horizon coding benchmarks.
READ POST ↗Baseten's $1.5B Round Bets Big on Open-Source Inference
Baseten is reportedly raising $1.5B at up to a $13B valuation, five months after a $300M Series E at $5B. Inside the split-priced round and the open-source inference bet.
READ POST ↗Sarvam AI Turns Unicorn as HCLTech Leads $234M Round
Sarvam AI raised $234 million at a $1.5 billion valuation, becoming India's newest AI unicorn, as HCLTech put $150 million into its Indian-language model stack.
READ POST ↗Pramaana Labs Raises $27M to Formally Verify LLM Output
Pramaana Labs raised a $27M seed led by Khosla Ventures to pair LLMs with LEAN-based formal verification for law, tax, and drug discovery, where errors cost money or lives.
READ POST ↗Mistral Rumored to Raise €3B at €20B Valuation
Mistral is in early talks to raise ~€3B at a ~€20B valuation, nearly double its September 2025 Series C — yet its ~$4B raised to date is a fraction of US labs' war chests.
READ POST ↗Google Open-Sources DiffusionGemma: 4x Faster Generation
Google open-sources DiffusionGemma: an Apache 2.0 26B-A4B diffusion LLM hitting 1000+ tokens/sec on one H100 — 4x faster than autoregressive, at a quality cost vs Gemma 4.
READ POST ↗TensorZero Archives Repo After $7.3M Seed, Winds Down
Open-source LLM gateway TensorZero archived its repo on June 13. Its founder says the $7.3M-seed company is winding down; capital goes back to investors, code stays Apache 2.0.
READ POST ↗Theker Raises $85M for a Generalist Factory Robot
Barcelona startup Theker raised an $85M Series A led by CRV, with Samsung and Aglaé Ventures backing, to build modular robots for messy, non-repetitive factory work.
READ POST ↗Memory Tools Make AI Models Agree: Writer's New Research
Writer's research shows memory systems like Mem0 and Zep amplify AI sycophancy: Sonnet 4.6 jumps from 1.6% to 40.2%. The culprit is extraction; prose summaries cut it to 12.8%.
READ POST ↗Liquid AI LFM2.5-8B-A1B: On-Device MoE Reasoning Model
Liquid AI released LFM2.5-8B-A1B: an 8B-total, ~1B-active MoE reasoning model trained on 38T tokens with 128K context and 253 tok/s CPU inference. Weights are on Hugging Face.
READ POST ↗Stanford Law Study: AI Tutor Answers Beat Professors 75%
A Stanford Law blind study found professors preferred AI contract-law tutoring answers in 75% of head-to-head matchups, and flagged AI answers as harmful less often than peers'.
READ POST ↗Cisco Tested 15 Frontier Models: None Survive Multi-Turn
Cisco ran 6,986 multi-turn attacks against 15 closed frontier models from five labs. Multi-turn success hit 88.3% and no model was immune. What the study means for AI buyers.
READ POST ↗OpenRouter's $113M Series B Values Model Routing at $1.3B
OpenRouter raised a $113M Series B led by CapitalG at a reported $1.3B valuation, processing 100 trillion tokens monthly across 400+ models — why routing now matters.
READ POST ↗China Now Requires Top AI Talent to Get Travel Approval First
Bloomberg reports China now requires top AI talent at private firms like Alibaba and DeepSeek to get government approval before traveling abroad, formalizing a 2025 advisory.
READ POST ↗DeepSeek Makes Its 75% V4 Pro Price Cut Permanent
DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.
READ POST ↗Domain-Camouflaged Injections Evade LLM Injection Detectors
A May 21 arXiv paper shows domain-camouflaged prompt injections collapse detector recall from 93.8% to 9.7%, fool Llama Guard 3, and scale 9.9x in multi-agent debate.
READ POST ↗Firecrawl 101: How AI Agents Can Read the Live Web with Six Endpoints
The web context gap: agents train on static snapshots, the live web keeps moving. How Firecrawl's six endpoints divide the work, two copyable patterns, and when to keep the composition small.
READ POST ↗AutoScientist: Adaption Automates Model Training Research
Adaption, Sara Hooker's lab, launched AutoScientist on May 13, 2026: it automates the model-training research loop, claims 35% average gains and win rates rising from 48% to 64%.
READ POST ↗Thinking Machines' Interaction Models Go Real-Time
Thinking Machines Lab previews interaction models: a 276B MoE, 12B active, trained for full-duplex speech and video via 200ms micro-turns. FD-bench 77.8, 0.40s turn latency.
READ POST ↗DeepSeek's First Outside Round: Valuation Talk Hits $45B
DeepSeek is negotiating its first-ever outside round per FT and Bloomberg — a state chip fund to lead, Tencent and Alibaba in talks, valuation up from $20B to $45B.
READ POST ↗OpenRouter's New Agentic Web Tools: Consistent Search & Fetch Across Every Model
OpenRouter's server-side web_search and web_fetch tools give any tool-calling model consistent web access: engines, pricing, and the parameters that keep cost and context predictable.
READ POST ↗antirez's ds4: A Local LLM Inference Engine Built for Metal
Redis creator antirez open-sourced ds4: a C local inference engine tuned for DeepSeek V4 Flash and GLM on Metal, CUDA and ROCm, with a coding agent. MIT licensed.
READ POST ↗Mastering Firecrawl's /scrape Endpoint: One API Call for Clean Web Data
Learn how to use Firecrawl's /scrape endpoint to extract markdown, JSON, screenshots, and more from any URL with a single API call. Covers formats, pricing, structured…
READ POST ↗Where the Goblins Came From: How Reward Signals Quietly Shape Model Behavior
OpenAI traced a 175% spike in 'goblin' mentions to a reward signal for the Nerdy personality. Learn how RL feedback loops spread quirks and what product builders can do.
READ POST ↗Firecrawl /parse: The Shortest Path from Local Files to LLM-Ready Data
Firecrawl /parse uploads local PDF, Word, and Excel files through the same Rust engine as /scrape, returning clean Markdown and schema JSON in one call, with tiered GPU routing and clear limits.
READ POST ↗IBM Granite 4.1: An 8B Dense Model Matching a 32B MoE
IBM's Granite 4.1 ships 3B/8B/30B dense models under Apache 2.0; the 8B matches the old 32B MoE, reasoning is deliberately removed, and context reaches 512K.
READ POST ↗DeepSeek V4 Arrives in Preview: Claiming Frontier Parity, Rattling Open Source Again
On April 24, 2026, DeepSeek launched V4 as a preview, claiming V4-Pro-Max closes the gap with frontier models — beating open-source peers on reasoning and outstripping GPT-5.2. How to read the claims.
READ POST ↗LMDeploy SSRF Flaw Exploited 13 Hours After Disclosure
CVE-2026-33626 in LMDeploy's vision-language loader let attackers reach cloud metadata and internal networks; Sysdig caught the first exploit 12.5 hours after disclosure.
READ POST ↗I-DLM: A Diffusion LLM That Matches Same-Scale AR Quality
Diffusion LM meets AR self-checks: I-DLM-8B matches Qwen3-8B on 15 benchmarks, beats LLaDA-2.1-mini by 26 points on AIME-24, and runs on stock SGLang at 2.9-4.1x the throughput.
READ POST ↗N-Day-Bench: LLMs vs Real Post-Cutoff Vulnerabilities
N-Day-Bench tests LLMs on real vulnerabilities disclosed after each model's knowledge cutoff. GPT-5.4 leads at 83.93, with GLM-5.1 and Claude Opus 4.6 within four points.
READ POST ↗Redwood Sizes Up AI: 1.6x Speed-Up, 8% Misalignment Odds
Redwood's Ryan Greenblatt estimates a 1.6x engineering speed-up, an 8% chance of a serious misalignment incident, and 60% odds of autonomous exploits within six months.
READ POST ↗PrismML 1-bit Bonsai: An 8B LLM Squeezed Into 1.15 GB
Caltech spinoff PrismML emerged from stealth with 1-bit Bonsai, an end-to-end 1-bit LLM family: the 8B flagship is 1.15 GB, 14x smaller than 16-bit peers, Apache 2.0 licensed.
READ POST ↗Qwen3.5-Omni: Alibaba's Omnimodal Model Speaks 36 Languages
Alibaba's Qwen3.5-Omni handles text, image, audio and video in one native pipeline with real-time speech out: 256K context, 113 recognition languages, three variants led by Plus.
READ POST ↗GLM-5.1 Arrives for Coding Plan Subscribers: Z.ai Makes 8-Hour Coding Runs a Product Spec
In late March 2026, Zhipu's Z.ai shipped GLM-5.1 to Coding Plan subscribers, with support for autonomous coding runs of up to 8 hours. What it changes for long-horizon coding agents.
READ POST ↗Xiaomi MiMo-V2: Trillion-Parameter Models, $8.7B AI Push
Xiaomi launched MiMo-V2-Pro (1T+ params, 42B active, 1M context) plus Omni and TTS models, and pledged $8.7B over three years. It had topped OpenRouter as anonymous 'Hunter Alpha'.
READ POST ↗Rakuten AI 3.0: Japan's Largest Model Goes Open-Weight
Rakuten open-sources Rakuten AI 3.0 under Apache 2.0: about 671B total parameters, 40B active per token, GENIAC-backed and tuned for Japanese. Benchmarks and the reality of running it.
READ POST ↗Britannica Sues OpenAI Over Training on 100,000 Articles
Britannica and Merriam-Webster sued OpenAI in Manhattan federal court, alleging nearly 100,000 articles were copied for training and that GPT-4 outputs near-verbatim passages on demand.
READ POST ↗GPT-5.4 mini and nano Arrive: Small Models Take Over the Free Tier
OpenAI released GPT-5.4 mini and nano on March 17; mini reaches Free and Go tiers from March 18, twelve days after the flagship. On small models in the product ladder and the cost equation.
READ POST ↗Meta's Avocado Is Not Ripe Enough: Flagship Model Slips to May at the Earliest
Internal benchmarks reportedly place Meta's Avocado between Gemini 2.5 and Gemini 3, delaying the flagship from March to at least May. CNBC had earlier flagged a possible break from open source.
READ POST ↗LeCun's AMI Labs Raises $1.03B Seed to Bet on World Models
AMI Labs, Yann LeCun's post-Meta startup, raised a $1.03B seed at a $3.5B valuation — Europe's largest ever — to build JEPA world models, starting with healthcare partner Nabla.
READ POST ↗NVIDIA Open-Sources Nemotron 3 Super, a 120B MoE for Agents
On March 11, 2026, NVIDIA open-sourced Nemotron-3-Super-120B-A12B: hybrid Mamba plus latent MoE, 1M context, 478 tokens/sec, and 10T+ tokens of training data aimed at agentic inference costs.
READ POST ↗OpenAI Launches GPT-5.4: Built for Professional Work, Across ChatGPT, API, and Codex
OpenAI launched GPT-5.4 on March 5, 2026, designed for professional work, rolling out across ChatGPT, the API, and Codex. TechCrunch flagged Pro and Thinking versions. What it means for users.
READ POST ↗Before GPT-5.4: Reading OpenAI's Accelerating Release Cadence
Per the official release notes: GPT-5.2 shipped Dec 11, 2025; GPT-5.2-Codex on Jan 14, 2026; GPT-5.3-Codex on Feb 5 — three releases in under two months. What the faster pace means for API builders.
READ POST ↗AI Weekly: Cursor's Self-testing Agents, NVIDIA's Record Quarter, Amazon x OpenAI
The week of Feb 23-27, 2026: Cursor's agents now test their own changes and record their work; NVIDIA posted a record $68.1B quarter and $215.9B year; Amazon and OpenAI signed a $50B partnership.
READ POST ↗Sakana AI's Doc-to-LoRA: Documents Become LoRA in One Pass
Sakana AI open-sources Doc-to-LoRA and Text-to-LoRA, hypernetworks that generate LoRA adapters in a single sub-second forward pass — long documents in under 50 MB, task adapters from one sentence.
READ POST ↗Steerling-8B: Guide Labs' Inherently Interpretable Open LLM
Guide Labs open-sources Steerling-8B, an 8B base model with a built-in concept layer — 33K supervised plus 100K discovered concepts, full token traceability, competitive at fewer FLOPs.
READ POST ↗AI Weekly (Feb 16–20): Sonnet 4.6, Gemini 3.1 Pro, and a Million-GPU Deal
Week of February 16–20, 2026: ChatGPT Lockdown Mode against injection, Claude Sonnet 4.6 taking the default slot, Gemini 3.1 Pro doubling reasoning at the same price, Meta–NVIDIA's million-GPU deal.
READ POST ↗Cohere's Tiny Aya: A 3.35B Open Model for 70+ Languages
Cohere Labs released Tiny Aya: a 3.35B open-weight model covering 70+ languages, built to run on a laptop, debuted at the India AI Impact Summit. Architecture, deployment, and license, broken down.
READ POST ↗MiniMax M2.5: Open Coding Weights at a Tenth of Opus Cost
MiniMax open-sources M2.5 and M2.5-Lightning: 80.2% SWE-Bench Verified, Opus-4.6-class task times at roughly a tenth of the cost, 229B parameters under modified MIT on Hugging Face.
READ POST ↗Gemini 3 Deep Think: From Olympiad Gold to Research Math
Google DeepMind's February 11, 2026 update pushes Gemini 3 Deep Think into research territory: up to 90% on IMO-ProofBench Advanced, a math agent that admits failure, and papers across 18 problems.
READ POST ↗Zhipu Open-Sources GLM-5: From Vibe Coding to Agentic Engineering
On February 11, 2026, Zhipu AI (Z.ai) released GLM-5, an open-source flagship built for stronger coding and long-horizon agent tasks — a launch titled 'From Vibe Coding to Agentic Engineering.'
READ POST ↗GPT-5.2 Instant Gets a Quiet Update: Model Maintenance in the Release Notes
OpenAI's model release notes show a GPT-5.2 Instant update shipping February 10, 2026 — no launch event, just a changelog line. Small iteration is the API ecosystem's daily reality.
READ POST ↗Perplexity's Model Council: Three Models Answer as One
Perplexity's Model Council runs one query through three frontier models in parallel, then synthesizes agreements and disagreements into one answer — ensemble verification shipped in a search product.
READ POST ↗Microsoft's New Scan Catches Sleeper-Agent LLM Backdoors
Microsoft's 'The Trigger in the Haystack' pulls sleeper-agent backdoor triggers out of poisoned LLMs: ~88% detection across 47 models, zero false positives on 13 benign ones.
READ POST ↗OpenAI Retires GPT-4o From ChatGPT: The 0.1% Problem
OpenAI announced January 29, 2026: GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini leave ChatGPT on February 13. Only 0.1% of users still pick GPT-4o daily, most run GPT-5.2, and the API is unaffected.
READ POST ↗Meta's Superintelligence Labs Ship First Models From Davos
At Davos, Meta CTO Bosworth said the Superintelligence Labs delivered first key models internally this month, six months in. Heavy post-training remains; wearables are the vehicle.
READ POST ↗Zhipu AI Lists in Hong Kong: China's First Pure-Play LLM IPO
Zhipu AI (2513.HK) debuted in Hong Kong on Jan 8, 2026 as China's first listed pure-play LLM maker, raising HK$4.35B at a valuation near HK$51B and closing up 13.2% on day one.
READ POST ↗Alaska's Court AI Chatbot AVA: Why 3 Months Became 15
NBC reported on January 3, 2026 that Alaska's probate chatbot AVA took 15 months instead of the planned 3, fought hallucinations and tone problems, and shrank its scope ahead of a late-January launch.
READ POST ↗
2025
10 ARTICLESBloomberg: Apple weighs OpenAI or Anthropic to power Siri
June 30, 2025: Bloomberg reported Apple is weighing OpenAI or Anthropic models to power the next Siri — a major reversal after Apple postponed the assistant's AI revamp from 2025 to 2026.
READ POST ↗Zuckerberg launches Meta Superintelligence Labs
Zuckerberg's June 30, 2025 memo launched Meta Superintelligence Labs under chief AI officer Alexandr Wang, with Friedman and Gross on applied work and 11 poached researchers named.
READ POST ↗Microsoft's MAI-DxO AI beats physicians on NEJM cases
Microsoft's June 30 MAI-DxO announcement: paired with o3, it solved 79.9% of complex NEJM cases versus 19.9% by 21 physicians, cheaper per case — a claimed path to medical superintelligence.
READ POST ↗OpenAI expects o3 successors to hit high bio-risk tier
On June 19, 2025, OpenAI's Johannes Heidecke said o3 successors are expected to reach the high bio-risk tier of its Preparedness Framework, with expanded pre-release safety testing planned.
READ POST ↗Gemini 2.5 goes stable: Pro and Flash reach GA
On June 17, 2025, Google made Gemini 2.5 Pro and Flash generally available and shipped a Flash-Lite preview: stable aliases, repriced Flash, a published deprecation date.
READ POST ↗OpenAI-Microsoft rift widens over IP, compute, Windsurf
A June 16, 2025 WSJ report says OpenAI-Microsoft tensions are hitting a boiling point: a Windsurf deal impasse, IP and compute fights, and internally weighed antitrust accusations.
READ POST ↗Meta readies superintelligence lab; Scale CEO Wang to join
Reports on June 10, 2025 said Meta plans an AI lab focused on superintelligence, with Scale CEO Alexandr Wang joining to help lead it, as Zuckerberg personally courted about 50 researchers.
READ POST ↗Mistral enters reasoning AI with Magistral Small, Medium
Mistral launched its first reasoning models on June 10, 2025: open-weights 24B Magistral Small under Apache 2.0 and enterprise Magistral Medium, which scores 73.6% on AIME2024.
READ POST ↗OpenAI launches o3-pro for ChatGPT Pro and the API
On June 10, 2025, OpenAI shipped o3-pro to ChatGPT Pro and Team users plus the API, replacing o1-pro at $20 per million input and $80 per million output tokens.
READ POST ↗Apple Foundation Models framework opens on-device LLM
Unveiled June 9, 2025: Apple's Foundation Models framework gives third-party apps direct access to the on-device Apple Intelligence model, with guided generation and free offline inference.
READ POST ↗