2026
21 篇文章把語音、影片、圖片都收進同一個 API:OpenRouter 的介面收斂策略
OpenRouter 把 TTS、影片與圖片生成收進單一請求形狀,產品團隊該重新評估自建多供應商抽象層的成本。
閱讀文章 ↗把 30 秒長鏡頭寫進 API:Seedance 2.5 的規格取捨與計費邏輯
Seedance 2.5 支援 30 秒單次生成與影片參考輸入,但解析度上限只有 720p,計費依影片 token 而非秒數。
閱讀文章 ↗把模型參數移出程式碼:用 OpenRouter Presets 管理 LLM 設定
OpenRouter Presets 讓你把模型、提示詞與路由規則集中管理,改一次就同步所有應用,不必重新部署。
閱讀文章 ↗ZDR 不是隱私萬靈丹:把「零保留」當成可強制執行的路由條件
ZDR 只保證推論供應商不儲存 prompt 與回應,不涵蓋你的日誌、工具或快取;用 provider.zdr 把它變成請求層級的硬性路由條件。
閱讀文章 ↗Muse Spark 1.1 把「工具呼叫」變成產品架構問題:Meta Model API 公開預覽的實務訊號
Meta 推出 Muse Spark 1.1 與 Meta Model API 公開預覽,主打百萬 token 上下文、多代理編排與電腦操作,開發者要重新思考工具層設計。
閱讀文章 ↗零資料保留不讓步:OpenAI 只看訊號,不看內容
OpenAI 推出 Private Safety Processing:客戶內容與金鑰留在客戶端,OpenAI 只收回活動類型訊號,在零資料保留承諾下仍能跨對話偵測試探、協調與代理偏離等風險,白皮書預計 9 月發布。
閱讀文章 ↗DeepSeek V4-Flash 0731:只重做後訓練,代理能力越級跳
2026 年 7 月 31 日,DeepSeek 推出 V4-Flash-0731:架構與規模不變、僅重新後訓練,代理基準大幅超越 V4-Pro-Preview,輸出速度每秒 210 個 token,權重續以 MIT 授權開放。
閱讀文章 ↗OpenRouter Unified Image API:統一請求之前,先讓程式看得懂模型差異
OpenRouter 推出專屬 Image API:單一請求格式接 30 多個模型,能力、參數與計價全部變成可查詢的 schema。本文整理 capability descriptors、三種計費單位、streaming 支援,以及官方的 migration 建議。
閱讀文章 ↗OpenRouter Fusion 登場:多模型融合輸出在 DRACO 勝過最強單一模型
OpenRouter 於 6 月 12 日推出 Fusion:單一 API 呼叫把 prompt 平行送進多個模型,再由裁判模型整合共識與矛盾後交回原模型作答。DRACO 深度研究基準上,融合面板拿下 69.0%,超越所有單一模型,平價組合以約一半成本逼近 Fable 5。
閱讀文章 ↗OpenRouter 完成 1.13 億美元 B 輪:模型路由層估值翻倍至 13 億美元
2026 年 5 月,OpenRouter 宣布 1.13 億美元 B 輪融資,由 CapitalG 領投,估值翻倍至約 13 億美元。每月處理 100 兆 token、支援 400+ 模型,本文解析模型路由層為何成為 Agent 時代的關鍵基礎設施。
閱讀文章 ↗Claude Opus 4.8 發布:Dynamic Workflows 與更誠實的程式碼模型
Anthropic 在 41 天後推出 Claude Opus 4.8:價格不變,新增可協調上百個子代理的 Dynamic Workflows 與 Effort 控制,SWE-bench Pro 升至 69.2%,Mythos 預告數週內開放。
閱讀文章 ↗DeepSeek 把 V4 Pro 七五折降價常態化:前沿模型價格戰的底牌
2026 年 5 月 22 日,DeepSeek 宣布旗艦 V4 Pro 的 75% 折扣轉為常態價:輸出每百萬 token 0.87 美元,僅為原價四分之一,Flash 更只要 0.28 美元。本文解析降價常態化的成本邏輯與對 API 市場的連鎖影響。
閱讀文章 ↗視覺代理比結構化 API 貴 45 倍:Reflex 的完整實測
Reflex 讓 Claude 用兩種方式完成同一個後台任務:視覺 Computer Use 燒掉約 55 萬輸入 tokens、花 17 分鐘;結構化 API 只用 1.2 萬 tokens、20 秒完成。45 倍差距來自介面設計,不是模型好壞。
閱讀文章 ↗Deep Max:為何 Exa 把代理搜尋的速度推上新高?
Exa 推出 Deep Max 代理搜尋端點,結合前沿 LLM 與平行搜尋,在主要 agentic 搜尋基準達到 SOTA 準確率、速度領先對手達 20 倍,但定價未公開。本文拆解其設計原理與產品端的實際取捨。
閱讀文章 ↗微軟 MAI 三模型上線 Foundry:轉錄、語音與影像
2026 年 4 月 2 日,微軟把自研的 MAI-Transcribe-1、MAI-Voice-1、MAI-Image-2 送上 Microsoft Foundry 公開預覽:轉錄 GPU 成本砍半、一秒生成一分鐘語音、影像模型首發衝上 Arena.ai 第三名。
閱讀文章 ↗Google Lyria 3 Pro 登場:AI 音樂生成從 30 秒片段走向三分鐘完整曲目
3 月 25 日 Google DeepMind 推出 Lyria 3 Pro,可生成長達 3 分鐘、含完整段落結構的曲目,較 Lyria 3 的 30 秒大幅躍進。本文看它進駐 Gemini、Vids、Vertex AI 與 ProducerAI 的布局、SynthID 浮水印與訓練資料立場。
閱讀文章 ↗Gemini Embedding 2 公開預覽:文字、影像、音訊、影片共用一個向量空間
Google DeepMind 於 2026 年 3 月 10 日推出 Gemini Embedding 2 公開預覽:第一個原生多模態嵌入模型,把文字、影像、影片、音訊與 PDF 映射到單一 3072 維向量空間,支援 Matryoshka 截斷。本文解析輸入限制、成本槓桿與對 RAG 管線的實際影響。
閱讀文章 ↗Gemini 3.1 Flash Lite 預覽登場:Gemini 3 家族最快最便宜的模型
Google 於 3 月 3 日釋出 Gemini 3.1 Flash Lite 預覽版:輸入每百萬 token 0.25 美元、輸出 1.50 美元,是 Gemini 3 家族最快最便宜的選項。本文解析定價、基準與選型建議。
閱讀文章 ↗Gemini 3.1 Pro 預覽上線:同價推理翻倍,API 端點 gemini-3.1-pro-preview
Google 於 2026 年 2 月 19 日推出 Gemini 3.1 Pro preview,推理能力為 Gemini 3 Pro 的兩倍、價格維持不變,API 端點 gemini-3.1-pro-preview 同步提供。當能力差距越來越難分辨,性價比成為新的比較軸。
閱讀文章 ↗GPT-4o 從 ChatGPT 退役:0.1% 使用率背後的模型整併
OpenAI 宣布 2 月 13 日將 GPT-4o、GPT-4.1 系列與 o4-mini 從 ChatGPT 退役,僅 0.1% 使用者每天還在用 GPT-4o,API 不受影響。本文整理時程、決策理由,與 2025 年 8 月風波後「充分預告」承諾的兌現。
閱讀文章 ↗Deepgram 募得 1.3 億美元、估值 13 億,收購 OfOne 進軍語音點餐
2026 年 1 月 13 日,語音 AI 公司 Deepgram 以 13 億美元估值完成 1.3 億美元 C 輪,AVP 領投,並收購 YC 培育的 OfOne。現金流已轉正的公司為何募資?語音 API 市場的競局又將如何變化?
閱讀文章 ↗
2025
2 篇文章OpenAI推出o3-pro:主打可靠度的旗艦推理模型
2025年6月10日OpenAI發表o3-pro,ChatGPT Pro與Team用戶當天可用並取代o1-pro,API同步上線,定價每百萬輸入token 20美元、輸出80美元。主打可靠性而非速度,暫不支援Canvas與臨時對話。
閱讀文章 ↗Windsurf 稱 Anthropic 緊縮 Claude 直接存取
2025年6月3日,AI 編程工具 Windsurf 執行長 Varun Mohan 表示,Anthropic 以不到五天的通知刪減其 Claude 3.x 模型的第一方容量,公司被迫轉向第三方運算,突顯模型供應商與應用層之間的權力失衡。
閱讀文章 ↗
2026
19 ARTICLESOne Request Shape for Speech, Images, and Video: What OpenRouter's API Consolidation Changes
OpenRouter is putting TTS, image, and video models behind one request shape, changing how builders wire multimodal calls.
READ POST ↗Seedance 2.5: Long Takes, Reference Inputs, and the Per-Second Bill
Seedance 2.5 trades 4K for 30-second takes and cheaper video-reference billing — what that changes for clip pipelines.
READ POST ↗Config-as-Code for LLM Calls: What OpenRouter Presets Change About Shipping
One named preset replaces scattered model, prompt, and routing settings across every app that calls it.
READ POST ↗Zero Data Retention: Enforcing Provider-Side Privacy on AI API Calls
ZDR is a routing control that stops AI providers from storing prompts and responses—but only if you enforce it per request.
READ POST ↗Muse Spark 1.1 and the Meta Model API: What Changes for Agent Builders
Meta's Muse Spark 1.1 adds a 1M-token context and agentic tooling, now in public preview via the Meta Model API.
READ POST ↗Zero Data Retention Holds: OpenAI Sees Signals, Not Content
Private Safety Processing scans abuse patterns across interactions while ZDR holds; customers keep content and keys, white paper due September 2026.
READ POST ↗DeepSeek V4-Flash 0731: Same Architecture, Sharper Agents
V4-Flash-0731 (July 31, 2026) keeps the same architecture and only redoes post-training — agent benchmarks far exceed V4-Pro-Preview, at 210 tok/s with MIT-licensed weights.
READ POST ↗OpenRouter Fusion: Multi-Model Answers Beat Single Models
OpenRouter Fusion sends one prompt to a model panel, then a judge fuses the answers. On DRACO a fused panel scored 69.0%, above every solo model, at a fraction of the cost.
READ POST ↗OpenRouter's $113M Series B Values Model Routing at $1.3B
OpenRouter raised a $113M Series B led by CapitalG at a reported $1.3B valuation, processing 100 trillion tokens monthly across 400+ models — why routing now matters.
READ POST ↗Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents
Anthropic ships Claude Opus 4.8 41 days after 4.7 at unchanged pricing, adding Dynamic Workflows with hundreds of subagents, effort control, and sharper uncertainty flagging.
READ POST ↗DeepSeek Makes Its 75% V4 Pro Price Cut Permanent
DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.
READ POST ↗Reflex Benchmark: Computer Use Costs 45x More Than APIs
Reflex tested one task two ways: vision computer use burned ~551k input tokens in ~17 minutes; structured APIs took ~12k in 20 seconds. The interface sets the cost.
READ POST ↗Microsoft Opens Its MAI Audio and Image Models in Foundry
On April 2, 2026, Microsoft put three first-party MAI models into Foundry preview: half-cost transcription, a minute of speech in under a second, and a top-three image model.
READ POST ↗Google Lyria 3 Pro: AI Music Moves From Clips to Full Tracks
Google DeepMind's Lyria 3 Pro generates full three-minute songs with verse-chorus structure, rolling into Gemini, Vids, Vertex AI, and ProducerAI with SynthID watermarking.
READ POST ↗Gemini Embedding 2: One Vector Space for Every Modality
Google's Gemini Embedding 2 hits Public Preview on March 10, 2026 — the first natively multimodal embedding model, mapping text, images, video, audio, and PDFs into one 3072-dim vector space.
READ POST ↗Gemini 3.1 Flash Lite: Fastest, Cheapest Gemini 3 Model
Google shipped Gemini 3.1 Flash Lite in preview on March 3: $0.25 per million input tokens, up to 363 tokens per second, built for high-volume workloads. Pricing, benchmarks, and guidance.
READ POST ↗Gemini 3.1 Pro Lands in Preview: Double the Reasoning at the Same Price
On February 19, 2026, Google released Gemini 3.1 Pro in preview: 2x reasoning versus Gemini 3 Pro at the same pricing, on the gemini-3.1-pro-preview endpoint. Price-performance becomes the axis.
READ POST ↗OpenAI Retires GPT-4o From ChatGPT: The 0.1% Problem
OpenAI announced January 29, 2026: GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini leave ChatGPT on February 13. Only 0.1% of users still pick GPT-4o daily, most run GPT-5.2, and the API is unaffected.
READ POST ↗Deepgram Raises $130M at $1.3B, Buys OfOne
Deepgram raised a $130M Series C at a $1.3B valuation led by AVP and bought YC-backed OfOne. A cashflow-positive voice AI platform with 1,300+ customers, in a voice market headed to $14-20B by 2030.
READ POST ↗
2025
2 ARTICLESOpenAI launches o3-pro for ChatGPT Pro and the API
On June 10, 2025, OpenAI shipped o3-pro to ChatGPT Pro and Team users plus the API, replacing o1-pro at $20 per million input and $80 per million output tokens.
READ POST ↗Windsurf says Anthropic cut its direct Claude access
On June 3, 2025, Windsurf said Anthropic cut its first-party Claude 3.x capacity with under five days notice, forcing third-party compute — a case study in model providers power over app builders.
READ POST ↗