2026
10 篇文章選影像模型,終於不用只看範例圖了
OpenRouter 推出 Visual Image Benchmarks,用七類挑戰性 Prompt 比較 39 個影像模型,補上文字基準之外的空白。產品工作者可以先用它篩選,再拿自己的內容驗證。
閱讀文章 ↗FLUX 3 登場:影片、圖像、聲音一個骨幹全包
Black Forest Labs 發表多模態基礎模型 FLUX 3:最長 20 秒含原生音訊的影片生成、多語言文字渲染,並推出以 FLUX 3 為骨幹的機器人動作模型 FLUX-mimic。
閱讀文章 ↗Google Images 25 週年:從文字搜尋到視覺生成
Google Images 迎來 25 週年,推出可瀏覽的首頁與 AI Overviews 影像生成功能。本文回顧視覺搜尋的演進,並探討對產品建構者的啟示。
閱讀文章 ↗Nano Banana 2 Lite 登場:4 秒生圖、單張 3.4 美分的高速影像路線
2026 年 6 月 30 日,Google 發表 Nano Banana 2 Lite:約 4 秒完成文字生圖、每張 0.034 美元,同場推出 Gemini Omni Flash 影片模型,用「快出圖、再動圖」的串接工作流把低成本高速生成推向主流。
閱讀文章 ↗Adobe 收購 Topaz Labs:裝置端 AI 影像增強補強 Firefly
Adobe 於 2026 年 6 月 25 日宣布收購影像增強公司 Topaz Labs,交易預計下半年完成。Topaz 的 Astra、Wonder 模型與 Neurostream 本地運算技術將整合進 Firefly 與創作套件,把大型視訊模型搬進消費級 GPU。
閱讀文章 ↗OpenRouter Unified Image API:統一請求之前,先讓程式看得懂模型差異
OpenRouter 推出專屬 Image API:單一請求格式接 30 多個模型,能力、參數與計價全部變成可查詢的 schema。本文整理 capability descriptors、三種計費單位、streaming 支援,以及官方的 migration 建議。
閱讀文章 ↗Midjourney Medical 全身超音波掃描儀:60 秒成像的豪賭與爭議
Midjourney 於 2026 年 6 月 18 日發表 Medical 全身超音波掃描儀:受測者浸入水槽,環形探頭 60 秒取滿全身資料,由 AI 重建成類 CT 影像。官方稱速度近百倍於 MRI,並引用「避免 30% 死亡」的說法,引發放射科醫師強烈質疑。
閱讀文章 ↗ChatGPT Images 2.0 的實用觀察:從精準控制到多語言排版
OpenAI 在 2026 年 4 月 21 日推出 ChatGPT Images 2.0,主打更精準的圖像控制、多語言文字排版與更真實的人物呈現。本文從產品建構者的角度整理重點與限制。
閱讀文章 ↗微軟 MAI 三模型上線 Foundry:轉錄、語音與影像
2026 年 4 月 2 日,微軟把自研的 MAI-Transcribe-1、MAI-Voice-1、MAI-Image-2 送上 Microsoft Foundry 公開預覽:轉錄 GPU 成本砍半、一秒生成一分鐘語音、影像模型首發衝上 Arena.ai 第三名。
閱讀文章 ↗Nano Banana 2 登場:Pro 級能力、Flash 級速度的圖像生成
2026 年 2 月 26 日,Google 推出 Nano Banana 2 圖像生成模型,主打 Pro 級能力、Flash 級速度,已在 Gemini API 開放預覽。同場還有 Lyria 3 音樂生成與 Ultra 用戶的 Deep Think。本文看這波二月更新的策略。
閱讀文章 ↗
2025
1 篇文章2026
9 ARTICLESOpenRouter's Image Benchmarks: A Practical Guide for Product Builders
OpenRouter launches visual image benchmarks to help developers compare 39 image models across 7 challenge categories, with practical advice for product builders.
READ POST ↗FLUX 3: One Backbone for Video, Images, Audio — and Robots
Black Forest Labs launched FLUX 3, a multimodal model with 20-second native-audio video, plus FLUX-mimic, a robot action model built on the same backbone.
READ POST ↗Google Images at 25: From Text Links to AI Image Generation in Search
Google Images turns 25 with a new browseable homepage and AI image generation in AI Overviews. Here's what product builders need to know.
READ POST ↗Google's Nano Banana 2 Lite: 4-Second, 3.4-Cent Images
On June 30, 2026, Google launched Nano Banana 2 Lite — images in about 4 seconds at $0.034 each — plus the Gemini Omni Flash video model, pairing fast stills with cheap video.
READ POST ↗Adobe Acquires Topaz Labs to Bring On-Device AI to Firefly
Adobe is acquiring Topaz Labs, whose Astra and Wonder models and Neurostream on-device tech will join Firefly and the creative apps; the deal closes in H2 2026.
READ POST ↗Midjourney Medical: 60-Second Full-Body Ultrasound Gamble
Midjourney's Medical scanner dips users in a water tank, fires a ring of transducers for 60 seconds, and AI-reconstructs CT-like images. Radiologists push back on health claims.
READ POST ↗ChatGPT Images 2.0: What Builders Should Know About Text, Multilingual Layouts, and Realism
OpenAI's ChatGPT Images 2.0 brings better text rendering, multilingual support, and realistic styles. Here's what product builders need to know.
READ POST ↗Microsoft Opens Its MAI Audio and Image Models in Foundry
On April 2, 2026, Microsoft put three first-party MAI models into Foundry preview: half-cost transcription, a minute of speech in under a second, and a top-three image model.
READ POST ↗Nano Banana 2 Arrives: Pro-tier Image Generation at Flash-tier Speed
Google launched Nano Banana 2 on Feb 26, 2026: Pro-tier image generation at Flash-tier speed, in preview in the Gemini API, alongside Lyria 3 music generation and Deep Think for Ultra users.
READ POST ↗