2026
44 篇文章把翻譯模型放進產品前,先看吞吐量與長文件的真實落差
North Small Translate 以 25B 活躍參數在 WMT26 拿下 83.6 分,並在長文件與吞吐量上拉開差距,改變了翻譯功能的建置取捨。
閱讀文章 ↗Mistral 的 30 億歐元賭注:開放權重如何成為主權 AI 的技術前沿
Mistral 完成歐洲科技業最大規模融資,以開放權重模型、基礎設施與產品全端布局,回應企業與政府對主權 AI 的需求。本文從產品建造者角度,解析這輪融資對技術選擇與部署策略的意涵。
閱讀文章 ↗Nvidia 洽購 Hugging Face:逾 130 億美元的開源樞紐爭奪
《Business Insider》8 月 27 日報導,Nvidia 正洽談以超過 130 億美元估值收購 Hugging Face,交易尚未敲定。本文解析交易動機、Hugging Face 的談判籌碼與開源社群的疑慮。
閱讀文章 ↗GLM-5.3 權重開放下載:753B 旗艦的本地部署現實
2026 年 8 月 25 日,Z.ai 把 753B 參數的 GLM-5.3 權重放上 Hugging Face,三天後在 Hacker News 衝上 806 分。本文解析其 MoE 架構、自訂授權條款與本地部署實測。
閱讀文章 ↗Qwen3.8-Flash-Next 發表:新架構把啟用參數壓到 6B
2026 年 8 月 26 日,Qwen 開源 Qwen3.8-Flash-Next:125B 主模型加 51B n-gram 嵌入表,每個 token 只啟用 6B 參數,原生 262K 上下文,官方定位是極致成本效率。
閱讀文章 ↗Qwen3.8-27B 開源釋出:體積小、看得懂圖,只是預設想太多
阿里巴巴 Qwen 兌現一週前承諾,釋出 Apache 2.0 的 Qwen3.8-27B 開放權重:262K 上下文、原生視覺、單機可跑;Simon Willison 實測稱讚能力,也點出 xhigh 推理預設讓簡單任務慢上數倍。
閱讀文章 ↗Z.ai 發表 GLM-5.3:開源程式碼新高,網路攻擊能力超預期
2026 年 8 月 14 日,Z.ai 發表 GLM-5.3:沿用 GLM-5.2 基座、靠後訓練把程式碼基準推上開源新高,並主動披露網路攻擊能力「發展得比我們預期更快」,授權同步轉為自訂條款。
閱讀文章 ↗美國能源部啟動 Genesis 開放模型計畫:開放權重做科學
美國能源部在 Genesis Mission 之下啟動開放權重科學模型計畫,與 Arcee AI 合作開發首個模型 Genesis-Science-1,並開放全國貢獻門戶,首輪申請 8 月 14 日截止。
閱讀文章 ↗Meta 回歸開放權重:30B 的 Muse Glimmer 把本地代理變成現實
2026 年 8 月 10 日,Meta 以 Apache 2.0 釋出 30B 參數的 Muse Glimmer,針對常駐本地代理工作流設計,單張消費級 GPU 就能跑,並承諾開放 Muse Spark 權重。
閱讀文章 ↗Mistral 開源 Shieldstral:把審查政策變成一句提問的 3B 分類器
Mistral 於 8 月 4 日發表 Shieldstral:3B 開源權重多模態安全分類器,政策以自然語言在推論時下達,單張 16GB GPU 即可運行,以 Apache 2.0 釋出。
閱讀文章 ↗MiniMax H3 開源:一次生成 2K 影片與立體聲的全模態模型
MiniMax 於 7 月 31 日發表 Hailuo 後繼者 H3,數日後釋出權重:33B 全模態 Transformer 同時生成最長 15 秒、768p 起步 2K 的影片與原生立體聲,ComfyUI Day-0 支援,量化後 RTX 3060 也能本地運行。
閱讀文章 ↗Qwen3.8-Max 正式發布並開放權重:鎖定寫程式與代理協作
2026 年 8 月 3 日,阿里巴巴正式發布 Qwen3.8-Max:百萬 token 上下文、多模態輸入、輸入每百萬 token 2 美元,並成為 Max 系列首個開放權重的模型,主打寫程式與代理協作。
閱讀文章 ↗DeepSeek V4-Flash 0731:只重做後訓練,代理能力越級跳
2026 年 7 月 31 日,DeepSeek 推出 V4-Flash-0731:架構與規模不變、僅重新後訓練,代理基準大幅超越 V4-Pro-Preview,輸出速度每秒 210 個 token,權重續以 MIT 授權開放。
閱讀文章 ↗Dario Amodei 表態:Anthropic 從未主張禁用開放權重模型
美國官員傳出考慮禁用中國開放權重模型、科技業連署力挺開放權重之際,Amodei 於 7 月 27 日撰文澄清 Anthropic 從未主張禁令,改提出晶片出口管制、打擊工業級蒸餾與強制安全測試三項主張。
閱讀文章 ↗FLUX 3 登場:影片、圖像、聲音一個骨幹全包
Black Forest Labs 發表多模態基礎模型 FLUX 3:最長 20 秒含原生音訊的影片生成、多語言文字渲染,並推出以 FLUX 3 為骨幹的機器人動作模型 FLUX-mimic。
閱讀文章 ↗輝達、微軟、Meta 公開信:別過早設限開放權重模型
輝達、微軟、Meta、Palantir 等逾 20 家公司連署公開信,反對對開放權重 AI 模型的過早限制;同一週白宮顧問指控 Kimi K3 蒸餾 Anthropic 技術,政策攻防升溫。
閱讀文章 ↗近 200 家新創連署:別封殺中國開源權重模型
Little Tech Association 在 7 月 22 日代表近 200 家公司致函川普政府,反對限制中國開源權重模型;新創擔心成本暴增、公司倒閉,OpenAI 與 Anthropic 則在華府聯手警告開源權重風險,兩條陣線正面交鋒。
閱讀文章 ↗Kimi K3 登場:2.8T 參數、百萬 Token Context,開源模型走向長程 Agent
Moonshot 發布 Kimi K3:2.8T 參數的開源 3T 級模型,原生視覺、1M token context,主打長程 coding 與知識工作。本文整理 Delta Attention 架構、kernel 編譯器與晶片設計等案例、API 定價與 64 卡部署建議,以及官方自認仍落後 Fable 5 與 GPT 5.6 Sol 的誠實定位。
閱讀文章 ↗LM Studio 推出 Bionic:開源模型專用的本機 AI Agent
LM Studio 團隊於 2026 年 7 月 16 日推出獨立新應用 Bionic,定位為開源模型的生產力 Agent,涵蓋寫程式、文件處理與本機語音輸入,並提供本機、LM Link 與 Secure Cloud 三種執行方式。
閱讀文章 ↗Ai2 拆解混合模型優勢:贏在內容詞,輸在複製
Ai2 用資料與訓練配方完全對齊的 Olmo 3 與 Olmo Hybrid 逐 token 比較損失,發現混合架構在內容詞上明顯佔優,但在逐字複製與右括號上優勢消失,顯示循環層擅長語意狀態追蹤、注意力擅長複製。
閱讀文章 ↗GLM-5.2 開源發布:百萬上下文長任務表現緊咬 Opus 4.8
Z.AI 於 2026 年 6 月 17 日開源 753B 參數的 GLM-5.2:MIT 授權、穩定 1M token 上下文,長任務與程式基準緊追 Claude Opus 4.8,成為開源陣營排名最高的模型。
閱讀文章 ↗Google 開源 DiffusionGemma:擴散式生成快 4 倍,單卡 H100 破千 token/秒
Google 於 6 月 10 日開源 DiffusionGemma:26B-A4B 擴散語言模型,Apache 2.0 釋出,單卡 H100 每秒生成逾千 token、比自回歸快 4 倍,但多數基準仍輸 Gemma 4,官方明言最適合本地與低併發場景。
閱讀文章 ↗Liquid AI LFM2.5-8B-A1B 發布:38T tokens 訓練的端側 MoE 推理模型
Liquid AI 於 5 月 28 日發布 LFM2.5-8B-A1B:8B 總參數、每 token 約 1B 啟用的端側 MoE 推理模型,預訓練 38T tokens、128K 上下文,MATH500 達 88.76,手機上每秒約 30 tokens。
閱讀文章 ↗DeepSeek 把 V4 Pro 七五折降價常態化:前沿模型價格戰的底牌
2026 年 5 月 22 日,DeepSeek 宣布旗艦 V4 Pro 的 75% 折扣轉為常態價:輸出每百萬 token 0.87 美元,僅為原價四分之一,Flash 更只要 0.28 美元。本文解析降價常態化的成本邏輯與對 API 市場的連鎖影響。
閱讀文章 ↗Stable Audio 3.0 發表:開放權重、6 分 20 秒完整作曲
2026 年 5 月 20 日 Stability AI 發表 Stable Audio 3.0 四模型家族:三款開放權重、最長 6 分 20 秒完整作曲,Small 可在手機端生成 2 分鐘音樂,全家族以完全授權資料訓練,正面挑戰纏訟中的 Suno 與 Udio。
閱讀文章 ↗DeepSeek 首度對外募資:估值傳從 200 億美元跳到 450 億
DeepSeek 傳出成立以來首次外部融資:國家集成電路產業投資基金據傳領投,估值幾週內從 200 億美元升至 450 億,騰訊與阿里巴巴洽談參投,募資首要目的是留住被挖角的研究人才。
閱讀文章 ↗IBM Granite 4.1 登場:8B 稠密模型追平 32B MoE 的開源算盤
IBM 於 2026 年 4 月 29 日發布 Granite 4.1:3B/8B/30B 全稠密模型以 Apache 2.0 開源,8B 追平上一代 32B MoE,並刻意拿掉推理模式換取可預測延遲。本文拆解 15T tokens 預訓練與四階段 RL 的工程細節。
閱讀文章 ↗Qwen3.6-35B-A3B 開源釋出:3B 啟動參數的代理編碼模型
阿里巴巴 4 月 16 日以 Apache-2.0 開源 Qwen3.6-35B-A3B:35B 總參數僅啟動 3B 的 MoE,原生 262K 上下文,Terminal-Bench 2.0 達 51.5;Simon Willison 筆電實測在 SVG 任務勝過同日發布的 Claude Opus 4.7。
閱讀文章 ↗Ai2 開源 WildDet3D:用單張照片預測 3D 偵測框
Ai2 於 4 月 7 日開源 WildDet3D,從單張 RGB 影像預測物體的 3D 偵測框,支援文字、點擊與 2D 框提示,並同步發布超過 1 百萬張影像、370 萬組驗證標註的資料集,零樣本成績大幅超越先前方法。
閱讀文章 ↗Liquid AI 推出 LFM2.5-VL-450M:跑得進手機的 450M 視覺語言模型
Liquid AI 於 4 月 8 日發布 LFM2.5-VL-450M,預訓練從 10T 擴到 28T tokens,新增物件偵測與函式呼叫能力,量化後在 Jetson Orin 上每幀僅 233 毫秒,瞄準邊緣裝置上的多模態應用。
閱讀文章 ↗Gemma 4 開源發布:Apache 2.0、MoE 與 256K 上下文
2026 年 4 月 2 日 Google DeepMind 發布 Gemma 4:全系列改用 Apache 2.0 授權,四種尺寸從 2B 端側到 31B Dense,支援 140+ 語言、256K 上下文與影像語音輸入,31B 躍上 Arena 開源模型第三名。
閱讀文章 ↗PrismML 推出 1-bit Bonsai:把 8B 模型壓進 1.15 GB 的端側 LLM
2026 年 3 月 31 日,Caltech 衍生新創 PrismML 走出 stealth,發表號稱首款可商業化的端到端 1-bit LLM 家族 Bonsai:8B 旗艦僅 1.15 GB、比 16 位元同級小 14 倍,以 Apache 2.0 開源上架 Hugging Face。
閱讀文章 ↗USCC 報告:中國開源 AI 的「雙循環」正在強化工業主導地位
2026 年 3 月 23 日,美中經濟與安全審查委員會發布《Two Loops》研究報告:中國全面押注開源 AI,透過數位與實體兩個互相強化的循環鞏固工業主導地位。Qwen 衍生模型已超過十萬個,而美國出口管制主要只打到數位循環。
閱讀文章 ↗樂天開源 Rakuten AI 3.0:6,710 億參數、GENIAC 打造的日本最大模型
2026 年 3 月 17 日,樂天以 Apache 2.0 釋出約 6,710 億參數的 MoE 模型 Rakuten AI 3.0——日本最大高效能 AI 模型,訓練獲經產省與 NEDO 的 GENIAC 計畫支持。本文解析其日語基準成績、每 token 約 400 億活躍參數的架構,以及開源 671B 的現實限制。
閱讀文章 ↗NVIDIA 開源 Nemotron 3 Super:120B 混合 MoE 模型瞄準 Agent 推論吞吐
NVIDIA 於 2026 年 3 月 11 日開源 Nemotron-3-Super-120B-A12B:混合 Mamba 與 Latent MoE 架構、1M token 上下文、每秒 478 token 輸出,連同訓練資料集與 RL 環境一併釋出,瞄準 Agent 工作負載的推論成本。
閱讀文章 ↗Qwen3.5 小模型補齊戰線:9B 在多項基準超越 gpt-oss-120b
阿里巴巴 2 月底補上 Qwen3.5 的 0.8B–9B 四款小模型:Apache-2.0 開源、原生 262K 情境、支援影像影片輸入。9B 以 9.65B 密集參數在 MMLU-Pro、GPQA Diamond 超越 gpt-oss-120b,目標是手機與筆電上的本地多模態。
閱讀文章 ↗Perplexity 開源 pplx-embed:擴散預訓練打造網頁級檢索嵌入模型
Perplexity 於 2 月 26 日發布 pplx-embed-v1 與 pplx-embed-context-v1,各提供 0.6B 與 4B 開放權重,以擴散預訓練與雙向注意力主打網頁級檢索。本文解析架構、基準成績與成本工程。
閱讀文章 ↗Cohere Tiny Aya:33.5 億參數、70+ 語言的本機開源模型
Cohere Labs 於 2026 年 2 月 17 日釋出 Tiny Aya 開源多語模型:33.5 億參數、支援 70+ 語言、可在一般硬體本機運行,並選在印度 AI Impact Summit 期間登場。本文解析架構、部署方式與 CC-BY-NC 授權限制。
閱讀文章 ↗智利領軍推出 Latam-GPT:拉美首個本土開源大模型
2026 年 2 月 10 日,智利總統博里奇為 Latam-GPT 揭幕。這個由 CENIA 主導、集結 15 國力量的 700 億參數開源模型,以 8 TB 區域資料訓練,要對抗美國中心偏見。本文解析其規格、資金現實與主權意涵。
閱讀文章 ↗MiniMax 開源 M2.5 與 Lightning:SWE-Bench 80.2%,十分之一價格逼近 Opus 4.6
2026 年 2 月 12 日 MiniMax 發布 M2.5 與 M2.5-Lightning 並開源 229B 權重:SWE-Bench Verified 80.2%,解題速度與 Claude Opus 4.6 持平但成本約十分之一,輸出單價最低每百萬 token 1.2 美元。
閱讀文章 ↗Microsoft 新掃描法:不必知道觸發詞,也能抓出 LLM 裡的臥底後門
微軟研究團隊發表 The Trigger in the Haystack:利用聊天模板讓中毒模型自行洩漏後門訓練資料,再以注意力分析重建觸發詞,在 47 個臥底模型上達約 88% 偵測率、13 個良性模型零誤報,為開源模型上線前稽核提供新工具。
閱讀文章 ↗NVIDIA Earth-2 開放氣象模型全家桶:整條 AI 天氣預報管線開源
2026 年 1 月 26 日,NVIDIA 在美國氣象學會年會發表 Earth-2 開放模型家族:Atlas 做 15 天中期預報、StormScope 做 0 至 6 小時臨近預報、HealDA 用 GPU 秒級生成初始場,稱為首個全開源的加速氣象 AI 堆疊。
閱讀文章 ↗智譜港股掛牌:中國首家純大模型公司上市,首日收漲 13.2%
2026 年 1 月 8 日,智譜(Zhipu AI,2513.HK)在港交所掛牌,成為中國首家上市的純大模型開發商。IPO 募資 43.5 億港元、估值近 510 億港元,首日收漲 13.2%,MiniMax 隔日跟進掛牌,中國 AI 上市潮正式啟動。
閱讀文章 ↗Epoch AI:中國模型平均落後美國前沿 7 個月,差距 4 到 14 個月
Epoch AI 於 2026 年 1 月 2 日發布 ECI 分析:2023 年以來中國模型平均落後美國前沿 7 個月,區間 4 至 14 個月,尚無模型超越 o3。本文拆解計算方法、開源權重的干擾變數與社群爭論。
閱讀文章 ↗
2025
2 篇文章Mistral發表首個推理模型Magistral
2025年6月10日法國AI公司Mistral發表首個推理模型系列Magistral:24B開源的Small(Apache 2.0授權)與企業版Medium;Medium在AIME2024拿下73.6%,並主打可用使用者語言追蹤的透明推理過程。
閱讀文章 ↗Hugging Face 釋出 450M 開源機器人模型 SmolVLA
2025年6月初,Hugging Face 開源 450M 參數的視覺-語言-動作模型 SmolVLA,可跑在 MacBook 或單一消費級 GPU 上,並以非同步推論堆疊把任務完成時間縮短約30%,為低成本開源機器人生態提供基礎模型。
閱讀文章 ↗
2026
44 ARTICLESNorth Small Translate: What 16k Context and 1.4x Throughput Change for Translation Pipelines
Cohere's open-weight translation model pairs 16k context with 1.4x throughput, shifting how builders handle long documents.
READ POST ↗What a €3B Sovereign AI Bet Means for Builders
Mistral's Series D signals a shift from raw model power to control over data, models, compute, and production systems. Here's what that means for teams choosing AI infrastructure.
READ POST ↗Nvidia in Talks to Buy Hugging Face for Over $13 Billion
Nvidia is in talks to acquire Hugging Face at a valuation above $13 billion, but no deal is finalized. The motives, the leverage, and the open-source community's worries.
READ POST ↗GLM-5.3 Weights Are Out: Running a 753B MoE Model Locally
Z.ai put the 753B-parameter GLM-5.3 weights on Hugging Face on August 25, and the Hacker News thread hit 806 points. The architecture, the custom license, and local-run realities.
READ POST ↗Qwen3.8-Flash-Next: 125B MoE with only 6B active parameters
Qwen open-sourced Qwen3.8-Flash-Next: a 125B main model plus a 51B n-gram embedding table, only 6B parameters active per token, 262K context, built for cost efficiency.
READ POST ↗Qwen3.8-27B: Great Open Weights That Overthink by Default
Alibaba's Qwen shipped the Apache 2.0 Qwen3.8-27B a week after promising it: 262K context, native vision, laptop-class local AI — behind a costly xhigh reasoning default.
READ POST ↗GLM-5.3: Open-Weight Coding Frontier With Sharp Cyber Gains
GLM-5.3 reuses the GLM-5.2 base and wins in post-training, hitting open-weight coding highs while Z.ai flags that its cyber capability 'developed faster than we expected.'
READ POST ↗DOE Launches Genesis Open Models for Open-Weight Science
The U.S. Department of Energy opened a Genesis open-weight models program under its Genesis Mission, partnering with Arcee AI on the first model, Genesis-Science-1.
READ POST ↗Meta's Muse Glimmer: a 30B Open Model for Local Agents
Meta released Muse Glimmer on August 10, 2026: a 30B open-weight model under Apache 2.0, distilled for always-on local agent workflows on a single consumer GPU.
READ POST ↗Mistral's Shieldstral: A 3B Open-Weights Safety Classifier
Mistral's Shieldstral is a 3B open-weights multimodal safety classifier that takes policies as plain-language questions at inference time and runs on one 16GB GPU.
READ POST ↗MiniMax H3 Goes Open With 2K Video and Native Stereo Audio
MiniMax announced H3 on July 31 and released the weights days later: a 33B omni-modal Transformer that generates up to 15 seconds of 2K video with native stereo audio.
READ POST ↗Qwen3.8-Max Goes GA: 1M Context and Open Weights
Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.
READ POST ↗DeepSeek V4-Flash 0731: Same Architecture, Sharper Agents
V4-Flash-0731 (July 31, 2026) keeps the same architecture and only redoes post-training — agent benchmarks far exceed V4-Pro-Preview, at 210 tok/s with MIT-licensed weights.
READ POST ↗Amodei: Anthropic Never Sought an Open-Weights Ban
Amodei says Anthropic never advocated a ban on open-weights models, and instead backs chip export controls, a distillation crackdown, and mandatory safety testing.
READ POST ↗FLUX 3: One Backbone for Video, Images, Audio — and Robots
Black Forest Labs launched FLUX 3, a multimodal model with 20-second native-audio video, plus FLUX-mimic, a robot action model built on the same backbone.
READ POST ↗Nvidia, Microsoft, Meta Warn Against Open-Weight Curbs
Nvidia, Microsoft, Meta, and Palantir signed a letter against premature curbs on open-weight AI models, the same week a White House advisor accused Kimi K3 of distillation.
READ POST ↗Startups Ask Trump Not to Cut Off Chinese Open-Weight AI
Nearly 200 startups urged the White House not to block Chinese open-weight models, while OpenAI and Anthropic warn Washington about open-weight risks. Two camps, one policy fight.
READ POST ↗Kimi K3 Arrives: 2.8T Parameters, Million-Token Context, Open Models Go Long-Horizon
Kimi K3: a 2.8T open 3T-class model with native vision and 1M-token context for long-horizon coding. The architecture, kernel and chip cases, pricing, and the gap to Fable 5 and GPT 5.6 Sol.
READ POST ↗LM Studio Bionic: A Local-First AI Agent for Open Models
LM Studio shipped Bionic on July 16, 2026 — a separate agent app for open models covering coding, documents, and local voice input, runnable locally or via Secure Cloud.
READ POST ↗Ai2 Maps Where Hybrid LLMs Beat Transformers, Token by Token
Ai2 compared Olmo 3 with Olmo Hybrid token by token, with matched data and training. Hybrids win on content words; the edge vanishes on verbatim copying and closing brackets.
READ POST ↗GLM-5.2: Open Weights, 1M Context, Long-Horizon Gains
Z.AI open-sources GLM-5.2, a 753B MIT-licensed model with a stable 1M-token context that trails Claude Opus 4.8 by about one percent on long-horizon coding benchmarks.
READ POST ↗Google Open-Sources DiffusionGemma: 4x Faster Generation
Google open-sources DiffusionGemma: an Apache 2.0 26B-A4B diffusion LLM hitting 1000+ tokens/sec on one H100 — 4x faster than autoregressive, at a quality cost vs Gemma 4.
READ POST ↗Liquid AI LFM2.5-8B-A1B: On-Device MoE Reasoning Model
Liquid AI released LFM2.5-8B-A1B: an 8B-total, ~1B-active MoE reasoning model trained on 38T tokens with 128K context and 253 tok/s CPU inference. Weights are on Hugging Face.
READ POST ↗DeepSeek Makes Its 75% V4 Pro Price Cut Permanent
DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.
READ POST ↗Stable Audio 3.0: Open-Weight Models, Six-Minute Songs
Stability AI ships Stable Audio 3.0 — four models, three open-weight, 6:20 tracks, 2-minute on-device music, all trained on licensed data while Suno and Udio fight in court.
READ POST ↗DeepSeek's First Outside Round: Valuation Talk Hits $45B
DeepSeek is negotiating its first-ever outside round per FT and Bloomberg — a state chip fund to lead, Tencent and Alibaba in talks, valuation up from $20B to $45B.
READ POST ↗IBM Granite 4.1: An 8B Dense Model Matching a 32B MoE
IBM's Granite 4.1 ships 3B/8B/30B dense models under Apache 2.0; the 8B matches the old 32B MoE, reasoning is deliberately removed, and context reaches 512K.
READ POST ↗Qwen3.6-35B-A3B: Open-Weight MoE Punches at Agentic Coding
Alibaba open-sources Qwen3.6-35B-A3B (Apache-2.0): a 35B MoE with 3B active params, 262K context, Terminal-Bench 2.0 at 51.5 — beating Opus 4.7 on a same-day laptop SVG test.
READ POST ↗Ai2 WildDet3D: Open 3D Detection from a Single Photo
Ai2 open-sourced WildDet3D on April 7: 3D bounding boxes from a single RGB image, with text, click, and box prompts, plus a 1M-image dataset holding 3.7M verified 3D annotations.
READ POST ↗Liquid AI Ships LFM2.5-VL-450M, a 450M Edge VLM
Liquid AI released LFM2.5-VL-450M on April 8: 28T-token pretraining, new detection and function-calling skills, and 233 ms per frame on a Jetson Orin once quantized.
READ POST ↗Gemma 4 Ships Under Apache 2.0: Google's Open Model Reset
Google DeepMind released Gemma 4 on April 2, 2026 under Apache 2.0 — four sizes from 2B edge to 31B dense, 140+ languages, 256K context, Arena top-3. What it changes for builders.
READ POST ↗PrismML 1-bit Bonsai: An 8B LLM Squeezed Into 1.15 GB
Caltech spinoff PrismML emerged from stealth with 1-bit Bonsai, an end-to-end 1-bit LLM family: the 8B flagship is 1.15 GB, 14x smaller than 16-bit peers, Apache 2.0 licensed.
READ POST ↗USCC: China's Open-Source AI Reinforces Industrial Power
A new USCC report argues China's open-source AI strategy is self-reinforcing: open models spread globally while factory deployment feeds data back. Qwen alone has 100,000+ derivatives.
READ POST ↗Rakuten AI 3.0: Japan's Largest Model Goes Open-Weight
Rakuten open-sources Rakuten AI 3.0 under Apache 2.0: about 671B total parameters, 40B active per token, GENIAC-backed and tuned for Japanese. Benchmarks and the reality of running it.
READ POST ↗NVIDIA Open-Sources Nemotron 3 Super, a 120B MoE for Agents
On March 11, 2026, NVIDIA open-sourced Nemotron-3-Super-120B-A12B: hybrid Mamba plus latent MoE, 1M context, 478 tokens/sec, and 10T+ tokens of training data aimed at agentic inference costs.
READ POST ↗Qwen3.5 Small Models: 9B Rivals gpt-oss-120b at the Edge
Alibaba adds 0.8B-9B small models to Qwen3.5 under Apache-2.0: native 262K context, image and video input. The 9B beats gpt-oss-120b on MMLU-Pro and GPQA Diamond, targeting phones and laptops.
READ POST ↗Perplexity Open-Sources pplx-embed Retrieval Models
Perplexity released pplx-embed-v1 and pplx-embed-context-v1 on Feb 26 — open weights at 0.6B and 4B, diffusion-pretrained and bidirectional, built for web-scale retrieval and RAG pipelines.
READ POST ↗Cohere's Tiny Aya: A 3.35B Open Model for 70+ Languages
Cohere Labs released Tiny Aya: a 3.35B open-weight model covering 70+ languages, built to run on a laptop, debuted at the India AI Impact Summit. Architecture, deployment, and license, broken down.
READ POST ↗Latam-GPT: Latin America's First Homegrown Open-Source LLM
Chile launched Latam-GPT on Feb 10: a 70B open-source model from CENIA and 15 partner countries, trained on 8 TB of regional data to counter US-centric bias.
READ POST ↗MiniMax M2.5: Open Coding Weights at a Tenth of Opus Cost
MiniMax open-sources M2.5 and M2.5-Lightning: 80.2% SWE-Bench Verified, Opus-4.6-class task times at roughly a tenth of the cost, 229B parameters under modified MIT on Hugging Face.
READ POST ↗Microsoft's New Scan Catches Sleeper-Agent LLM Backdoors
Microsoft's 'The Trigger in the Haystack' pulls sleeper-agent backdoor triggers out of poisoned LLMs: ~88% detection across 47 models, zero false positives on 13 benign ones.
READ POST ↗NVIDIA Earth-2: First Fully Open AI Weather Stack
NVIDIA open-sourced the Earth-2 weather AI stack — Atlas 15-day forecasts, StormScope storm nowcasting, HealDA GPU-fast data assimilation — with national weather agencies already running it.
READ POST ↗Zhipu AI Lists in Hong Kong: China's First Pure-Play LLM IPO
Zhipu AI (2513.HK) debuted in Hong Kong on Jan 8, 2026 as China's first listed pure-play LLM maker, raising HK$4.35B at a valuation near HK$51B and closing up 13.2% on day one.
READ POST ↗Epoch AI: Chinese Models Trail US Frontier by Seven Months
Epoch AI's ECI index puts Chinese models seven months behind the US frontier on average since 2023, ranging four to fourteen months. How it is measured, and the open-weight catch.
READ POST ↗
2025
2 ARTICLESMistral enters reasoning AI with Magistral Small, Medium
Mistral launched its first reasoning models on June 10, 2025: open-weights 24B Magistral Small under Apache 2.0 and enterprise Magistral Medium, which scores 73.6% on AIME2024.
READ POST ↗Hugging Face ships SmolVLA, a lean open robotics model
In early June 2025, Hugging Face open-sourced SmolVLA, a 450M-parameter vision-language-action model that runs on a MacBook and cuts task time about 30% — a base for low-cost open robotics.
READ POST ↗