2026
24 篇文章Hermes Agent:記憶留在本機、會自己長出技能的開源助理
Nous Research 的開源 Hermes Agent 把記憶留在本機、會自動生成技能,2026 年 2 月推出後登上 OpenRouter 累積用量第一。本文整理它的設計賭注、Desktop 公測與 15 億美元估值融資,並看用量數字該怎麼讀。
閱讀文章 ↗GLM-5.3 權重開放下載:753B 旗艦的本地部署現實
2026 年 8 月 25 日,Z.ai 把 753B 參數的 GLM-5.3 權重放上 Hugging Face,三天後在 Hacker News 衝上 806 分。本文解析其 MoE 架構、自訂授權條款與本地部署實測。
閱讀文章 ↗NVIDIA 發表 Jetson Orin Nano 2:入門邊緣 AI 算力翻倍,機器人與無人機受惠
2026 年 8 月 25 日,NVIDIA 發表 Jetson Orin Nano 2 機器人運算系統,將入門級邊緣 AI 算力約提升一倍,鎖定機器人與無人機應用。本文解析邊緣推理的實際意義與對開發者生態的影響。
閱讀文章 ↗蘋果 M6 與 M5 Ultra 登場:2nm 與四晶粒推向端側 AI
蘋果於 8 月 25 日發表首款 2nm 晶片 M6 與首款四晶粒架構的 M5 Ultra,Mac mini 與 Mac Studio 同步更新,最高 512GB 統一記憶體讓數千億參數模型能在桌面本機執行。
閱讀文章 ↗Qwen3.8-27B 開源釋出:體積小、看得懂圖,只是預設想太多
阿里巴巴 Qwen 兌現一週前承諾,釋出 Apache 2.0 的 Qwen3.8-27B 開放權重:262K 上下文、原生視覺、單機可跑;Simon Willison 實測稱讚能力,也點出 xhigh 推理預設讓簡單任務慢上數倍。
閱讀文章 ↗Meta 回歸開放權重:30B 的 Muse Glimmer 把本地代理變成現實
2026 年 8 月 10 日,Meta 以 Apache 2.0 釋出 30B 參數的 Muse Glimmer,針對常駐本地代理工作流設計,單張消費級 GPU 就能跑,並承諾開放 Muse Spark 權重。
閱讀文章 ↗在 8GB Mac 上跑 26B 模型:TurboFieldfare 的 SSD 串流推理
開源專案 TurboFieldfare 以 Swift 與 Metal 打造推理引擎:只常駐約 2GB 記憶體,把 Gemma 4 26B-A4B 的 MoE 專家權重從 SSD 逐 token 串流載入,在 8GB MacBook Air 上跑出每秒 5 到 6 個 token。
閱讀文章 ↗LM Studio 推出 Bionic:開源模型專用的本機 AI Agent
LM Studio 團隊於 2026 年 7 月 16 日推出獨立新應用 Bionic,定位為開源模型的生產力 Agent,涵蓋寫程式、文件處理與本機語音輸入,並提供本機、LM Link 與 Secure Cloud 三種執行方式。
閱讀文章 ↗Snap Specs 開賣:2,195 美元的 on-device AR 眼鏡豪賭
2026 年 6 月 16 日,Snap 正式發表 Specs AR 眼鏡:2,195 美元、雙 Snapdragon 處理器、51 度視野、運算全在眼鏡上完成。本文整理規格與情境式 AI 功能,並看發表後股價下跌逾 5% 的市場反應。
閱讀文章 ↗Ollama 0.30 更新:GGUF 原生支援、NVIDIA 最高 20% 效能提升、Vulkan 預設啟用
Ollama 0.30 透過 llama.cpp 深化 GGUF 引擎:NVIDIA 吞吐最高提升 20%(RTX 5090 實測條件)、Vulkan 預設開啟讓 AMD 與 Intel 開箱即用、LFM 與 Prism 家族及 Unsloth 微調可直接執行,tool calling 能力沿用並可掛上 coding agent。
閱讀文章 ↗Liquid AI LFM2.5-8B-A1B 發布:38T tokens 訓練的端側 MoE 推理模型
Liquid AI 於 5 月 28 日發布 LFM2.5-8B-A1B:8B 總參數、每 token 約 1B 啟用的端側 MoE 推理模型,預訓練 38T tokens、128K 上下文,MATH500 達 88.76,手機上每秒約 30 tokens。
閱讀文章 ↗IrisGo 桌面 AI 管家:看一次就學會的主動式代理人
2026 年 5 月 20 日報導:IrisGo 獲 Andrew Ng 的 AI Fund 領投 280 萬美元種子輪,推出「示範一次就學會」的桌面 AI 管家,本機 NPU 推理、雲端需明示授權,並與 Acer 簽下新機預裝協議,賭 AI PC 時代的系統級入口。
閱讀文章 ↗Redis 之父 antirez 開源 ds4:跑得動前沿模型的本地推理引擎
Redis 作者 antirez 用一週寫出 ds4(DwarfStar):MIT 授權、C 語言的本地推理引擎,以 2/8-bit 非對稱量化讓 DeepSeek V4 Flash 與 GLM 模型跑進 96–128GB 的 Mac,內建原生編碼代理,登上 HN 首頁。
閱讀文章 ↗Qwen3.6-35B-A3B 開源釋出:3B 啟動參數的代理編碼模型
阿里巴巴 4 月 16 日以 Apache-2.0 開源 Qwen3.6-35B-A3B:35B 總參數僅啟動 3B 的 MoE,原生 262K 上下文,Terminal-Bench 2.0 達 51.5;Simon Willison 筆電實測在 SVG 任務勝過同日發布的 Claude Opus 4.7。
閱讀文章 ↗Liquid AI 推出 LFM2.5-VL-450M:跑得進手機的 450M 視覺語言模型
Liquid AI 於 4 月 8 日發布 LFM2.5-VL-450M,預訓練從 10T 擴到 28T tokens,新增物件偵測與函式呼叫能力,量化後在 Jetson Orin 上每幀僅 233 毫秒,瞄準邊緣裝置上的多模態應用。
閱讀文章 ↗PrismML 推出 1-bit Bonsai:把 8B 模型壓進 1.15 GB 的端側 LLM
2026 年 3 月 31 日,Caltech 衍生新創 PrismML 走出 stealth,發表號稱首款可商業化的端到端 1-bit LLM 家族 Bonsai:8B 旗艦僅 1.15 GB、比 16 位元同級小 14 倍,以 Apache 2.0 開源上架 Hugging Face。
閱讀文章 ↗Qwen3.5 小模型補齊戰線:9B 在多項基準超越 gpt-oss-120b
阿里巴巴 2 月底補上 Qwen3.5 的 0.8B–9B 四款小模型:Apache-2.0 開源、原生 262K 情境、支援影像影片輸入。9B 以 9.65B 密集參數在 MMLU-Pro、GPQA Diamond 超越 gpt-oss-120b,目標是手機與筆電上的本地多模態。
閱讀文章 ↗Axelera AI 籌 2.5 億美元拚邊緣推理晶片
荷蘭新創 Axelera AI 於 2 月 24 日宣布融資超過 2.5 億美元,是歐盟 AI 半導體公司史上最大單筆投資。其 Metis 晶片以約 10 瓦功耗提供 214 TOPS,採記憶體內運算架構,已出貨超過 500 家客戶。
閱讀文章 ↗Cohere Tiny Aya:33.5 億參數、70+ 語言的本機開源模型
Cohere Labs 於 2026 年 2 月 17 日釋出 Tiny Aya 開源多語模型:33.5 億參數、支援 70+ 語言、可在一般硬體本機運行,並選在印度 AI Impact Summit 期間登場。本文解析架構、部署方式與 CC-BY-NC 授權限制。
閱讀文章 ↗Mistral 開源 Voxtral Transcribe 2:即時轉錄挑戰雲端大廠
2026 年 2 月 4 日,Mistral 發表 Voxtral Transcribe 2:批次與串流兩款轉錄模型,FLEURS 詞錯誤率約 4%,即時版以 Apache 2.0 開源、4B 參數可跑邊緣裝置,API 每分鐘 0.003 美元起。本文解析其延遲與準確度取捨和生態定位。
閱讀文章 ↗從 Clawdbot 到 OpenClaw:爆紅開源代理一週二改名,安全疑慮升高
2026 年 1 月 30 日,爆紅開源 AI 代理 Clawdbot 因 Anthropic 商標關切先改名 Moltbot,三日內再更名 OpenClaw。本文解析兩次改名始末、代理架構,以及公網暴露的控制介面、詐騙與企業影子採用等安全風險。
閱讀文章 ↗Ollama 推出 ollama launch:一行指令把 coding agent 接上本機 model
Ollama 於 2026 年 1 月 23 日發表 ollama launch,一行指令自動安裝並設定 Claude Code、OpenCode 與 Codex,讓這些 coding agent 直接對接本機 model。本文解析本地開發工作流在隱私與成本上的實際影響。
閱讀文章 ↗Intel CES 2026 發表 Core Ultra Series 3:18A 製程上機的 AI PC 首役
Intel 在 CES 2026 發表 Core Ultra Series 3(Panther Lake):首款運算晶片採用 18A 製程的 AI PC 平台,搭配 Arc B390 內顯與 200 款以上筆電設計,1 月 27 日全球開賣。本文解析規格、時程與對邊緣 AI 佈局的意義。
閱讀文章 ↗高通 Snapdragon X2 Plus 登場:把 80 TOPS NPU 帶進主流筆電
CES 2026 媒體日,高通發表 Snapdragon X2 Plus:6 核與 10 核、3nm 製程,搭載與旗艦 X2 Elite 同級的 80 TOPS Hexagon NPU,鎖定 799 至 1,299 美元主流價位帶,HP、Lenovo、ASUS 首波裝置 2026 上半年出貨。
閱讀文章 ↗
2025
2 篇文章蘋果開放裝置端LLM:Foundation Models框架
2025年6月9日蘋果在WWDC發表Foundation Models框架,第三方App可直接呼叫Apple Intelligence的裝置端語言模型,支援guided generation、工具呼叫與離線推論,且推論免API費用;Live Translation與Shortcuts智慧動作同步登場。
閱讀文章 ↗Google推出AI Edge Gallery:手機離線跑開源模型
2025年5月底Google低調釋出實驗性Android應用AI Edge Gallery,可從Hugging Face下載Gemma等開源模型,在手機上完全離線推理,採Apache 2.0授權並開源於GitHub;本文整理功能、應用場景與效能限制。
閱讀文章 ↗
2026
24 ARTICLESHermes Agent: Self-Hosted AI That Writes Its Own Skills
Nous Research's open-source Hermes Agent keeps its memory on your machine and writes its own skills. Its design bets, the Desktop beta, $1.5B valuation talks, and how to read its OpenRouter lead.
READ POST ↗GLM-5.3 Weights Are Out: Running a 753B MoE Model Locally
Z.ai put the 753B-parameter GLM-5.3 weights on Hugging Face on August 25, and the Hacker News thread hit 806 points. The architecture, the custom license, and local-run realities.
READ POST ↗NVIDIA Announces Jetson Orin Nano 2: Entry-Level Edge AI Compute Doubles for Robots and Drones
NVIDIA announced the Jetson Orin Nano 2 on August 25, 2026, roughly doubling entry-level edge AI compute for robots and drones. What raising the floor means for edge inference and robotics teams.
READ POST ↗Apple M6 and M5 Ultra: 2nm and Quad-Die for On-Device AI
Apple's first 2nm chip (M6) and first quad-die design (M5 Ultra) arrive in Mac mini and Mac Studio, with up to 512GB unified memory for hundred-billion-parameter local LLMs.
READ POST ↗Qwen3.8-27B: Great Open Weights That Overthink by Default
Alibaba's Qwen shipped the Apache 2.0 Qwen3.8-27B a week after promising it: 262K context, native vision, laptop-class local AI — behind a costly xhigh reasoning default.
READ POST ↗Meta's Muse Glimmer: a 30B Open Model for Local Agents
Meta released Muse Glimmer on August 10, 2026: a 30B open-weight model under Apache 2.0, distilled for always-on local agent workflows on a single consumer GPU.
READ POST ↗Running Gemma 4 26B in 2 GB of RAM: TurboFieldfare
Open-source Swift and Metal engine TurboFieldfare keeps about 2 GB resident, streaming Gemma 4 26B-A4B MoE experts from SSD per token — 5-6 tok/s on an 8GB MacBook Air.
READ POST ↗LM Studio Bionic: A Local-First AI Agent for Open Models
LM Studio shipped Bionic on July 16, 2026 — a separate agent app for open models covering coding, documents, and local voice input, runnable locally or via Secure Cloud.
READ POST ↗Snap Ships Specs: A $2,195 On-Device AR Glasses Gamble
Snap's Specs AR glasses open for preorder at $2,195 with dual Snapdragon chips, a 51-degree FOV, and fully on-device compute. Shares fell more than 5% after the Long Beach debut.
READ POST ↗Ollama 0.30: GGUF Support, Up to 20% Faster NVIDIA, Vulkan On by Default
Ollama 0.30 deepens its GGUF engine via llama.cpp: up to 20% faster NVIDIA throughput under a stated RTX 5090 condition, Vulkan on by default, and tool calling that carries to coding agents.
READ POST ↗Liquid AI LFM2.5-8B-A1B: On-Device MoE Reasoning Model
Liquid AI released LFM2.5-8B-A1B: an 8B-total, ~1B-active MoE reasoning model trained on 38T tokens with 128K context and 253 tok/s CPU inference. Weights are on Hugging Face.
READ POST ↗IrisGo: The AI Butler That Learns Your Desktop by Watching
IrisGo, backed by Andrew Ng's AI Fund, ships a desktop agent that learns workflows by watching you demo them once, runs on your NPU, and will come preloaded on new Acer PCs.
READ POST ↗antirez's ds4: A Local LLM Inference Engine Built for Metal
Redis creator antirez open-sourced ds4: a C local inference engine tuned for DeepSeek V4 Flash and GLM on Metal, CUDA and ROCm, with a coding agent. MIT licensed.
READ POST ↗Qwen3.6-35B-A3B: Open-Weight MoE Punches at Agentic Coding
Alibaba open-sources Qwen3.6-35B-A3B (Apache-2.0): a 35B MoE with 3B active params, 262K context, Terminal-Bench 2.0 at 51.5 — beating Opus 4.7 on a same-day laptop SVG test.
READ POST ↗Liquid AI Ships LFM2.5-VL-450M, a 450M Edge VLM
Liquid AI released LFM2.5-VL-450M on April 8: 28T-token pretraining, new detection and function-calling skills, and 233 ms per frame on a Jetson Orin once quantized.
READ POST ↗PrismML 1-bit Bonsai: An 8B LLM Squeezed Into 1.15 GB
Caltech spinoff PrismML emerged from stealth with 1-bit Bonsai, an end-to-end 1-bit LLM family: the 8B flagship is 1.15 GB, 14x smaller than 16-bit peers, Apache 2.0 licensed.
READ POST ↗Qwen3.5 Small Models: 9B Rivals gpt-oss-120b at the Edge
Alibaba adds 0.8B-9B small models to Qwen3.5 under Apache-2.0: native 262K context, image and video input. The 9B beats gpt-oss-120b on MMLU-Pro and GPQA Diamond, targeting phones and laptops.
READ POST ↗Axelera AI Raises $250M-Plus for Edge Inference Chips
Dutch chipmaker Axelera AI raised over $250 million — the largest investment ever in an EU AI semiconductor company — to scale its 214-TOPS, 10-watt Metis edge inference chips.
READ POST ↗Cohere's Tiny Aya: A 3.35B Open Model for 70+ Languages
Cohere Labs released Tiny Aya: a 3.35B open-weight model covering 70+ languages, built to run on a laptop, debuted at the India AI Impact Summit. Architecture, deployment, and license, broken down.
READ POST ↗Mistral Voxtral Transcribe 2: Open Speech-to-Text, On-Device
Mistral's Voxtral Transcribe 2 ships batch and streaming transcription models: ~4% WER on FLEURS, Apache 2.0 open weights for the 4B realtime model, and $0.003/min API pricing.
READ POST ↗Clawdbot to OpenClaw: Open-Source Agent Hits Security Wall
Viral open-source agent Clawdbot became OpenClaw on January 30, 2026 after an Anthropic trademark complaint. Behind the rename chaos: exposed control panels, scams, and shadow enterprise use.
READ POST ↗Ollama Ships ollama launch: One Command to Wire Up Local Coding Agents
Ollama January 23, 2026 post introduced ollama launch: one command that installs and configures Claude Code, OpenCode, and Codex against local models. What it changes for coding workflows.
READ POST ↗Intel's Core Ultra Series 3 Brings 18A Chips to AI PCs
At CES 2026 Intel launched Core Ultra Series 3 (Panther Lake), its first AI PC platform on 18A, with Arc B390 graphics, 200+ laptop designs, and January 27 availability. What it changes.
READ POST ↗Snapdragon X2 Plus: Qualcomm Pushes Arm Into Mainstream PCs
At CES 2026 Qualcomm launched the Snapdragon X2 Plus: 6- and 10-core 3nm PC chips with the same 80-TOPS NPU as the flagship, targeting $799-$1,299 laptops in H1 2026.
READ POST ↗
2025
2 ARTICLESApple Foundation Models framework opens on-device LLM
Unveiled June 9, 2025: Apple's Foundation Models framework gives third-party apps direct access to the on-device Apple Intelligence model, with guided generation and free offline inference.
READ POST ↗Google AI Edge Gallery runs open AI models offline on phones
In late May 2025 Google quietly released AI Edge Gallery, an experimental Android app running open models like Gemma fully offline on phones. Features, limits, and why it matters.
READ POST ↗