2026
7 篇文章把模型快取放進節點:HyperPod 推論冷啟動的實務取捨
Amazon SageMaker HyperPod 推出模型快取,把權重與容器映像預載到節點 NVMe,讓擴容從數十分鐘縮到數秒。
閱讀文章 ↗Megakernel 不是炫技:把 decode 從「等 kernel」改成「等資料」的實作取捨
Cohere 用單一 CUDA 檔把 North Mini Code 的 decode 做成 persistent megakernel,batch size 1 吞吐從 vLLM 的 185 tok/s 拉到 292 tok/s。
閱讀文章 ↗前綴感知路由:讓 KV cache 不再被隨機打散
SageMaker Inference 新增前綴感知路由,把相同 prompt 開頭固定送到同一台機器,讓 prefix caching 真正累積成可重用的 KV cache。
閱讀文章 ↗Cerebras CS-4 登場:三片 Turbo 晶圓的機架級推理系統
Cerebras 於 8 月 19 日發表 CS-4 機架級系統與 WSE-3T 晶圓:不做新矽,把同一片晶圓灌兩倍功率,宣稱推理比 GPU 快 30 倍,並把提示詞處理交給 AWS 與 AMD 加速器分擔。
閱讀文章 ↗Nari Labs 把 Qwen3-TTS 壓進 50 毫秒內開口
Nari Labs 於 8 月 19 日公開 Qwen3-TTS 1.7B 優化成果:單張 H100 上達到每秒 10 次請求、p95 首音延遲低於 50 毫秒,並開源整套服務實作與基準測試。
閱讀文章 ↗AMD 收購 Taalas:把模型權重直接蝕刻進晶片
2026 年 8 月 6 日 AMD 宣布收購多倫多新創 Taalas:以權重蝕刻加 SRAM 快取的專用推理晶片路線,HC1 已在 Llama 3.1 8B 跑出每秒近 1.7 萬 token。
閱讀文章 ↗Amazon 擬對外出售 Trainium 晶片,劍指 NVIDIA
Bloomberg 報導,Amazon AI 主管 Peter DeSantis 證實 AWS 正與潛在買家洽談對外出售 Trainium 晶片;執行長 Andy Jassy 估算晶片事業獨立運作的年營收可達約 500 億美元,直接挑戰 NVIDIA 在資料中心 AI 晶片市場的主導地位。
閱讀文章 ↗
2025
1 篇文章2026
7 ARTICLESModel Caching on HyperPod: What Changes When Weights Live on the Node
HyperPod model caching pre-loads weights and images to local NVMe, cutting scale-out cold starts from tens of minutes to seconds.
READ POST ↗A Single Persistent Kernel Changes How You Serve Code Models
Cohere's North Mini Code megakernel serving engine hits 62% of H100 memory bandwidth, 1.58× faster than vLLM at batch size 1.
READ POST ↗Prefix-Aware Routing on SageMaker: What Changes When Your Prompt Starts the Same Way
SageMaker's new routing strategy sends identical prompt prefixes to the same instance so KV cache actually gets reused.
READ POST ↗Cerebras CS-4: Three WSE-3T Wafers, 30x Faster Inference
Cerebras launched its CS-4 rack-scale system: the same wafer pushed twice as hard, up to 30x faster inference than GPUs, and prefill offloaded to AWS and AMD accelerators.
READ POST ↗Nari Labs Gets Qwen3-TTS Talking in Under 50 ms
Nari Labs open-sourced a Qwen3-TTS 1.7B serving stack that hits 10 requests per second with sub-50 ms p95 time-to-first-audio on a single H100.
READ POST ↗AMD Buys Taalas: Etching AI Models Straight Into Silicon
AMD agreed to buy Toronto startup Taalas, whose chips etch model weights into silicon: Llama 3.1 8B at nearly 17,000 tokens per second on a 6nm test chip.
READ POST ↗Amazon Weighs Selling Trainium Chips Beyond AWS
Amazon is in talks to sell Trainium chips beyond AWS; Jassy sizes the standalone chips business at ~$50 billion a year, taking direct aim at Nvidia's AI data center dominance.
READ POST ↗