2026
8 篇文章Claude Fable 推翻雅可比猜想:數學界震動的一週
Anthropic 研究員 Levent Alpöge 與 Claude Fable 5 在世界盃決賽夜提出雅可比猜想反例,隔天完成 Lean 驗證、陶哲軒撰文解析,AI 正在改寫數學研究的驗證流程。
閱讀文章 ↗Qwen 3.8 登場:阿里 2.4 兆參數旗艦先上訂閱,權重還沒開
2026 年 7 月 19 日,阿里巴巴發布 Qwen 3.8,2.4 兆參數旗艦以 Qwen3.8-Max-Preview 形態先在訂閱方案上線,開放權重尚未釋出,被視為對 Kimi K3 的回應。
閱讀文章 ↗Meta 自導式測試時訓練:讓 LLM 自己挑該學的長上下文證據
Meta AI 提出自導式測試時訓練 S-TTT,先讓模型自行挑選與問題相關的證據段落、再只對這些段落做適應訓練,在 LongBench-v2 與 LongBench-Pro 上帶來最高 15% 的相對準確率提升。
閱讀文章 ↗OpenAI 模型推翻 Erdős 單位距離猜想:AI 數學的首座里程碑
2026 年 5 月 20 日,OpenAI 宣布一個通用推理模型自主推翻 Erdős 1946 年的單位距離猜想,建構出超過 n^1.014 個單位距離點對,Will Sawin 補上明確指數。本文解析證明思路與對 AI 數學研究的意義。
閱讀文章 ↗Grok 4.20 公測:四個代理人先辯論再回話
xAI 於 2026 年 2 月 17 日推出 Grok 4.20 公開測試:內建 Grok、Harper、Benjamin、Lucas 四個代理人,推理時即時辯論、交叉驗證後才輸出。eWeek 報導幻覺下降約 65%,本文拆解運作機制、成本與對開發者的意義。
閱讀文章 ↗智利領軍推出 Latam-GPT:拉美首個本土開源大模型
2026 年 2 月 10 日,智利總統博里奇為 Latam-GPT 揭幕。這個由 CENIA 主導、集結 15 國力量的 700 億參數開源模型,以 8 TB 區域資料訓練,要對抗美國中心偏見。本文解析其規格、資金現實與主權意涵。
閱讀文章 ↗Qwen3-Max-Thinking 登場:阿里巴巴的兆級參數專有推理 model
阿里巴巴 Qwen 團隊於 2026 年 1 月下旬(約 25 日)發表 Qwen3-Max-Thinking:兆級參數的專有推理旗艦,主打自適應工具使用。InfoWorld 認為它讓企業的 model 選項再添一員。本文解析其定位與選型意義。
閱讀文章 ↗Epoch AI:中國模型平均落後美國前沿 7 個月,差距 4 到 14 個月
Epoch AI 於 2026 年 1 月 2 日發布 ECI 分析:2023 年以來中國模型平均落後美國前沿 7 個月,區間 4 至 14 個月,尚無模型超越 o3。本文拆解計算方法、開源權重的干擾變數與社群爭論。
閱讀文章 ↗
2026
9 ARTICLESClaude Fable Disproves the Jacobian Conjecture
Anthropic researcher Levent Alpoge and Claude Fable 5 produced a Jacobian conjecture counterexample on World Cup final night; it was machine-checked in Lean by morning.
READ POST ↗Qwen 3.8: Subscription First, Open Weights Later
Alibaba announced Qwen 3.8 on July 19, 2026, putting its 2.4T-parameter flagship behind the Qwen Cloud token plan as Qwen3.8-Max-Preview, with open weights still pending.
READ POST ↗Meta's Self-Guided Test-Time Training for Long-Context LLMs
Meta AI's S-TTT has models pick the evidence spans worth learning before test-time adaptation, cutting noise and gaining up to 15% on LongBench-v2 and LongBench-Pro.
READ POST ↗Claude Fable 5 Review: An Autonomous Worker for Long Tasks, Not a Stronger Chatbot
Re-examining Claude Fable 5 against the official announcement and prompting guide: positioning, long-task evidence like Stripe's one-day migration, cost and safety boundaries, and deployment changes.
READ POST ↗OpenAI Model Disproves Erdos Unit Distance Conjecture
OpenAI said a general reasoning model disproved Erdos's 1946 unit distance conjecture; Sawin's refinement yields n^1.014 unit-distance pairs. A milestone for AI in mathematics.
READ POST ↗Grok 4.20 Beta: Four Agents That Debate Before They Answer
xAI's Grok 4.20 beta ships four named agents that debate in real time before answering. eWeek reports hallucinations down about 65%. Here is how the loop works and what it costs builders.
READ POST ↗Latam-GPT: Latin America's First Homegrown Open-Source LLM
Chile launched Latam-GPT on Feb 10: a 70B open-source model from CENIA and 15 partner countries, trained on 8 TB of regional data to counter US-centric bias.
READ POST ↗Qwen3-Max-Thinking Arrives: Alibaba's Trillion-Parameter Proprietary Reasoner
Alibaba Qwen team released Qwen3-Max-Thinking in late January 2026: a trillion-parameter proprietary reasoning flagship with adaptive tool use, aimed at enterprise model selection.
READ POST ↗Epoch AI: Chinese Models Trail US Frontier by Seven Months
Epoch AI's ECI index puts Chinese models seven months behind the US frontier on average since 2023, ranging four to fourteen months. How it is measured, and the open-weight catch.
READ POST ↗