此主題的繁體中文文章,最新優先。English posts in this topic, newest first.
Exa 用兩個只有搜尋後端不同的 RL agent 做對照:Exa 訓練的 4B 在六個 benchmark 全數領先 SERP 版,還常超過未訓練的 235B,並以 0.58B tokens 追平 1.89B 的表現。本文拆解更密 reward signal 的機制、47 萬次搜尋呼叫的穩定性數據,以及這份廠商研究的限制。
Exa's controlled experiment shows that changing the search backend during RL training significantly impacts agent performance and efficiency.