Apple《The Illusion of Thinking》:Reasoning Model 的優勢只存在於中等複雜度區間
Apple Machine Learning Research 的 NeurIPS 2025 論文用可控 puzzle environments 檢驗 Large Reasoning Models,提出三個複雜度區間:低複雜度標準 LLM 可能更好;中等複雜度 LRM 有優勢;高複雜度兩者都崩潰,且 LRM 會在仍有 token budget 時降低 reasoning effort。
2506.06941v3(2025-11-20 更新)。作者:Parshin Shojaee、Iman Mirzadeh、Keivan Alizadeh、Maxwell Horton、Samy Bengio、Mehrdad Farajtabar。論文標題為 The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity。近期 frontier models 推出會產生詳細 thinking process 的 LRMs,但既有 benchmark 多看 final answer accuracy,容易受 contamination 影響,也無法分析 reasoning traces 的結構與品質。
作者使用可控 puzzle environments,能精準調整 compositional complexity,同時維持邏輯結構一致。這讓研究者不只看答案對錯,也能檢查模型內部 reasoning trace 如何展開。
論文包含 Tower of Hanoi、Checker Jumping、River Crossing、Blocks World 等 puzzle。這些任務的好處是規則明確、可模擬驗證、可控制難度,且較不依賴既有數學 / coding benchmark 記憶。
Frontier LRMs 在某些複雜度後出現 complete accuracy collapse;更反直覺的是,reasoning effort 會隨複雜度上升到某點後下降,即使仍有足夠 token budget。
| 複雜度區間 | 觀察結果 | 解讀 |
|---|---|---|
| 低複雜度 | standard LLM surprisingly outperform LRMs | 額外 thinking 可能是 overhead,簡單題不一定需要長推理。 |
| 中等複雜度 | LRMs demonstrate advantage | thinking / self-reflection 在需要多步組合時能帶來實際收益。 |
| 高複雜度 | both models experience complete collapse | 當 compositional depth 超出能力,更多 token 不等於更會算;模型可能無法穩定執行明確演算法。 |