6 juni 2025
17 min
https://machinelearning.apple.com/research/illusion-of-thinking
The document investigates the capabilities and limitations of Large Reasoning Models (LRMs), a new generation of language models designed for complex problem-solving. It critiques current evaluation methods, which often rely on mathematical benchmarks prone to data contamination, and instead proposes using controllable puzzle environments to systematically analyze model behavior. The research identifies three distinct performance regimes based on problem complexity: standard models may outperform LRMs at low complexity, LRMs show an advantage at medium complexity, but both collapse at high complexity. Crucially, LRMs exhibit a counter-intuitive decline in reasoning effort as problems become overwhelmingly difficult, despite having available token budgets, and also demonstrate surprising limitations in executing exact algorithms and inconsistent reasoning across different puzzle types.
Lyssna på fler avsnitt från
KnowledgeDB.ai
Visar 1–10 av 37 avsnitt
26 juli 2026
44 min
18 juni 2026
18 min
2 oktober 2025
15 min
30 augusti 2025
6 min
26 juli 2025
23 min
4 juli 2025
19 min
24 juni 2025
11 min
23 juni 2025
22 min
5 juni 2025
23 min
3 juni 2025
20 min