Loading timeline…
20242020s
Reasoning Models (Test-Time Compute)
Models that spend compute thinking before answering — the shift from scaling training to scaling inference.
Why It Was Important
Instead of answering immediately, a reasoning model generates a hidden chain of thought, evaluating and backtracking before it responds, with reinforcement learning used to train the thinking itself. OpenAI's o1 established the approach in 2024 and DeepSeek-R1 reproduced it openly in 2025. It opened a second scaling axis: capability now rises with inference compute, not only with training compute, which reshaped both model design and the economics of serving them.
Who Invented It
OpenAI, DeepSeek
The o1 team at OpenAI, reproduced openly by DeepSeek's R1.
Applications
- OpenAI o1
- Advanced Mathematics Verification
- Scientific coding tasks
Key Papers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · arXiv · 2025