Test-time compute
Extra computation spent when a model answers a question, at inference time, rather than only during training; a lever for harder reasoning tasks.
Test-time compute means giving a model more processing when you ask it a question, not just when you train it. Reasoning systems such as OpenAI o1 and DeepSeek-R1 spend longer on chains of thought at answer time, trading speed and cost for accuracy on hard tasks.
It is the complement to scaling training data and parameters: a second axis of progress after raw model size. Competition benchmarks such as the IMO gold result in 2025 showed how far that lever had moved in a single year.