May 2026
Intermediate
376 pages
9h 52m
English
This chapter introduces the concept of test-time compute, and it surveys state-of-the-art SLMs and libraries. It also provides a complete commodity-hardware example, showing how you can apply the Group Relative Policy Optimization (GRPO) technique used to train the DeepSeek-R1 models to specialize an SLM for a given domain.
Test-time compute (TTC) is a new concept for LLMs—it emerged in 2024 and refers to the computational ...
Read now
Unlock full access