AI Reasoning Models: The Building Blocks for the Next Generation of AI Applications

AI Reasoning Models: The Building Blocks for the Next Generation of AI Applications

1h 17mAdvanced2026-09-11

Authors

Sebastian Raschka

Sebastian Raschka

Course details

Want to explore the emerging class of reasoning-enhanced large language models (LLMs), otherwise known as "reasoning models?" In this course, learn how these models differ from conventional LLMs and why they are particularly well-suited for complex tasks like coding challenges, puzzles, and other areas that require more advanced problem-solving. Through real-world case studies, such as DeepSeek-R1, the course breaks down three main methods for building and improving reasoning capabilities in LLMs. Along the way, get hands-on experience with inference-time scaling, reinforcement learning, and distillation. You'll also learn about the trade-offs between qualitative performance, efficiency, and development cost when using reasoning models versus conventional models. By the end of this course, you'll be prepared to build and use reasoning models on your next project.

Learning objectives
Describe the differences between conventional LLMs and reasoning-enhanced LLM models.
Identify the advantages and disadvantages of using reasoning models compared to conventional LLMs to make informed choices.
Explain the four main approaches to building reasoning models and when each is most appropriate.
Compare the trade-offs between inference-time scaling and training methods when developing reasoning LLMs.
Identify cost-effective strategies for developing reasoning models on limited budgets, such as distillation.

Concepts

Introduction

  • Understanding LLMs and reasoning models

What Are Reasoning Models

  • Developing LLMs - A brief overview
  • LLMs vs. reasoning models
  • What reasoning means in LLMs
  • Use cases for reasoning models

Improving Reasoning with Inference-Time Scaling

  • What is inference-time scaling
  • Hands-on example - Chain-of-thought (CoT) prompting
  • Self-consistency and majority voting
  • Hands-on example - Using LLMs via GitHub models
  • Hands-on example - Inference-time scaling via API calls
  • Test-time scaling vs. training scaling - Performance and cost trade-offs

Distillation for Reasoning Models

  • What is supervised fine-tuning (SFT)
  • What is distillation in LLMs
  • Types of SFT data for reasoning
  • Case study - DeepSeek R1-Distill
  • Cost-efficient distillation projects
  • Hands-on - A distilled model in action

Reinforcement Learning for Reasoning

  • Intro to reinforcement learning (RL) in LLMs
  • Training signals - Rule-based vs. human feedback
  • Case study - DeepSeek-R1-Zero and reasoning with pure RL
  • RL and instruction fine-tuning (SFT)
  • Hands-on example - Comparing RL-based reasoning models
  • Distillation vs. reinforcement learning

Conclusion

  • Next steps with reasoning models
40,000 Toman