Reduce Agent Costs with Semantic Caching

Reduce Agent Costs with Semantic Caching

1h 1mIntermediate2026-08-13

Authors

Samuel Agbede

Samuel Agbede

Course details

Explore the art of semantic caching with RedisVL to supercharge your LLM-powered applications. Learn how to create a semantic cache using RedisVL, enabling reductions in cost and response time by harnessing the power of vector embeddings. Implement metadata filtering and scoped retrieval to ensure responses maintain contextual accuracy. Gain practical experience by building an LLM application around your cache, including connecting an app to existing cache systems. Discover the nuances of agent-based decisions in caching and learn to evaluate performance through various metrics like hit rate and latency. This course is designed for intermediate to advanced developers and AI engineers looking to advance their skills in deploying effective agent-driven applications. By the end, you'll be more equipped to support metadata filtering, assess cache performance, and optimize agent deployments.

Learning objectives
Implement a semantic caching layer using RedisVL to map vectorized user queries to reusable LLM responses.
Integrate semantic caching into an agent-based LLM application to reduce repeated model calls and improve response latency.
Apply metadata filtering and scoped retrieval to ensure cached responses remain contextually accurate across roles or regions.
Evaluate semantic cache performance by measuring hit rate, latency, and cost tradeoffs to optimize production agent deployments.

Concepts

0. Introduction

  • 01 - How semantic caching streamlines agents

1. Implement Semantic Caching with RedisVL

  • 02 - Build your first semantic cache using RedisVL
  • 03 - Build an LLM app around your cache
  • 04 - Support and validate metadata filtering
  • 05 - Support for tool-calling
  • 06 - Evaluate cache performance

Conclusion

40,000 Toman