AI tech talk series: Semantic caching
Oct 07, 20262:00 PM – 2:45 PM BST
97% believe in it. 4% have built it. New research: State of context engineering.
Every repeat query hitting the LLM at full price adds up fast, and it’s usually finance who notices first, after the AI line item has already tripled. Most teams treat this as a budget problem to fix later, but this is actually a design decision that belongs in the architecture from the start.
This session adds Redis semantic caching, so repeat and near-duplicate queries get served from cache instead of hitting the model again.
We’ll cover how semantic matching decides what counts as a repeat query, where the savings show up first, and the reduction in LLM spend.

Redis
Developer Advocate
Speak to a Redis expert and learn more about enterprise-grade Redis today.