# Prompt caching in Strom

> A quick primer about how we use caching at Strom, and how to optimize your requests for maximum cache reuse.

Shashank Aggarwal · Oct 10, 2026 · 1 min read · Engineering
Canonical: https://platform.uprelic.com/blog/prompt-caching

---

There are 2 levels of caching in play for a request fired at Strom:
1. **Within-request caching**: Caching which happens within different parts of a single API request
2. **Across-request caching**: When subsequent requests share a common prefix

## 1. Within-request caching

A request with this structure:
```
State + Question 1 + Question 2 + ...
```

runs internally as one model call per question (concurrently):
```
State + Question 1
State + Question 2
...
```
Clearly, we can cache the state since it's a common prefix. We perform an explicit "priming" step to cache the state, and every question then starts from that cached state. Priming costs an extra call, so we only do it when it pays for itself: when there are enough questions sharing a long enough state. Requests with images aren't primed.


## 2. Across-request caching

This is the more straightforward one. If a subsequent request shares a common prefix (state and optionally questions) with a previous request, Strom will automatically reuse the common prefix, resulting in lower latency.


## Tips for getting the most out of caching
  - Ask many questions in one request rather than one request per question.
  - Keep the state byte-identical between requests that should share it. Standard LLM caching guidelines apply here, such as moving the dynamic part of the state (e.g., context or current date) to the end.
  - The amount of cache reuse also depends on the load level on our servers.
  - Caching only affects latency. The number of billed tokens is only dependent on the request's payload size.
  - Every response reports how many input tokens came from the cache in `usage.cached_input_tokens`.
