Overview
Prompt caching stores responses for identical requests, returning cached results instantly without making a new API call. This is ideal for:- Repeated queries (FAQ bots, common questions)
- Development and testing
- High-traffic endpoints with predictable inputs
Enable Caching
Add thecache parameter to your request:
How It Works
- OpenModex hashes the request (model + messages + parameters)
- If a cached response exists and hasn’t expired, it’s returned immediately
- If no cache exists, the request goes to the provider and the response is cached
Cache Key
The cache key is computed from:- Model name
- Messages array (content and roles)
- Temperature, max_tokens, and other generation parameters
TTL (Time to Live)
Cost Savings
Cached responses are significantly cheaper:- Cached input tokens are billed at ~50% of the normal rate
- No output token cost for cached responses
- Zero latency — responses are returned in <10ms