System Overview
OpenModex sits between your application and AI model providers. Every request flows through a pipeline of intelligent middleware before reaching the optimal provider.Request Flow
1
Authentication
Your API key is validated and mapped to your account, team, and rate limits.
2
Rate Limiting
Request is checked against your tier’s rate limits (per-key, per-minute). Uses Redis sorted sets for precise sliding-window enforcement.
3
Cache Check
If prompt caching is enabled, OpenModex checks for an exact-match cached response. Cache hits return instantly with zero provider cost.
4
Routing Decision
The routing engine selects the optimal provider based on your strategy (cost, latency, or quality), current provider health, and fallback configuration.
5
Provider Request
The request is translated to the target provider’s format and forwarded. OpenModex supports real-time SSE streaming pass-through.
6
Response + Metadata
The provider response is normalized to OpenAI format, enriched with OpenModex metadata (routing info, cost, latency), and returned to your app.
Core Components
API Gateway
- OpenAI-compatible REST API — drop-in replacement for OpenAI’s API
- SSE streaming — real-time token-by-token delivery
- Idempotency — safe retries with
Idempotency-Keyheader (24h TTL)
Routing Engine
The routing engine maintains a real-time view of provider health and performance:Circuit Breaker
Provider health is monitored automatically:- Closed — requests route normally
- Half-Open — after consecutive failures, traffic is throttled
- Open — provider is temporarily removed from the routing pool
- Recovery — periodic health probes restore the provider
Provider Adapters
Each AI provider is wrapped in a standardized adapter that handles:- Request format translation (OpenAI format → provider-native format)
- Response normalization (provider-native → OpenAI format)
- Error mapping to consistent error codes
- Streaming protocol translation
Caching Layer
- Exact-match caching — hash of (model + messages + parameters)
- Configurable TTL — 60s to 24h per request
- Redis-backed — sub-millisecond cache lookups
- Cost savings — cached responses billed at ~50% input rate, zero output cost
Billing Pipeline
- Every request emits a billing event to Kafka
- Events are consumed asynchronously and persisted
- Account balances are updated atomically via Redis Lua scripts
- Analytics are flushed to Apache Doris for dashboards and reporting
Infrastructure
Security
- API keys are hashed at rest; only the prefix is stored for lookup
- BYOK provider keys are encrypted with AWS KMS envelope encryption
- Rate limiting uses atomic Redis operations to prevent race conditions
- All traffic is encrypted with TLS 1.3
- Passwords are hashed with Argon2id