Skip to main content

System Overview

OpenModex sits between your application and AI model providers. Every request flows through a pipeline of intelligent middleware before reaching the optimal provider.

Request Flow

1

Authentication

Your API key is validated and mapped to your account, team, and rate limits.
2

Rate Limiting

Request is checked against your tier’s rate limits (per-key, per-minute). Uses Redis sorted sets for precise sliding-window enforcement.
3

Cache Check

If prompt caching is enabled, OpenModex checks for an exact-match cached response. Cache hits return instantly with zero provider cost.
4

Routing Decision

The routing engine selects the optimal provider based on your strategy (cost, latency, or quality), current provider health, and fallback configuration.
5

Provider Request

The request is translated to the target provider’s format and forwarded. OpenModex supports real-time SSE streaming pass-through.
6

Response + Metadata

The provider response is normalized to OpenAI format, enriched with OpenModex metadata (routing info, cost, latency), and returned to your app.

Core Components

API Gateway

  • OpenAI-compatible REST API — drop-in replacement for OpenAI’s API
  • SSE streaming — real-time token-by-token delivery
  • Idempotency — safe retries with Idempotency-Key header (24h TTL)

Routing Engine

The routing engine maintains a real-time view of provider health and performance:

Circuit Breaker

Provider health is monitored automatically:
  1. Closed — requests route normally
  2. Half-Open — after consecutive failures, traffic is throttled
  3. Open — provider is temporarily removed from the routing pool
  4. Recovery — periodic health probes restore the provider

Provider Adapters

Each AI provider is wrapped in a standardized adapter that handles:
  • Request format translation (OpenAI format → provider-native format)
  • Response normalization (provider-native → OpenAI format)
  • Error mapping to consistent error codes
  • Streaming protocol translation
Supported Providers:

Caching Layer

  • Exact-match caching — hash of (model + messages + parameters)
  • Configurable TTL — 60s to 24h per request
  • Redis-backed — sub-millisecond cache lookups
  • Cost savings — cached responses billed at ~50% input rate, zero output cost

Billing Pipeline

  1. Every request emits a billing event to Kafka
  2. Events are consumed asynchronously and persisted
  3. Account balances are updated atomically via Redis Lua scripts
  4. Analytics are flushed to Apache Doris for dashboards and reporting

Infrastructure

Security

  • API keys are hashed at rest; only the prefix is stored for lookup
  • BYOK provider keys are encrypted with AWS KMS envelope encryption
  • Rate limiting uses atomic Redis operations to prevent race conditions
  • All traffic is encrypted with TLS 1.3
  • Passwords are hashed with Argon2id

Data Flow Diagram