Skip to main content

Overview

Streaming allows you to receive response tokens as they’re generated, rather than waiting for the complete response. This dramatically improves perceived latency for your users.

When to Use Streaming

  • Chat interfaces — show typing indicators and real-time text
  • Long responses — don’t make users wait for the full response
  • Code generation — display code as it’s written

How It Works

  1. Set stream: true in your request
  2. OpenModex opens an SSE (Server-Sent Events) connection
  3. Tokens are sent as data: events as they’re generated
  4. The stream ends with data: [DONE]

Examples

Streaming + Fallbacks

Streaming works seamlessly with fallbacks. If the primary model fails before streaming starts, OpenModex automatically retries with the fallback model: