Skip to main content
POST
Create chat completion

Request Body

string
required
Model ID to use (e.g., gpt-4o, claude-3.5-sonnet, gemini-2.0-flash).
array
required
A list of messages comprising the conversation.
boolean
default:false
If true, returns a stream of Server-Sent Events (SSE).
number
default:1
Sampling temperature between 0 and 2. Lower values are more deterministic.
integer
Maximum number of tokens to generate.
number
default:1
Nucleus sampling parameter.
string | string[]
Stop sequence(s) that will halt generation.
number
default:0
Penalty for token frequency (-2.0 to 2.0).
number
default:0
Penalty for token presence (-2.0 to 2.0).
array
A list of tools the model may call (function calling).
string | object
Controls which tool is called. auto, none, or a specific tool.
object
Force a specific output format (e.g., {"type": "json_object"}).
integer
Deterministic sampling seed for reproducibility.

OpenModex Extensions

object
Configure intelligent routing for this request.
object
Configure prompt caching.

Response

Example

Authorizations

Authorization
string
header
required

API key authentication. Pass your OpenModex API key as a Bearer token in the Authorization header: Authorization: Bearer omx_sk_...

Body

application/json

Request body for creating a chat completion.

model
string
required

ID of the model to use.

messages
object[]
required

A list of messages comprising the conversation so far.

stream
boolean
default:false

If true, partial message deltas will be sent as server-sent events.

temperature
number

Sampling temperature between 0 and 2. Higher values make output more random.

Required range: 0 <= x <= 2
top_p
number

Nucleus sampling parameter. Only consider tokens with top_p probability mass.

n
integer

How many completions to generate for each prompt.

max_tokens
integer

The maximum number of tokens to generate in the completion.

stop

Up to 4 sequences where the API will stop generating further tokens.

frequency_penalty
number

Penalizes new tokens based on their existing frequency in the text so far.

Required range: -2 <= x <= 2
presence_penalty
number

Penalizes new tokens based on whether they appear in the text so far.

Required range: -2 <= x <= 2
logit_bias
object

Modify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.

tools
object[]

A list of tools the model may call.

tool_choice

Controls which tool is called by the model. Can be 'none', 'auto', 'required', or a specific tool object.

response_format
object

An object specifying the format the model must output (e.g., {"type": "json_object"}).

seed
integer

If specified, the system will attempt to sample deterministically.

user
string

A unique identifier representing your end-user, for abuse monitoring.

routing
object

Configuration for intelligent request routing.

cache
object

Configuration for response caching.

Response

Chat completion response. When stream=true, returns SSE stream of ChatCompletionChunk objects.

Response from a chat completion request.

id
string

A unique identifier for the completion.

object
enum<string>

The object type, always 'chat.completion'.

Available options:
chat.completion
created
integer

Unix timestamp (in seconds) of when the completion was created.

model
string

The model used for the completion.

choices
object[]

A list of completion choices.

usage
object

Token usage statistics for a completion request.

system_fingerprint
string

A fingerprint representing the backend configuration.

openmodex
object

OpenModex-specific metadata included in completion responses.