Skip to main content
POST
Create chat completion

Corps de la requête

string
requis
Identifiant du modèle à utiliser (par exemple, gpt-4o, claude-3.5-sonnet, gemini-2.0-flash).
array
requis
Une liste de messages composant la conversation.
boolean
défaut:false
Si true, renvoie un flux de Server-Sent Events (SSE).
number
défaut:1
Température d’échantillonnage entre 0 et 2. Des valeurs plus basses sont plus déterministes.
integer
Nombre maximum de tokens à générer.
number
défaut:1
Paramètre d’échantillonnage par noyau.
string | string[]
Séquence(s) d’arrêt qui interrompront la génération.
number
défaut:0
Pénalité de fréquence des tokens (-2.0 à 2.0).
number
défaut:0
Pénalité de présence des tokens (-2.0 à 2.0).
array
Une liste d’outils que le modèle peut appeler (appel de fonctions).
string | object
Contrôle quel outil est appelé. auto, none ou un outil spécifique.
object
Forcer un format de sortie spécifique (par exemple, {"type": "json_object"}).
integer
Graine d’échantillonnage déterministe pour la reproductibilité.

Extensions OpenModex

object
Configurer le routage intelligent pour cette requête.
object
Configurer la mise en cache des prompts.

Réponse

Exemple

Autorisations

Authorization
string
header
requis

API key authentication. Pass your OpenModex API key as a Bearer token in the Authorization header: Authorization: Bearer omx_sk_...

Corps

application/json

Request body for creating a chat completion.

model
string
requis

ID of the model to use.

messages
object[]
requis

A list of messages comprising the conversation so far.

stream
boolean
défaut:false

If true, partial message deltas will be sent as server-sent events.

temperature
number

Sampling temperature between 0 and 2. Higher values make output more random.

Plage requise: 0 <= x <= 2
top_p
number

Nucleus sampling parameter. Only consider tokens with top_p probability mass.

n
integer

How many completions to generate for each prompt.

max_tokens
integer

The maximum number of tokens to generate in the completion.

stop

Up to 4 sequences where the API will stop generating further tokens.

frequency_penalty
number

Penalizes new tokens based on their existing frequency in the text so far.

Plage requise: -2 <= x <= 2
presence_penalty
number

Penalizes new tokens based on whether they appear in the text so far.

Plage requise: -2 <= x <= 2
logit_bias
object

Modify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.

tools
object[]

A list of tools the model may call.

tool_choice

Controls which tool is called by the model. Can be 'none', 'auto', 'required', or a specific tool object.

response_format
object

An object specifying the format the model must output (e.g., {"type": "json_object"}).

seed
integer

If specified, the system will attempt to sample deterministically.

user
string

A unique identifier representing your end-user, for abuse monitoring.

routing
object

Configuration for intelligent request routing.

cache
object

Configuration for response caching.

Réponse

Chat completion response. When stream=true, returns SSE stream of ChatCompletionChunk objects.

Response from a chat completion request.

id
string

A unique identifier for the completion.

object
enum<string>

The object type, always 'chat.completion'.

Options disponibles:
chat.completion
created
integer

Unix timestamp (in seconds) of when the completion was created.

model
string

The model used for the completion.

choices
object[]

A list of completion choices.

usage
object

Token usage statistics for a completion request.

system_fingerprint
string

A fingerprint representing the backend configuration.

openmodex
object

OpenModex-specific metadata included in completion responses.