Skip to main content
POST
Create chat completion

Cuerpo de la Solicitud

string
requerido
ID del modelo a usar (por ejemplo, gpt-4o, claude-3.5-sonnet, gemini-2.0-flash).
array
requerido
Una lista de mensajes que componen la conversación.
boolean
predeterminado:false
Si es true, devuelve un stream de Server-Sent Events (SSE).
number
predeterminado:1
Temperatura de muestreo entre 0 y 2. Valores más bajos son más determinísticos.
integer
Número máximo de tokens a generar.
number
predeterminado:1
Parámetro de muestreo nucleus.
string | string[]
Secuencia(s) de parada que detendrán la generación.
number
predeterminado:0
Penalización por frecuencia de token (-2.0 a 2.0).
number
predeterminado:0
Penalización por presencia de token (-2.0 a 2.0).
array
Una lista de herramientas que el modelo puede llamar (function calling).
string | object
Controla qué herramienta se llama. auto, none o una herramienta específica.
object
Forzar un formato de salida específico (por ejemplo, {"type": "json_object"}).
integer
Semilla de muestreo determinístico para reproducibilidad.

Extensiones OpenModex

object
Configurar enrutamiento inteligente para esta solicitud.
object
Configurar almacenamiento en caché de prompts.

Respuesta

Ejemplo

Autorizaciones

Authorization
string
header
requerido

API key authentication. Pass your OpenModex API key as a Bearer token in the Authorization header: Authorization: Bearer omx_sk_...

Cuerpo

application/json

Request body for creating a chat completion.

model
string
requerido

ID of the model to use.

messages
object[]
requerido

A list of messages comprising the conversation so far.

stream
boolean
predeterminado:false

If true, partial message deltas will be sent as server-sent events.

temperature
number

Sampling temperature between 0 and 2. Higher values make output more random.

Rango requerido: 0 <= x <= 2
top_p
number

Nucleus sampling parameter. Only consider tokens with top_p probability mass.

n
integer

How many completions to generate for each prompt.

max_tokens
integer

The maximum number of tokens to generate in the completion.

stop

Up to 4 sequences where the API will stop generating further tokens.

frequency_penalty
number

Penalizes new tokens based on their existing frequency in the text so far.

Rango requerido: -2 <= x <= 2
presence_penalty
number

Penalizes new tokens based on whether they appear in the text so far.

Rango requerido: -2 <= x <= 2
logit_bias
object

Modify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.

tools
object[]

A list of tools the model may call.

tool_choice

Controls which tool is called by the model. Can be 'none', 'auto', 'required', or a specific tool object.

response_format
object

An object specifying the format the model must output (e.g., {"type": "json_object"}).

seed
integer

If specified, the system will attempt to sample deterministically.

user
string

A unique identifier representing your end-user, for abuse monitoring.

routing
object

Configuration for intelligent request routing.

cache
object

Configuration for response caching.

Respuesta

Chat completion response. When stream=true, returns SSE stream of ChatCompletionChunk objects.

Response from a chat completion request.

id
string

A unique identifier for the completion.

object
enum<string>

The object type, always 'chat.completion'.

Opciones disponibles:
chat.completion
created
integer

Unix timestamp (in seconds) of when the completion was created.

model
string

The model used for the completion.

choices
object[]

A list of completion choices.

usage
object

Token usage statistics for a completion request.

system_fingerprint
string

A fingerprint representing the backend configuration.

openmodex
object

OpenModex-specific metadata included in completion responses.