Skip to content
繁體中文
Chat

Create Chat Completion ​

OpenAI-compatible Chat Completions API for text generation and multi-turn conversations.

POST /v1/chat/completions

Creates a model response for the given chat conversation. Returns a chat completion object, or a streamed sequence of chat completion chunk objects if the request is streamed.

  • Uses the OpenAI Chat Completions API request format
  • Plain text messages and multi-turn conversations
  • Streaming and non-streaming responses

For image, audio, and file analysis, see File Analysis. For tool use, see Tool Calling.

Endpoint ​

text
https://api.tokatlas.ai/v1/chat/completions

Authentication ​

All endpoints require Bearer Token authentication.

http
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

Important

Never commit a real API key to a repository or expose it in client-side code.

Request Body ​

ParameterTypeRequiredDefaultDescription
modelstringYes-Model ID used to generate the response
messagesarrayYes-List of messages comprising the conversation
temperaturenumberNo1.0Sampling temperature, range 0–2
top_pnumberNo1.0Nucleus sampling parameter, range 0–1
max_tokensintegerNo-Maximum tokens to generate (deprecated for o-series; use max_completion_tokens)
max_completion_tokensintegerNo-Upper bound for generated tokens, including reasoning tokens
streambooleanNofalseWhether to stream the response via SSE
stream_optionsobjectNo-Options for streaming (only when stream: true)
stopstring or arrayNo-Up to 4 sequences where generation stops
nintegerNo1Number of chat completion choices to generate
frequency_penaltynumberNo0Frequency penalty, range -2.0 to 2.0
presence_penaltynumberNo0Presence penalty, range -2.0 to 2.0
response_formatobjectNo-Output format: text or JSON schema
reasoning_effortstringNo-Reasoning effort for reasoning models
metadataobjectNo-Up to 16 key-value pairs attached to the object
storebooleanNo-Whether to store the output for later retrieval

model ​

First list available models and confirm Chat Completions support. Current upstream models are listed by OpenAI, Claude, and Gemini.

For example, GPT-6.1 Sol and GPT-6 Astra support text requests in Chat, but tool calls require Responses. Other providers also require protocol conversion on the Tokatlas route; a model name alone does not establish compatibility.

Current upstream IDs include:

  • OpenAI: gpt-6.1-sol, gpt-6-astra, gpt-6-luna
  • Anthropic: claude-opus-5-5, claude-sonnet-5-5, claude-fable-5-1
  • Google: gemini-3.8-flash, gemini-3.5-flash-lite

messages ​

A list of messages comprising the conversation so far. Each message includes role and content.

RoleDescription
developerDeveloper-provided instructions. With o1 models and newer, developer messages replace system messages
systemSystem prompt to set AI behavior (use developer for o1 and newer)
userMessages sent by the end user
assistantMessages sent by the model in response to user messages
toolTool results returned by your application with the matching tool_call_id

Basic user message:

json
[{"role": "user", "content": "Hello!"}]

Developer / system prompt:

json
[
  {"role": "developer", "content": "You are a helpful assistant."},
  {"role": "user", "content": "Hello!"}
]

Multi-turn conversation:

json
[
  {"role": "user", "content": "Hello"},
  {"role": "assistant", "content": "Hi! How can I help you?"},
  {"role": "user", "content": "Tell me about AI"}
]

For text messages, content is a plain string. It may also be an array of text content parts:

json
{"role": "user", "content": [{"type": "text", "text": "Hello!"}]}

temperature ​

What sampling temperature to use, between 0 and 2. Higher values like 0.8 make output more random; lower values like 0.2 make it more focused. Recommend using either temperature or top_p, not both.

top_p ​

Nucleus sampling parameter, range 0–1.

max_tokens / max_completion_tokens ​

  • max_tokens — maximum tokens to generate. Deprecated in favor of max_completion_tokens; not compatible with o-series models.
  • max_completion_tokens — upper bound for generated tokens, including visible output and reasoning tokens.

stream ​

If set to true, the model response is streamed to the client as it is generated using Server-Sent Events. Each chunk has object set to chat.completion.chunk.

  • true — Streaming response
  • false — Complete response at once (default)

stream_options ​

Options for streaming response. Only set when stream: true.

FieldTypeDescription
include_usagebooleanStream a final chunk with token usage before data: [DONE]
include_obfuscationbooleanAdd obfuscation fields to normalize payload sizes (default true)

stop ​

Up to 4 sequences where the API stops generating further tokens. Not supported with latest reasoning models o3 and o4-mini.

n ​

How many chat completion choices to generate for each input message. Keep n as 1 to minimize costs.

frequency_penalty / presence_penalty ​

Number between -2.0 and 2.0.

  • frequency_penalty — penalizes tokens based on their existing frequency in the text
  • presence_penalty — penalizes tokens based on whether they appear in the text so far

response_format ​

An object specifying the format the model must output.

Plain text (default):

json
{"type": "text"}

Structured JSON output (JSON Schema):

json
{
  "type": "json_schema",
  "json_schema": {
    "name": "math_response",
    "schema": {
      "type": "object",
      "properties": {
        "steps": {"type": "array", "items": {"type": "string"}},
        "final_answer": {"type": "string"}
      },
      "required": ["steps", "final_answer"],
      "additionalProperties": false
    },
    "strict": true
  }
}

reasoning_effort ​

Constrains effort on reasoning for reasoning models. Supported values: none, minimal, low, medium, high, xhigh, max.

json
{
  "model": "o3-mini",
  "messages": [{"role": "user", "content": "Solve this math problem"}],
  "reasoning_effort": "high"
}

metadata ​

Up to 16 key-value pairs (keys max 64 chars, values max 512 chars) for storing additional information about the object.

Response ​

The table describes the protocol response object. Where the non-streaming examples use a { "code": 200, "data": { ... } } envelope, read the object in data; for routes returning the protocol object directly, read the root object. Parse streaming responses as events, not one JSON document. Native SDKs require a route that returns their expected protocol format directly.

FieldTypeDescription
idstringUnique identifier for the chat completion
objectstringObject type, always chat.completion
createdintegerUnix timestamp of when the completion was created
modelstringModel used for the chat completion
choicesarrayList of chat completion choices
usageobjectToken usage statistics
system_fingerprintstringBackend configuration fingerprint
service_tierstringProcessing tier used to serve the request

choices[] ​

FieldTypeDescription
indexintegerIndex of the choice in the list
messageobjectMessage generated by the model
finish_reasonstringWhy the model stopped generating tokens
logprobsobject or nullLog probability information

message

FieldTypeDescription
rolestringAlways assistant
contentstring or nullGenerated text content
refusalstring or nullRefusal message, if any

finish_reason possible values:

ValueDescription
stopNatural stop point or stop sequence reached
lengthMaximum token limit reached
content_filterContent omitted due to content filters
tool_callsThe model requested tool calls; the application must execute them and return results

usage ​

FieldTypeDescription
prompt_tokensintegerNumber of tokens in the prompt
completion_tokensintegerNumber of tokens in the generated completion
total_tokensintegerTotal tokens used (prompt + completion)
prompt_tokens_detailsobjectBreakdown including cached_tokens
completion_tokens_detailsobjectBreakdown including reasoning_tokens

Usage Examples ​

Basic Conversation ​

json
{
  "model": "gpt-5",
  "messages": [
    {"role": "developer", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Hello!"}
  ]
}

System Prompt ​

json
{
  "model": "gpt-5",
  "messages": [
    {"role": "system", "content": "You are a professional Python programming tutor"},
    {"role": "user", "content": "How should I learn programming?"}
  ]
}

Multi-turn Conversation ​

json
{
  "model": "gpt-5",
  "messages": [
    {"role": "user", "content": "What is machine learning?"},
    {"role": "assistant", "content": "Machine learning is a branch of artificial intelligence..."},
    {"role": "user", "content": "Can you give me a practical example?"}
  ]
}

Non-streaming Output ​

json
{
  "model": "gpt-5",
  "messages": [
    {"role": "user", "content": "Write a poem about spring"}
  ],
  "stream": false
}

Streaming Output ​

json
{
  "model": "gpt-5",
  "messages": [
    {"role": "developer", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Hello!"}
  ],
  "stream": true
}

Request Examples ​

See language setup. Set API_KEY and replace model, file URL, and ID placeholders first. Each version displays the raw response to the same request.

bash
curl --fail-with-body --silent --show-error --max-time 180 \
  --request POST \
  --url "https://api.tokatlas.ai/v1/chat/completions" \
  --header "Authorization: Bearer $API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
  "model": "gpt-5",
  "messages": [
    {
      "role": "developer",
      "content": "You are a helpful assistant."
    },
    {
      "role": "user",
      "content": "Hello!"
    }
  ],
  "stream": false
}'
python
import os
import requests

headers = {
    'Authorization': 'Bearer ' + os.environ["API_KEY"],
    'Content-Type': 'application/json',
}
payload = {'model': 'gpt-5',
 'messages': [{'role': 'developer', 'content': 'You are a helpful assistant.'},
              {'role': 'user', 'content': 'Hello!'}],
 'stream': False}
response = requests.request(
    'POST', 'https://api.tokatlas.ai/v1/chat/completions', headers=headers,
    json=payload,
    timeout=180,
)
response.raise_for_status()
print(response.text)
js
if (!process.env.API_KEY) throw new Error("Set API_KEY first.");
const response = await fetch("https://api.tokatlas.ai/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": "Bearer " + process.env.API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
  "model": "gpt-5",
  "messages": [
    {
      "role": "developer",
      "content": "You are a helpful assistant."
    },
    {
      "role": "user",
      "content": "Hello!"
    }
  ],
  "stream": false
}),
  signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());
java
import java.net.URI;
import java.net.http.*;
import java.time.Duration;

public class Example {
    public static void main(String[] args) throws Exception {
        String apiKey = System.getenv("API_KEY");
        if (apiKey == null || apiKey.isBlank()) {
            throw new IllegalArgumentException("Set API_KEY first.");
        }
        String payload = String.join("\n",
            "{",
            "  \"model\": \"gpt-5\",",
            "  \"messages\": [",
            "    {",
            "      \"role\": \"developer\",",
            "      \"content\": \"You are a helpful assistant.\"",
            "    },",
            "    {",
            "      \"role\": \"user\",",
            "      \"content\": \"Hello!\"",
            "    }",
            "  ],",
            "  \"stream\": false",
            "}"
        );
        HttpClient client = HttpClient.newBuilder()
            .connectTimeout(Duration.ofSeconds(30)).build();
        HttpRequest request = HttpRequest.newBuilder()
            .uri(URI.create("https://api.tokatlas.ai/v1/chat/completions"))
            .timeout(Duration.ofSeconds(180))
            .header("Authorization", "Bearer " + apiKey)
            .header("Content-Type", "application/json")
            .method("POST", HttpRequest.BodyPublishers.ofString(payload))
            .build();
        HttpResponse<String> response = client.send(
            request, HttpResponse.BodyHandlers.ofString());
        if (response.statusCode() < 200 || response.statusCode() >= 300) {
            throw new IllegalStateException("HTTP " + response.statusCode() + ": "
                + response.body());
        }
        System.out.println(response.body());
    }
}
go
package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
)

func main() {
    url := "https://api.tokatlas.ai/v1/chat/completions"

    payload := map[string]interface{}{
        "model": "gpt-5",
        "messages": []map[string]string{
            {"role": "developer", "content": "You are a helpful assistant."},
            {"role": "user", "content": "Hello!"},
        },
        "stream": false,
    }

    jsonData, _ := json.Marshal(payload)

    req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
    req.Header.Set("Authorization", "Bearer "+os.Getenv("API_KEY"))
    req.Header.Set("Content-Type", "application/json")

    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        panic(err)
    }
    defer resp.Body.Close()

    body, _ := io.ReadAll(resp.Body)
    fmt.Println(string(body))
}

Response Examples ​

Non-streaming (stream: false) ​

json
{
  "code": 200,
  "data": {
    "id": "chatcmpl-B9MBs8CjcvOU2jLn4n570S5qMJKcT",
    "object": "chat.completion",
    "created": 1741569952,
    "model": "gpt-5",
    "choices": [
      {
        "index": 0,
        "message": {
          "role": "assistant",
          "content": "Hello! How can I assist you today?",
          "refusal": null
        },
        "logprobs": null,
        "finish_reason": "stop"
      }
    ],
    "usage": {
      "prompt_tokens": 19,
      "completion_tokens": 10,
      "total_tokens": 29,
      "prompt_tokens_details": {
        "cached_tokens": 0
      },
      "completion_tokens_details": {
        "reasoning_tokens": 0
      }
    },
    "service_tier": "default"
  }
}

Streaming (stream: true) ​

When stream is true, the API returns a Server-Sent Events (SSE) stream with Content-Type: text/event-stream. Each event is a JSON object with object set to chat.completion.chunk. Concatenate the choices[].delta.content fields from each chunk to assemble the full response. The stream ends with data: [DONE].

text
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"Hello"},"logprobs":null,"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{"content":"! How can I assist you today?"},"logprobs":null,"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1694268190,"model":"gpt-5","choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}]}

data: [DONE]