API Reference
Complete reference for the Aion Labs REST API. All endpoints are under https://api.aionlabs.ai/v1/ and are compatible with the OpenAI client libraries.
Authentication
All endpoints except GET /v1/models require an API key.
Authorization: Bearer YOUR_API_KEYThe Api-Key YOUR_API_KEY header prefix is also accepted. Keys are issued and revoked in the dashboard.
Errors
All errors share the same envelope:
{
"error": {
"message": "Human-readable description",
"type": "<error_type>"
}
}| Status | Type | Cause |
|---|---|---|
| 400 | invalid_request_error | Unknown model or malformed request |
| 429 | rate_limit_error | Usage limit exceeded for your account |
| 502 | server_error | Temporary server-side error — retry the request |
Returns the catalog of available models in the standard OpenAI list envelope, so OpenAI-compatible clients can auto-discover models. No authentication required.
Response
{
"object": "list",
"data": [
{
"id": "aion-labs/aion-3.5",
"object": "model",
"created": 1790121600,
"name": "AionLabs: Aion 3.5",
"description": "...",
"date": "2026-09-23",
"context_length": 262144,
"max_completion_tokens": 32768,
"is_moderated": false,
"architecture": { "modality": "text->text" },
"pricing": {
"prompt": "0.000003",
"completion": "0.000006",
"input_cache_read": "0.00000075"
},
"reasoning": true,
"tool_calls": true,
"reasoning_effort": {
"supported": true,
"levels": ["low", "high", "max"],
"default": "high"
}
}
],
"models": [ ... ]
}Each entry under data carries the standard OpenAI object and created (Unix seconds, midnight UTC of the release date) fields alongside our extended metadata. The legacy models key holds the same entries without those two fields and is kept for existing integrations — new clients should read data.
Pricing fields are per-token. input_cache_read is only present on models with cache pricing.
Per-model parameter capabilities
Each entry states which request parameters the model accepts, so a client or gateway can translate requests from the catalog instead of hard-coding model families. The values mirror what the API enforces at request time.
| Field | Type | Description |
|---|---|---|
| reasoning | boolean | Whether the model reasons before answering (and so can return a reasoning field). |
| tool_calls | boolean | Whether the model accepts tools and can return tool_calls. |
| reasoning_effort.supported | boolean | Whether the model accepts the reasoning_effort parameter. When false, levels is empty and default is null; sending any value other than medium returns a 400. |
| reasoning_effort.levels | array | The exact reasoning_effort values the model accepts. Any other value returns a 400; values are never silently coerced. |
| reasoning_effort.default | string | What an omitted reasoning_effort means for this model. Always one of levels. |
OpenAI-compatible chat completions. Supports streaming and tool calls.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model ID. See Models. |
| messages | array | Yes | Conversation history. See message object below. |
| temperature | float | No | Sampling temperature. |
| max_tokens | integer | No | Maximum completion tokens (≥ 1). This cap bounds the total tokens the model generates for the request — the model's internal reasoning and the visible answer share this budget, and reasoning is generated first. A cap that is exhausted during reasoning returns the reasoning only with finish_reason=length and empty content (the reasoning is available in reasoning when reasoning_split is on). Size max_tokens to cover reasoning plus the expected answer; it is a ceiling, not a reservation, so unused budget is not billed. |
| stop | string[] | No | Stop sequences. |
| stream | boolean | No | Stream response as SSE. Default: false. Non-streaming responses must complete within 120 seconds or the edge returns HTTP 524; generation still completes server-side and is billed. Use stream: true for requests that may run longer (large max_tokens, reasoning models) — streamed responses have no total-duration limit, and reasoning chunks keep the connection active before visible content starts. Very long prompts count against the same limit: prefill alone can take well over a minute for prompts in the hundreds of thousands of tokens, so send those with stream: true. |
| stream_options | object | No | Streaming requests only. {"continuous_usage_stats": true} (vLLM convention) adds cumulative usage to every chunk — useful for billing partial output when a stream is cancelled mid-generation. include_usage is accepted for compatibility; the final chunk always carries usage. |
| tools | array | No | Tool definitions in OpenAI format. |
| reasoning_effort Aion models | string | No | Controls the reasoning level. Aion 3.5 / 3.5 Mini accept exactly low, high or max (default: high); reasoning cannot be disabled on these models, and any other value returns a 400. Aion 2.0 / 3.0 / 3.0 Mini accept none to disable reasoning, or low, medium, high or max to enable it (default: medium). The models are tuned to run with reasoning enabled; disabling it can degrade output. Not supported on aion-rp-llama-3.1-8b. Each model's accepted values and default are also published per entry in GET /v1/models (reasoning_effort.levels / reasoning_effort.default). |
| reasoning_split | boolean | No | Split <think> reasoning into a separate reasoning field. Defaults on for reasoning models. |
| metadata | object | No | Key-value pairs attached to the request: up to 16 keys, keys up to 64 characters, string values up to 512 characters (no nested objects). |
Message object
| Field | Type | Required | Description |
|---|---|---|---|
| role | string | Yes | system, user, assistant, or tool |
| content | string or parts | No | Message text, or a list of {"type", "text"} content parts. |
| tool_calls | array | No | Tool calls returned by the model. |
| tool_call_id | string | No | ID of the tool call being answered (role tool messages). |
Response (non-streaming)
{
"id": "chatcmpl_abc123",
"object": "chat.completion",
"created": 1700000000,
"model": "aion-labs/aion-2.0",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 8,
"total_tokens": 20,
"prompt_tokens_details": {
"cached_tokens": 0
}
}
}When reasoning_split is active, the message gains a reasoning field and content contains only the response text:
{
"message": {
"role": "assistant",
"reasoning": "Let me think through this...",
"content": "The answer is 42."
}
}Response (streaming)
When stream: true, the endpoint returns text/event-stream. Each event is a JSON delta, terminated by data: [DONE]:
data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" world"},"finish_reason":null}]}
data: [DONE]Alternative completion endpoint modelled on the OpenAI Responses API shape. Accepts the same model set and supports streaming.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model ID. |
| input | string or array | One of | Prompt string, or a list of {"role", "content"} message objects. Mutually exclusive with messages. |
| messages | array | One of | Chat history in the same format as /v1/chat/completions. Mutually exclusive with input. |
| temperature | float | No | Sampling temperature. |
| max_output_tokens | integer | No | Maximum completion tokens (≥ 1). |
| stream | boolean | No | Stream as SSE. Default: false. The same 120-second limit on non-streaming responses applies as for /v1/chat/completions. |
| reasoning_effort Aion models | string | No | Controls the reasoning level. Aion 3.5 / 3.5 Mini accept exactly low, high or max (default: high); reasoning cannot be disabled on these models, and any other value returns a 400. Aion 2.0 / 3.0 / 3.0 Mini accept none to disable reasoning, or low, medium, high or max to enable it (default: medium). The models are tuned to run with reasoning enabled; disabling it can degrade output. Not supported on aion-rp-llama-3.1-8b. Each model's accepted values and default are also published per entry in GET /v1/models (reasoning_effort.levels / reasoning_effort.default). |
| reasoning_split | boolean | No | Split reasoning into a separate content item. |
| metadata | object | No | Key-value pairs attached to the request: up to 16 keys, keys up to 64 characters, string values up to 512 characters (no nested objects). |
When using input as a list, each item's content may be a string or a list of {"type": "input_text", "text": "..."} parts.
Response (non-streaming)
{
"id": "resp_abc123",
"object": "response",
"created_at": 1700000000,
"model": "aion-labs/aion-2.0",
"status": "completed",
"output": [
{
"id": "msg_xyz",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{ "type": "output_text", "text": "Hello! How can I help?" }
]
}
],
"usage": {
"input_tokens": 12,
"output_tokens": 8,
"total_tokens": 20,
"input_tokens_details": {
"cached_tokens": 0
}
}
}When reasoning_split is active, a reasoning item is prepended to content:
"content": [
{ "type": "reasoning", "text": "Let me think through this..." },
{ "type": "output_text", "text": "The answer is 42." }
]