API Reference

Complete reference for the Aion Labs REST API. All endpoints are under https://api.aionlabs.ai/v1/ and are compatible with the OpenAI client libraries.

Authentication

All endpoints except GET /v1/models require an API key.

Authorization: Bearer YOUR_API_KEY

The Api-Key YOUR_API_KEY header prefix is also accepted. Keys are issued and revoked in the dashboard.

Errors

All errors share the same envelope:

{
  "error": {
    "message": "Human-readable description",
    "type": "<error_type>"
  }
}
Status Type Cause
400 invalid_request_error Unknown model or malformed request
429 rate_limit_error Usage limit exceeded for your account
502 server_error Temporary server-side error — retry the request
GET /v1/models

Returns the catalog of available models in the standard OpenAI list envelope, so OpenAI-compatible clients can auto-discover models. No authentication required.

Response

{
  "object": "list",
  "data": [
    {
      "id": "aion-labs/aion-3.5",
      "object": "model",
      "created": 1790121600,
      "name": "AionLabs: Aion 3.5",
      "description": "...",
      "date": "2026-09-23",
      "context_length": 262144,
      "max_completion_tokens": 32768,
      "is_moderated": false,
      "architecture": { "modality": "text->text" },
      "pricing": {
        "prompt": "0.000003",
        "completion": "0.000006",
        "input_cache_read": "0.00000075"
      },
      "reasoning": true,
      "tool_calls": true,
      "reasoning_effort": {
        "supported": true,
        "levels": ["low", "high", "max"],
        "default": "high"
      }
    }
  ],
  "models": [ ... ]
}

Each entry under data carries the standard OpenAI object and created (Unix seconds, midnight UTC of the release date) fields alongside our extended metadata. The legacy models key holds the same entries without those two fields and is kept for existing integrations — new clients should read data.

Pricing fields are per-token. input_cache_read is only present on models with cache pricing.

Per-model parameter capabilities

Each entry states which request parameters the model accepts, so a client or gateway can translate requests from the catalog instead of hard-coding model families. The values mirror what the API enforces at request time.

Field Type Description
reasoning boolean Whether the model reasons before answering (and so can return a reasoning field).
tool_calls boolean Whether the model accepts tools and can return tool_calls.
reasoning_effort.supported boolean Whether the model accepts the reasoning_effort parameter. When false, levels is empty and default is null; sending any value other than medium returns a 400.
reasoning_effort.levels array The exact reasoning_effort values the model accepts. Any other value returns a 400; values are never silently coerced.
reasoning_effort.default string What an omitted reasoning_effort means for this model. Always one of levels.
POST /v1/chat/completions

OpenAI-compatible chat completions. Supports streaming and tool calls.

Request body

Parameter Type Required Description
model string Yes Model ID. See Models.
messages array Yes Conversation history. See message object below.
temperature float No Sampling temperature.
max_tokens integer No Maximum completion tokens (≥ 1). This cap bounds the total tokens the model generates for the request — the model's internal reasoning and the visible answer share this budget, and reasoning is generated first. A cap that is exhausted during reasoning returns the reasoning only with finish_reason=length and empty content (the reasoning is available in reasoning when reasoning_split is on). Size max_tokens to cover reasoning plus the expected answer; it is a ceiling, not a reservation, so unused budget is not billed.
stop string[] No Stop sequences.
stream boolean No Stream response as SSE. Default: false. Non-streaming responses must complete within 120 seconds or the edge returns HTTP 524; generation still completes server-side and is billed. Use stream: true for requests that may run longer (large max_tokens, reasoning models) — streamed responses have no total-duration limit, and reasoning chunks keep the connection active before visible content starts. Very long prompts count against the same limit: prefill alone can take well over a minute for prompts in the hundreds of thousands of tokens, so send those with stream: true.
stream_options object No Streaming requests only. {"continuous_usage_stats": true} (vLLM convention) adds cumulative usage to every chunk — useful for billing partial output when a stream is cancelled mid-generation. include_usage is accepted for compatibility; the final chunk always carries usage.
tools array No Tool definitions in OpenAI format.
reasoning_effort Aion models string No Controls the reasoning level. Aion 3.5 / 3.5 Mini accept exactly low, high or max (default: high); reasoning cannot be disabled on these models, and any other value returns a 400. Aion 2.0 / 3.0 / 3.0 Mini accept none to disable reasoning, or low, medium, high or max to enable it (default: medium). The models are tuned to run with reasoning enabled; disabling it can degrade output. Not supported on aion-rp-llama-3.1-8b. Each model's accepted values and default are also published per entry in GET /v1/models (reasoning_effort.levels / reasoning_effort.default).
reasoning_split boolean No Split <think> reasoning into a separate reasoning field. Defaults on for reasoning models.
metadata object No Key-value pairs attached to the request: up to 16 keys, keys up to 64 characters, string values up to 512 characters (no nested objects).

Message object

Field Type Required Description
role string Yes system, user, assistant, or tool
content string or parts No Message text, or a list of {"type", "text"} content parts.
tool_calls array No Tool calls returned by the model.
tool_call_id string No ID of the tool call being answered (role tool messages).

Response (non-streaming)

{
  "id": "chatcmpl_abc123",
  "object": "chat.completion",
  "created": 1700000000,
  "model": "aion-labs/aion-2.0",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 8,
    "total_tokens": 20,
    "prompt_tokens_details": {
      "cached_tokens": 0
    }
  }
}

When reasoning_split is active, the message gains a reasoning field and content contains only the response text:

{
  "message": {
    "role": "assistant",
    "reasoning": "Let me think through this...",
    "content": "The answer is 42."
  }
}

Response (streaming)

When stream: true, the endpoint returns text/event-stream. Each event is a JSON delta, terminated by data: [DONE]:

data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" world"},"finish_reason":null}]}

data: [DONE]
POST /v1/responses

Alternative completion endpoint modelled on the OpenAI Responses API shape. Accepts the same model set and supports streaming.

Request body

Parameter Type Required Description
model string Yes Model ID.
input string or array One of Prompt string, or a list of {"role", "content"} message objects. Mutually exclusive with messages.
messages array One of Chat history in the same format as /v1/chat/completions. Mutually exclusive with input.
temperature float No Sampling temperature.
max_output_tokens integer No Maximum completion tokens (≥ 1).
stream boolean No Stream as SSE. Default: false. The same 120-second limit on non-streaming responses applies as for /v1/chat/completions.
reasoning_effort Aion models string No Controls the reasoning level. Aion 3.5 / 3.5 Mini accept exactly low, high or max (default: high); reasoning cannot be disabled on these models, and any other value returns a 400. Aion 2.0 / 3.0 / 3.0 Mini accept none to disable reasoning, or low, medium, high or max to enable it (default: medium). The models are tuned to run with reasoning enabled; disabling it can degrade output. Not supported on aion-rp-llama-3.1-8b. Each model's accepted values and default are also published per entry in GET /v1/models (reasoning_effort.levels / reasoning_effort.default).
reasoning_split boolean No Split reasoning into a separate content item.
metadata object No Key-value pairs attached to the request: up to 16 keys, keys up to 64 characters, string values up to 512 characters (no nested objects).

When using input as a list, each item's content may be a string or a list of {"type": "input_text", "text": "..."} parts.

Response (non-streaming)

{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1700000000,
  "model": "aion-labs/aion-2.0",
  "status": "completed",
  "output": [
    {
      "id": "msg_xyz",
      "type": "message",
      "status": "completed",
      "role": "assistant",
      "content": [
        { "type": "output_text", "text": "Hello! How can I help?" }
      ]
    }
  ],
  "usage": {
    "input_tokens": 12,
    "output_tokens": 8,
    "total_tokens": 20,
    "input_tokens_details": {
      "cached_tokens": 0
    }
  }
}

When reasoning_split is active, a reasoning item is prepended to content:

"content": [
  { "type": "reasoning", "text": "Let me think through this..." },
  { "type": "output_text", "text": "The answer is 42." }
]