Docs

The API is OpenAI‑compatible. If your client can talk to OpenAI, it can talk to us: change the base URL and the key.

Base URLhttps://api.agenttokens.dev/v1
Model idqwen3.8-27b
AuthAuthorization: Bearer at_…
EndpointsPOST /chat/completions (streaming and non‑streaming), GET /models

Quickstart

Create a key on the dashboard (new accounts get $1.00 of credit), then:

curl

curl https://api.agenttokens.dev/v1/chat/completions \
  -H "Authorization: Bearer at_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [
      {"role": "system", "content": "You are a terse coding assistant."},
      {"role": "user", "content": "Write a bash one-liner that counts lines of Go in this repo."}
    ],
    "stream": false
  }'

Python

pip install openai

from openai import OpenAI

client = OpenAI(
    base_url="https://api.agenttokens.dev/v1",
    api_key="at_YOUR_KEY",
)

stream = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Explain prefix caching in two sentences."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

Node

npm install openai

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.agenttokens.dev/v1",
  apiKey: process.env.OPENAI_API_KEY, // at_...
});

const res = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Explain prefix caching in two sentences." }],
});
console.log(res.choices[0].message.content);
console.log(res.usage); // prompt_tokens, completion_tokens, prompt_tokens_details.cached_tokens

Environment variables

Most CLIs and SDKs pick these up automatically:

export OPENAI_BASE_URL="https://api.agenttokens.dev/v1"
export OPENAI_API_KEY="at_YOUR_KEY"
export OPENAI_MODEL="qwen3.8-27b"

Agent harnesses

Copy‑paste configs for the tools we see most. Replace at_YOUR_KEY.

Add a provider to ~/.openclaw/openclaw.json and point the default agent at it. (Or run openclaw onboard --custom and paste the same values.)

{
  "models": {
    "providers": {
      "agenttokens": {
        "baseUrl": "https://api.agenttokens.dev/v1",
        "apiKey": "at_YOUR_KEY",
        "api": "openai-completions",
        "models": [
          { "id": "qwen3.8-27b", "name": "Qwen 3.8 27B", "contextWindow": 262144, "maxTokens": 32768 }
        ]
      }
    }
  },
  "agents": {
    "defaults": {
      "model": { "primary": "agenttokens/qwen3.8-27b" }
    }
  }
}

Billing & keys

Prepaid credit

You top up a balance (Stripe, one‑time payments, from $5) and each request deducts its cost when it completes. There is no subscription and nothing that can charge your card without you clicking a button. Credit never expires.

Cost per request is computed from the token counts the gateway reports:

cost = uncached_input × $0.12/M
     + cached_input   × $0.015/M
     + output         × $1.50/M

Cached tokens are the part of your prompt whose prefix was already in KV cache on the node that served you. The response usage object reports them under prompt_tokens_details.cached_tokens; prompt_tokens is the total including cached.

Keys

  • Format: at_ followed by 40 URL‑safe characters. Send it as Authorization: Bearer at_….
  • Shown once at creation. We store only a SHA‑256 hash; if you lose a key, revoke it and create another.
  • Up to 10 active keys per account. Use one per machine or harness so you can revoke individually.
  • Revoking takes effect immediately for new requests. In‑flight requests finish.

When the balance hits zero

Requests are rejected with HTTP 402 and an OpenAI‑style error body (see Errors). In‑flight requests are allowed to finish, so a single long completion can take your balance slightly below $0; that's on us, not you. Top up and requests resume immediately — no key change needed.

Headers

request
HeaderDirectionNotes
AuthorizationrequestBearer at_…. Required.
Content-Typeapplication/json.
Retry-AfterresponsePresent on 429 / 503. Seconds to wait before retrying.
X-Request-IdresponseQuote this when reporting a problem.

Limits

Context window262,144 tokens, input + output combined.
Rate limitsNone per minute or per day. Concurrency is shared fairly; under heavy load you may see 429s with Retry-After.
Request bodyUp to 8 MB of JSON.
StreamingSupported ("stream": true, SSE, same format as OpenAI). Use it for agents — time‑to‑first‑token is the metric that matters.
ToolsOpenAI‑style tools / tool_choice function calling is supported. Tool schemas count as prompt tokens and cache well.
Keys10 active per account.
Top‑ups$5–$2,000 per payment.

Errors

Errors use the OpenAI shape so existing client retry logic keeps working:

{
  "error": {
    "message": "Insufficient credit. Top up at https://agenttokens.dev/dashboard",
    "type": "insufficient_credit",
    "code": 402
  }
}
StatusMeaningWhat to do
400Malformed request, unknown model, or prompt + max_tokens exceeds the context window.Fix the request. Not retryable.
401Missing, malformed, or revoked key.Check the key. Not retryable.
402Balance is $0 or below.Top up on the dashboard, then retry.
429Temporarily shedding load.Back off for Retry-After seconds and retry.
500 / 503Something broke on our side.Retry with exponential backoff. You're not charged for failed requests.

Stuck? Email hello@agenttokens.dev with the X-Request-Id.