Docs
The API is OpenAI‑compatible. If your client can talk to OpenAI, it can talk to us: change the base URL and the key.
| Base URL | https://api.agenttokens.dev/v1 |
|---|---|
| Model id | qwen3.8-27b |
| Auth | Authorization: Bearer at_… |
| Endpoints | POST /chat/completions (streaming and non‑streaming), GET /models |
Quickstart
Create a key on the dashboard (new accounts get $1.00 of credit), then:
curl
curl https://api.agenttokens.dev/v1/chat/completions \
-H "Authorization: Bearer at_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [
{"role": "system", "content": "You are a terse coding assistant."},
{"role": "user", "content": "Write a bash one-liner that counts lines of Go in this repo."}
],
"stream": false
}'
Python
pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://api.agenttokens.dev/v1",
api_key="at_YOUR_KEY",
)
stream = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Explain prefix caching in two sentences."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
Node
npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.agenttokens.dev/v1",
apiKey: process.env.OPENAI_API_KEY, // at_...
});
const res = await client.chat.completions.create({
model: "qwen3.8-27b",
messages: [{ role: "user", content: "Explain prefix caching in two sentences." }],
});
console.log(res.choices[0].message.content);
console.log(res.usage); // prompt_tokens, completion_tokens, prompt_tokens_details.cached_tokens
Environment variables
Most CLIs and SDKs pick these up automatically:
export OPENAI_BASE_URL="https://api.agenttokens.dev/v1"
export OPENAI_API_KEY="at_YOUR_KEY"
export OPENAI_MODEL="qwen3.8-27b"
Agent harnesses
Copy‑paste configs for the tools we see most. Replace at_YOUR_KEY.
Add a provider to ~/.openclaw/openclaw.json and point the default agent at it. (Or run openclaw onboard --custom and paste the same values.)
{
"models": {
"providers": {
"agenttokens": {
"baseUrl": "https://api.agenttokens.dev/v1",
"apiKey": "at_YOUR_KEY",
"api": "openai-completions",
"models": [
{ "id": "qwen3.8-27b", "name": "Qwen 3.8 27B", "contextWindow": 262144, "maxTokens": 32768 }
]
}
}
},
"agents": {
"defaults": {
"model": { "primary": "agenttokens/qwen3.8-27b" }
}
}
}
Add to ~/.config/opencode/opencode.json (or opencode.json in the project root). Then /models in OpenCode to pick it.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"agenttokens": {
"npm": "@ai-sdk/openai-compatible",
"name": "AgentTokens",
"options": {
"baseURL": "https://api.agenttokens.dev/v1",
"apiKey": "at_YOUR_KEY"
},
"models": {
"qwen3.8-27b": {
"name": "Qwen 3.8 27B",
"limit": { "context": 262144, "output": 32768 }
}
}
}
},
"model": "agenttokens/qwen3.8-27b"
}
Qwen Code reads the standard OpenAI env vars. Put these in your shell profile or in ~/.qwen/.env.
export OPENAI_BASE_URL="https://api.agenttokens.dev/v1"
export OPENAI_API_KEY="at_YOUR_KEY"
export OPENAI_MODEL="qwen3.8-27b"
qwen
In Cline: Settings (gear) → API Provider → choose OpenAI Compatible, then fill in:
API Provider: OpenAI Compatible
Base URL: https://api.agenttokens.dev/v1
API Key: at_YOUR_KEY
Model ID: qwen3.8-27b
Leave the other model settings at their defaults. Context window is 262,144 tokens.
Billing & keys
Prepaid credit
You top up a balance (Stripe, one‑time payments, from $5) and each request deducts its cost when it completes. There is no subscription and nothing that can charge your card without you clicking a button. Credit never expires.
Cost per request is computed from the token counts the gateway reports:
cost = uncached_input × $0.12/M
+ cached_input × $0.015/M
+ output × $1.50/M
Cached tokens are the part of your prompt whose prefix was already in KV cache on the node that served you. The response usage object reports them under prompt_tokens_details.cached_tokens; prompt_tokens is the total including cached.
Keys
- Format:
at_followed by 40 URL‑safe characters. Send it asAuthorization: Bearer at_…. - Shown once at creation. We store only a SHA‑256 hash; if you lose a key, revoke it and create another.
- Up to 10 active keys per account. Use one per machine or harness so you can revoke individually.
- Revoking takes effect immediately for new requests. In‑flight requests finish.
When the balance hits zero
Requests are rejected with HTTP 402 and an OpenAI‑style error body (see Errors). In‑flight requests are allowed to finish, so a single long completion can take your balance slightly below $0; that's on us, not you. Top up and requests resume immediately — no key change needed.
Headers
| Header | Direction | Notes |
|---|---|---|
Authorization | request | Bearer at_…. Required. |
Content-Type | request | application/json. |
Retry-After | response | Present on 429 / 503. Seconds to wait before retrying. |
X-Request-Id | response | Quote this when reporting a problem. |
Limits
| Context window | 262,144 tokens, input + output combined. |
|---|---|
| Rate limits | None per minute or per day. Concurrency is shared fairly; under heavy load you may see 429s with Retry-After. |
| Request body | Up to 8 MB of JSON. |
| Streaming | Supported ("stream": true, SSE, same format as OpenAI). Use it for agents — time‑to‑first‑token is the metric that matters. |
| Tools | OpenAI‑style tools / tool_choice function calling is supported. Tool schemas count as prompt tokens and cache well. |
| Keys | 10 active per account. |
| Top‑ups | $5–$2,000 per payment. |
Errors
Errors use the OpenAI shape so existing client retry logic keeps working:
{
"error": {
"message": "Insufficient credit. Top up at https://agenttokens.dev/dashboard",
"type": "insufficient_credit",
"code": 402
}
}
| Status | Meaning | What to do |
|---|---|---|
400 | Malformed request, unknown model, or prompt + max_tokens exceeds the context window. | Fix the request. Not retryable. |
401 | Missing, malformed, or revoked key. | Check the key. Not retryable. |
402 | Balance is $0 or below. | Top up on the dashboard, then retry. |
429 | Temporarily shedding load. | Back off for Retry-After seconds and retry. |
500 / 503 | Something broke on our side. | Retry with exponential backoff. You're not charged for failed requests. |
Stuck? Email hello@agenttokens.dev with the X-Request-Id.