Qwen 3.8‑27B for coding agents.
Fast. Metered. Prepaid.
An OpenAI‑compatible endpoint on H200s with prefix caching, so your agent's 200k‑token context costs cents instead of dollars.
- bf16 weights, 262k context
- cached input $0.015/M
- no subscription, no surprise bills
export OPENAI_BASE_URL="https://api.agenttokens.dev/v1"
export OPENAI_API_KEY="at_YOUR_KEY"
export OPENAI_MODEL="qwen3.8-27b"
Works with anything that speaks the OpenAI chat completions API.
Plug into your harness
Pick your tool, paste, replace at_YOUR_KEY. That's the whole setup.
Add a provider to ~/.openclaw/openclaw.json and point the default agent at it. (Or run openclaw onboard --custom and paste the same values.)
{
"models": {
"providers": {
"agenttokens": {
"baseUrl": "https://api.agenttokens.dev/v1",
"apiKey": "at_YOUR_KEY",
"api": "openai-completions",
"models": [
{ "id": "qwen3.8-27b", "name": "Qwen 3.8 27B", "contextWindow": 262144, "maxTokens": 32768 }
]
}
}
},
"agents": {
"defaults": {
"model": { "primary": "agenttokens/qwen3.8-27b" }
}
}
}
Add to ~/.config/opencode/opencode.json (or opencode.json in the project root). Then /models in OpenCode to pick it.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"agenttokens": {
"npm": "@ai-sdk/openai-compatible",
"name": "AgentTokens",
"options": {
"baseURL": "https://api.agenttokens.dev/v1",
"apiKey": "at_YOUR_KEY"
},
"models": {
"qwen3.8-27b": {
"name": "Qwen 3.8 27B",
"limit": { "context": 262144, "output": 32768 }
}
}
}
},
"model": "agenttokens/qwen3.8-27b"
}
Qwen Code reads the standard OpenAI env vars. Put these in your shell profile or in ~/.qwen/.env.
export OPENAI_BASE_URL="https://api.agenttokens.dev/v1"
export OPENAI_API_KEY="at_YOUR_KEY"
export OPENAI_MODEL="qwen3.8-27b"
qwen
In Cline: Settings (gear) → API Provider → choose OpenAI Compatible, then fill in:
API Provider: OpenAI Compatible
Base URL: https://api.agenttokens.dev/v1
API Key: at_YOUR_KEY
Model ID: qwen3.8-27b
Leave the other model settings at their defaults. Context window is 262,144 tokens.
Pricing
Per 1M tokens, USD. Billed per request against your prepaid balance. No minimums, no tiers.
| AgentTokens | Typical OpenRouter median* | |
|---|---|---|
| Input (uncached) | $0.12 | $0.24 |
| Input (cached) | $0.015 | $0.24 |
| Output | $1.50 | $2.20 |
* Median list price across OpenRouter providers for this model class at time of writing; most don't discount cached input. Check current prices yourself.
Worked example: one agent session
A long coding session re‑sends the same system prompt, repo map and history on every turn. Say it totals 9M cached input tokens, 1.2M uncached input tokens and 350k output tokens.
Totals are computed from the current price table above.
Start with $1.00 free Top up from $5. Credit never expires.
Why it's fast
-
H200s, bf16
Full‑precision weights on H200 (141 GB HBM3e). No 4‑bit quantization, no quality tax, plenty of room for KV cache.
-
Prefix caching
Repeated prompt prefixes — system prompt, tool schemas, repo map — are served from KV cache. Cached input bills at $0.015/M and skips most of the prefill.
-
Sticky routing
Requests from the same key land on the same replica while it's warm, so your agent's prefix stays hot turn after turn instead of re‑computing on a cold node.
-
Low time‑to‑first‑token
Cached prefill plus continuous batching tuned for agent traffic: many medium‑sized requests, not a few huge ones. Streaming starts fast.
FAQ
What model exactly?
Qwen 3.8‑27B, served as model id qwen3.8-27b. Dense 27B, bf16 weights, 262,144‑token context. We run the released weights; no fine‑tune, no system prompt injection.
Is it quantized?
No. Weights are served in bf16, the precision they were released in. What you get from the endpoint is what the model card describes.
Context length and max output?
262,144 tokens total (input + output). Set max_tokens as you normally would; the gateway rejects requests that exceed the window with a 400 rather than silently truncating.
What do you log or store?
Prompts and completions are processed in memory and not persisted beyond the lifetime of the request. We store usage counts per request (token counts, model, timestamp, cost) so we can bill you and you can see your usage. Prompt caches live in GPU memory on the serving node and are evicted within minutes. See the privacy page.
Rate limits and concurrency?
There are no per‑minute request or token quotas. Run as many parallel agents as you like; capacity is shared fairly under load. If we ever have to shed load you'll get a 429 with a Retry‑After header, not a silent stall.
Refunds?
Credits are prepaid and non‑refundable except where the law requires it. Balance never expires. If we billed you wrong, email hello@agenttokens.dev and we'll fix it.
SLA and status?
No formal SLA on the pay‑as‑you‑go tier. We aim for high availability but this is a small service; if you need a contractual SLA, email us. You're only charged for requests that complete.
Have idle H100/H200s?
We buy spare capacity from operators who can run our serving stack. Tell us what you have.