Help Articles
Six short guides to the things that actually go wrong. Every figure here is taken from the running service, not from a plan for one.
How to Generate Your First API Key
Everything you need to make your first request, and the one thing about keys that surprises people.
- Sign in at the Dashboard. New accounts get 5,000 tokens free, which is enough to try every model we serve without paying anything.
- Under API Keys, choose Create Key and
give it a name you will recognise later —
laptop,ci,demo-app. - Copy the key immediately. It is shown once and never again.
- Send your first request.
curl https://your-endpoint/api/v1/chat \
-H "Authorization: Bearer sk-tai-..." \
-H "Content-Type: application/json" \
-d '{
"model": "tfmf",
"messages": [{"role": "user", "content": "Hello"}]
}'
Why you cannot see the key again. We store a SHA-256 fingerprint and a short display prefix, not the key itself. That means a stolen database cannot be turned into working keys — but it also means a lost key is gone. If you lose one, revoke it and make another; it takes seconds.
Keeping keys safe
- Put the key in an environment variable, never in source code. Anything committed to git is public sooner or later.
- Use a separate key per application. When one leaks you revoke one thing, not everything.
- Revoke keys you no longer use.
last_used_atin the dashboard shows which are actually still calling.
Understanding Pay-As-You-Go Billing
How credits work, and why the free allowance does not expire.
Every account starts with 5,000 free tokens. The free allowance is consumed before any credit, so you are never charged while it remains.
What you pay for
Tokens, priced per million, with input and output priced separately. Output costs more because generating a token is strictly more work than reading one.
The live table is on the Pricing page, which is the source of truth. As a sense of scale: TFMF input is ¥0.5 per million tokens and output ¥3 per million, so a typical few-hundred-token exchange costs a fraction of a fen.
Topping up
Credit is added with redeem codes. Redeeming a code applies the balance instantly; codes can also carry a plan upgrade.
Why codes, and not a card form? Because this is a research preview run from a home server by a student team, and we are not going to hold card details for it. A code moves the money and the platform only ever sees "this account has ¥20 of credit" — there is no payment instrument stored here to leak.
Checking usage
The Dashboard breaks usage down per model,
per month. The same numbers are available programmatically at
GET /api/usage, which is useful for catching a runaway
integration before it matters.
Common API Error Codes and Solutions
What each failure means, and what to change.
| Status | Code | Usually means |
|---|---|---|
| 401 | invalid_api_key |
The key is wrong, revoked, or the Bearer prefix is
missing. Check for whitespace when copying from a file. |
| 401 | not_authenticated |
No credentials at all. The header is absent rather than malformed. |
| 402 | insufficient_balance |
Free allowance and credit are both exhausted. Redeem a code to continue. |
| 404 | model_not_found |
The model name is not in the catalogue, or is not live yet. Call
GET /api/v1/models for the current list rather than
hardcoding one. |
| 422 | invalid_request |
A required field is missing, empty or over its length limit. The
response names the offending field in param. |
| 429 | rate_limited |
Too many requests too quickly. Back off and retry; see the rate limit article below. |
| 502 | backend_unavailable |
The model server could not be reached. This is ours to fix, not yours — retry with backoff. |
| 503 | not_configured |
The requested capability is not enabled on this deployment. |
Reading the response body
Errors are always JSON with a stable code and a human
message. Branch on the code, never on the message text — the
wording is for people and may change.
{
"code": "invalid_request",
"message": "`max_output_tokens` must be between 1 and 4096",
"param": "max_output_tokens"
}
The 502 is worth planning for. This service is a
research preview and genuinely does go down. Treat
backend_unavailable as expected, not exceptional, and have a
fallback or a queued retry. See Privacy &
Terms for how much downtime to expect.
Best Practices for API Integration
Four things that separate an integration that holds up from one that does not.
1. Stream by default
Set "stream": true. A complete answer takes as long as it
takes; streamed, the first words arrive in well under a second and the
user reads while the model writes. The API speaks
text/event-stream with these events:
message.start— the request was accepted, with an id.message.delta— the next piece of text. Append it.message.done— finished, with final usage figures.
for (const line of buffer.split("\n")) {
if (!line.startsWith("data:")) continue;
const evt = JSON.parse(line.slice(5));
if (evt.delta) append(evt.delta); // append, never replace
}
Append, do not replace. Every event carries only the new text. Replacing the whole message with each delta produces output that flickers and loses everything before the last token.
2. Retry the right failures
Retry 429 and 502 with exponential backoff. Do
not retry 401, 402 or
422 — the request is wrong and will stay wrong. Retrying
those just burns your own quota.
3. Set a timeout
A generation can legitimately take a while, especially with reasoning enabled. But without a timeout, a dropped connection hangs a worker forever. Give streaming requests a generous read timeout and a short connect timeout.
4. Do not log full prompts by default
Prompts contain whatever your users typed. If you log them, you have created a copy of user data in a second place with none of the controls the first place had.
Token Usage and Cost Optimization
Where the tokens actually go, and which levers matter.
What counts
Both directions are billed and both are reported in
usage on every response:
prompt_tokens— your system prompt plus the entire conversation history resent on every call.completion_tokens— the answer, and any reasoning the model produced before it.
The single biggest surprise: history is resent every time. A chat that feels like twenty messages is billed as twenty requests, each carrying all the messages before it. Cost grows quadratically with conversation length unless you do something about it.
Levers that work
- Trim history. Send the last few turns plus a summary of anything older. This is the largest saving available in most integrations.
- Cap
max_output_tokens. It is a ceiling, not a target, but it bounds the worst case. - Turn reasoning off when you do not need it. Thinking is genuinely useful for multi-step problems and pure cost for greetings. It is off by default for exactly this reason.
- Use a smaller model for small jobs. Classification and routing rarely need the largest model you have.
- Shorten the system prompt. It is billed on every single request. A 500-token system prompt costs more over a thousand calls than the answers do.
Watching it
GET /api/usage returns month-to-date totals and a per-model
breakdown. Alert on it rather than discovering the number at the end of
the month.
Rate Limits and How to Handle Them
Why the limit exists, and the backoff that respects it.
Requests are limited per account. The limit is there less to ration capacity than to stop one runaway loop from starving everyone else on a machine that is, frankly, one computer in someone's home.
When you exceed it the API returns 429 with
rate_limited. It does not queue the request; you retry it.
Exponential backoff with jitter
delay = min(cap, base * 2 ** attempt)
sleep(delay / 2 + random() * delay / 2) // full jitter
The jitter matters. Without it, every client that was rejected at the same moment retries at the same moment, and the retry becomes the next spike. Spreading them out is what makes the backoff work.
Practical rules
- Start with
base = 0.5sand a cap around 30 seconds. - Limit total attempts — five or so. Beyond that the user is waiting for something that is not coming.
- Back off on
502with the same code path; an unavailable backend and a throttled one want the same response from you. - Batch and cache where you can. The cheapest request is the one you do not make.
Still stuck?
If an article did not cover it, tell us — that is how the next one gets written.
