Help Articles

Six short guides to the things that actually go wrong. Every figure here is taken from the running service, not from a plan for one.

How to Generate Your First API Key

Everything you need to make your first request, and the one thing about keys that surprises people.

  1. Sign in at the Dashboard. New accounts get 5,000 tokens free, which is enough to try every model we serve without paying anything.
  2. Under API Keys, choose Create Key and give it a name you will recognise later — laptop, ci, demo-app.
  3. Copy the key immediately. It is shown once and never again.
  4. Send your first request.
curl https://your-endpoint/api/v1/chat \
  -H "Authorization: Bearer sk-tai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tfmf",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Why you cannot see the key again. We store a SHA-256 fingerprint and a short display prefix, not the key itself. That means a stolen database cannot be turned into working keys — but it also means a lost key is gone. If you lose one, revoke it and make another; it takes seconds.

Keeping keys safe

Understanding Pay-As-You-Go Billing

How credits work, and why the free allowance does not expire.

Every account starts with 5,000 free tokens. The free allowance is consumed before any credit, so you are never charged while it remains.

What you pay for

Tokens, priced per million, with input and output priced separately. Output costs more because generating a token is strictly more work than reading one.

The live table is on the Pricing page, which is the source of truth. As a sense of scale: TFMF input is ¥0.5 per million tokens and output ¥3 per million, so a typical few-hundred-token exchange costs a fraction of a fen.

Topping up

Credit is added with redeem codes. Redeeming a code applies the balance instantly; codes can also carry a plan upgrade.

Why codes, and not a card form? Because this is a research preview run from a home server by a student team, and we are not going to hold card details for it. A code moves the money and the platform only ever sees "this account has ¥20 of credit" — there is no payment instrument stored here to leak.

Checking usage

The Dashboard breaks usage down per model, per month. The same numbers are available programmatically at GET /api/usage, which is useful for catching a runaway integration before it matters.

Common API Error Codes and Solutions

What each failure means, and what to change.

StatusCodeUsually means
401invalid_api_key The key is wrong, revoked, or the Bearer prefix is missing. Check for whitespace when copying from a file.
401not_authenticated No credentials at all. The header is absent rather than malformed.
402insufficient_balance Free allowance and credit are both exhausted. Redeem a code to continue.
404model_not_found The model name is not in the catalogue, or is not live yet. Call GET /api/v1/models for the current list rather than hardcoding one.
422invalid_request A required field is missing, empty or over its length limit. The response names the offending field in param.
429rate_limited Too many requests too quickly. Back off and retry; see the rate limit article below.
502backend_unavailable The model server could not be reached. This is ours to fix, not yours — retry with backoff.
503not_configured The requested capability is not enabled on this deployment.

Reading the response body

Errors are always JSON with a stable code and a human message. Branch on the code, never on the message text — the wording is for people and may change.

{
  "code": "invalid_request",
  "message": "`max_output_tokens` must be between 1 and 4096",
  "param": "max_output_tokens"
}

The 502 is worth planning for. This service is a research preview and genuinely does go down. Treat backend_unavailable as expected, not exceptional, and have a fallback or a queued retry. See Privacy & Terms for how much downtime to expect.

Best Practices for API Integration

Four things that separate an integration that holds up from one that does not.

1. Stream by default

Set "stream": true. A complete answer takes as long as it takes; streamed, the first words arrive in well under a second and the user reads while the model writes. The API speaks text/event-stream with these events:

for (const line of buffer.split("\n")) {
  if (!line.startsWith("data:")) continue;
  const evt = JSON.parse(line.slice(5));
  if (evt.delta) append(evt.delta);   // append, never replace
}

Append, do not replace. Every event carries only the new text. Replacing the whole message with each delta produces output that flickers and loses everything before the last token.

2. Retry the right failures

Retry 429 and 502 with exponential backoff. Do not retry 401, 402 or 422 — the request is wrong and will stay wrong. Retrying those just burns your own quota.

3. Set a timeout

A generation can legitimately take a while, especially with reasoning enabled. But without a timeout, a dropped connection hangs a worker forever. Give streaming requests a generous read timeout and a short connect timeout.

4. Do not log full prompts by default

Prompts contain whatever your users typed. If you log them, you have created a copy of user data in a second place with none of the controls the first place had.

Token Usage and Cost Optimization

Where the tokens actually go, and which levers matter.

What counts

Both directions are billed and both are reported in usage on every response:

The single biggest surprise: history is resent every time. A chat that feels like twenty messages is billed as twenty requests, each carrying all the messages before it. Cost grows quadratically with conversation length unless you do something about it.

Levers that work

  1. Trim history. Send the last few turns plus a summary of anything older. This is the largest saving available in most integrations.
  2. Cap max_output_tokens. It is a ceiling, not a target, but it bounds the worst case.
  3. Turn reasoning off when you do not need it. Thinking is genuinely useful for multi-step problems and pure cost for greetings. It is off by default for exactly this reason.
  4. Use a smaller model for small jobs. Classification and routing rarely need the largest model you have.
  5. Shorten the system prompt. It is billed on every single request. A 500-token system prompt costs more over a thousand calls than the answers do.

Watching it

GET /api/usage returns month-to-date totals and a per-model breakdown. Alert on it rather than discovering the number at the end of the month.

Rate Limits and How to Handle Them

Why the limit exists, and the backoff that respects it.

Requests are limited per account. The limit is there less to ration capacity than to stop one runaway loop from starving everyone else on a machine that is, frankly, one computer in someone's home.

When you exceed it the API returns 429 with rate_limited. It does not queue the request; you retry it.

Exponential backoff with jitter

delay = min(cap, base * 2 ** attempt)
sleep(delay / 2 + random() * delay / 2)   // full jitter

The jitter matters. Without it, every client that was rejected at the same moment retries at the same moment, and the retry becomes the next spike. Spreading them out is what makes the backoff work.

Practical rules

Still stuck?

If an article did not cover it, tell us — that is how the next one gets written.