TAI API Reference
The TAI Assistant API is our own shape, not a clone of anyone else's. It gives you
assistants, persistent threads, one-off chat and streaming, authenticated with an API
key. This page documents every endpoint, and every example runs against
tai-sdk.
| Base URL | Where it comes from |
|---|---|
https://YOUR_API_HOST |
The SDK finds this for you. Resolution order: base_url=
argument → TAI_BASE_URL env var →
~/.tai/config.json →
api-endpoint.json (cached 12 h) → a built-in default.
|
Because the API host is a home server behind a tunnel, its address can change. That is what api-endpoint.json is for: it is published here on GitHub Pages, and every SDK client re-reads it, so a move is a one-file change.
Authentication
Create a key in the Dashboard, then send it as a bearer token. The full key is shown once, at creation; we store only a SHA-256 fingerprint and a short display prefix.
curl https://YOUR_API_HOST/api/v1/models \ -H "Authorization: Bearer sk-tai-..."
Keys are scoped to your account. Revoking one is immediate, and a revoked key returns
401 invalid_api_key.
Install the SDK
pip install tai-sdk
Python 3.9+. The only runtime dependency is httpx.
The key can also come from the environment, so it never has to appear in code:
export TAI_API_KEY="sk-tai-..."
Quickstart
from tai_sdk import TAI
client = TAI() # reads TAI_API_KEY, discovers the API host
reply = client.chat.create(
model="tfmf",
messages=[{"role": "user", "content": "Explain LoRA in two sentences."}],
)
print(reply.content) # shortcut for reply.output["content"]
print(reply.usage.input_tokens, reply.usage.output_tokens, reply.usage.cost_cny)
Every response carries its own usage, so you can cost
a call without asking a second endpoint.
Chat
Stateless: you send the whole history every time. POST /api/v1/chat
| Field | Type | Notes |
|---|---|---|
model | string | Required unless assistant_id is given. |
assistant_id | string | Supplies the model and the system instructions. |
messages | array | 1–100 items of {role, content}; role is user, assistant or system. |
temperature | float | 0.0–2.0, default 0.7. |
max_output_tokens | int | 1–8192, default 512. |
stream | bool | See Streaming. |
The response:
{
"id": "chat_...",
"object": "chat.completion",
"model": "tfmf",
"output": { "role": "assistant", "content": "..." },
"usage": { "input_tokens": 12, "output_tokens": 40,
"cost_cny": 0.000126, "estimated": true },
"created_at": "2026-10-01T09:00:00+00:00"
}
Streaming
Pass stream=True and you get server-sent events. The
SDK turns them into typed objects, so you never parse SSE yourself:
| Event | SDK type | Carries |
|---|---|---|
message.start | MessageStart | id, model, thread_id (null for stateless chat) |
message.delta | MessageDelta | delta — one text fragment |
message.done | MessageDone | content and usage |
error | raises | the matching exception class |
from tai_sdk.types import MessageDelta, MessageDone
stream = client.chat.stream(
model="tfmf",
messages=[{"role": "user", "content": "Count to five."}],
)
for event in stream:
if isinstance(event, MessageDelta):
print(event.delta, end="", flush=True)
elif isinstance(event, MessageDone):
print("\n%d tokens" % event.usage.output_tokens)
Breaking out of the loop early is safe: the HTTP response is closed for you, and any text already generated is still recorded and billed.
Assistants
An assistant stores a model plus system instructions, so you reuse a configuration
instead of resending it. POST /api/v1/assistants
assistant = client.assistants.create(
model="tfmf", # required
name="Tutor",
instructions="You are a patient tutor. Answer briefly.",
metadata={"team": "education"}, # optional, flat strings
)
print(assistant.id) # asst_...
print(assistant.created_at)
| Method | SDK | Endpoint |
|---|---|---|
| Create | client.assistants.create(model=, name=, instructions=, metadata=) | POST /api/v1/assistants |
| Retrieve | client.assistants.get(id) | GET /api/v1/assistants/{id} |
| Update | client.assistants.update(id, instructions=...) | PATCH /api/v1/assistants/{id} |
| List | client.assistants.list(limit=, after=) | GET /api/v1/assistants |
| Delete | client.assistants.delete(id) | DELETE /api/v1/assistants/{id} |
Deleting an assistant detaches it from any threads rather than deleting them. A thread keeps its own model, so it stays usable.
Threads
A thread is a stored conversation. You never resend history — that is the whole
point. POST /api/v1/threads
thread = client.threads.create(
assistant_id=assistant.id, # or model="tfmf"
title="LoRA questions",
)
print(thread.id) # thrd_...
| Method | SDK |
|---|---|
| Create | client.threads.create(assistant_id=, model=, title=, metadata=) |
| Retrieve | client.threads.get(id) |
| Update | client.threads.update(id, title=, assistant_id=, model=, metadata=) |
| List | client.threads.list(limit=, after=) |
| Delete | client.threads.delete(id) |
Messages
One call appends your message and generates the reply, returning both so you
do not need a second round trip.
POST /api/v1/threads/{id}/messages
exchange = client.threads.messages.create(
thread.id,
content="What does rank mean in LoRA?",
temperature=0.3,
max_output_tokens=256,
)
print(exchange.user_message.content)
print(exchange.assistant_message.content)
print(exchange.usage.cost_cny)
Or stream it, exactly like chat:
for event in client.threads.messages.create(thread.id, content="Go on", stream=True):
...
Reading history back:
page = client.threads.messages.list(thread.id, limit=20, order="asc")
for message in page.data:
print(message.role, message.content[:60])
Pagination
Every list endpoint takes limit (1–100, default 20)
and after (an object id), returns newest first, and
reports whether more exist:
page = client.assistants.list(limit=2)
print(len(page.data), page.has_more)
while page.has_more:
page = client.assistants.list(limit=2, after=page.last_id)
print(len(page.data), page.has_more)
Models
| Model | Description | Input / Output per 1M | Status |
|---|---|---|---|
GTC-2.5 minigtc-2.5-mini | Lightweight 64M base model | ¥0.5 / ¥1 | Live |
TFMFtfmf | General foundation model on Qwen 3.5-4B (INT8) | ¥0.5 / ¥3 | Live |
Gigglegiggle | General companion model (TFMF + LoRA) | ¥1 / ¥5 | In training |
Brobro | Companion model (TFMF + LoRA) | ¥1 / ¥5 | In training |
Sweetiesweetie | Companion model (TFMF + LoRA) | ¥1 / ¥5 | In training |
Query the live list at runtime rather than hard-coding prices:
for model in client.models.list(): # iterable, or use .data
print(model.id, model.live, model.context_window,
model.input_cny_per_1m, model.output_cny_per_1m)
tfmf = client.models.retrieve("tfmf") # raises ModelNotFoundError if unknown
Using a model that exists but is still training returns
409 model_not_available.
Usage
u = client.usage.retrieve()
print(u.requests_this_month, u.tokens_this_month)
print(u.cost_cny_this_month, u.balance_cny)
print(u.free_tokens) # {"granted": 5000, "used": .., "remaining": ..}
print(u.by_model) # per-model request, token and cost breakdown
Billing & Codes
The API is prepaid: every new account starts with 5,000 free tokens, then draws on its balance. Usage is billed per token at the prices above.
When the free tokens are gone and the balance is empty, generation stops with
402 insufficient_balance, and the response carries the
real numbers:
{
"code": "insufficient_balance",
"message": "Your free tokens are used up and your balance is empty. ...",
"balance_cny": 0.0,
"free_tokens_remaining": 0
}
Redeem codes are the simplest way to top up. They are sold or handed out however is convenient, and redeemed on the Dashboard:
TAI-XXXX-XXXX-XXXX
Codes are not case sensitive and ignore spaces and dashes, so
tai - xxxx - xxxx - xxxx works too.
Errors
Every failure returns a stable code next to a
human-readable message. The SDK maps each code to a typed exception, so you never
match on strings:
| HTTP | code | SDK exception | Means |
|---|---|---|---|
| 401 | missing_api_key / invalid_api_key | AuthenticationError | No key, or unknown / revoked. |
| 402 | insufficient_balance | InsufficientBalanceError | Free tokens used up and balance empty. |
| 403 | account_disabled | PermissionDeniedError | Account switched off. |
| 404 | not_found | NotFoundError | No such object, or it belongs to another account. |
| 404 | model_not_found | ModelNotFoundError | Unknown model id. |
| 409 | model_not_available | ModelNotAvailableError | Exists but still training. |
| 422 | invalid_request | InvalidRequestError | Bad field; param names it. |
| 429 | rate_limited | RateLimitError | Slow down. |
| 502 | backend_unavailable | BackendUnavailableError | The model server failed or is unreachable. |
from tai_sdk import TAI, InsufficientBalanceError, ModelNotAvailableError
try:
client.chat.create(model="giggle", messages=[{"role": "user", "content": "hi"}])
except ModelNotAvailableError as exc:
print(exc.code, exc.message) # model_not_available ...
except InsufficientBalanceError as exc:
print(exc.balance_cny) # whatever the server reported
Transport problems raise APIConnectionError or
APITimeoutError, and both name the base URL they
tried, which is what you want when the API host moves.
Limits & Reliability
- Context: 8,192 tokens per request. Prompt plus completion must fit.
- Message size: 32,000 characters; 100 messages per request.
- Metadata: 16 keys, flat string values.
- Retries: the SDK retries connection errors and 429/5xx up to
max_retries(default 2) with exponential backoff, honouringRetry-After. POSTs are only retried when the request provably never left the client, so a retry cannot double-charge.
Per-key rate limits are not switched on yet, and neither is automatic charging for overage — the platform is a research preview. Balance is enforced.
SDK Reference
| Namespace | Methods |
|---|---|
client.chat | create(), stream() |
client.assistants | create(), get(), update(), list(), delete() |
client.threads | create(), get(), update(), list(), delete() |
client.threads.messages | create() (alias add()), list() |
client.models | list(), retrieve() |
client.usage | retrieve() |
Constructor options:
TAI(
api_key=None, # default: TAI_API_KEY
base_url=None, # default: discovered (see Introduction)
timeout=30.0,
max_retries=2,
discover=True, # set False to skip the discovery document
discovery_url=None, # default: TAI_DISCOVERY_URL
discovery_timeout=3.0,
)
Async Client
AsyncTAI exposes the same methods and response types;
only the call sites change.
import asyncio
from tai_sdk import AsyncTAI
async def main():
async with AsyncTAI() as client:
reply = await client.chat.create(
model="tfmf",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.content)
# Note: no `await` here. The async client returns an async generator,
# so it is consumed with `async for` directly.
stream = client.chat.stream(
model="tfmf",
messages=[{"role": "user", "content": "Count to three."}],
)
async for event in stream:
...
asyncio.run(main())
Next Steps
- Grab a key in the Dashboard and make your first call.
- See Pricing for the full price list and a cost calculator.
- Read the SDK README for the same reference as a single document.