TAI API Reference

The TAI Assistant API is our own shape, not a clone of anyone else's. It gives you assistants, persistent threads, one-off chat and streaming, authenticated with an API key. This page documents every endpoint, and every example runs against tai-sdk.

Base URLWhere it comes from
https://YOUR_API_HOST The SDK finds this for you. Resolution order: base_url= argument → TAI_BASE_URL env var → ~/.tai/config.json → api-endpoint.json (cached 12 h) → a built-in default.

Because the API host is a home server behind a tunnel, its address can change. That is what api-endpoint.json is for: it is published here on GitHub Pages, and every SDK client re-reads it, so a move is a one-file change.

Authentication

Create a key in the Dashboard, then send it as a bearer token. The full key is shown once, at creation; we store only a SHA-256 fingerprint and a short display prefix.

curl https://YOUR_API_HOST/api/v1/models \
  -H "Authorization: Bearer sk-tai-..."

Keys are scoped to your account. Revoking one is immediate, and a revoked key returns 401 invalid_api_key.

Install the SDK

pip install tai-sdk

Python 3.9+. The only runtime dependency is httpx. The key can also come from the environment, so it never has to appear in code:

export TAI_API_KEY="sk-tai-..."

Quickstart

from tai_sdk import TAI

client = TAI()          # reads TAI_API_KEY, discovers the API host

reply = client.chat.create(
    model="tfmf",
    messages=[{"role": "user", "content": "Explain LoRA in two sentences."}],
)

print(reply.content)                    # shortcut for reply.output["content"]
print(reply.usage.input_tokens, reply.usage.output_tokens, reply.usage.cost_cny)

Every response carries its own usage, so you can cost a call without asking a second endpoint.

Chat

Stateless: you send the whole history every time. POST /api/v1/chat

FieldTypeNotes
modelstringRequired unless assistant_id is given.
assistant_idstringSupplies the model and the system instructions.
messagesarray1–100 items of {role, content}; role is user, assistant or system.
temperaturefloat0.0–2.0, default 0.7.
max_output_tokensint1–8192, default 512.
streamboolSee Streaming.

The response:

{
  "id": "chat_...",
  "object": "chat.completion",
  "model": "tfmf",
  "output":    { "role": "assistant", "content": "..." },
  "usage":     { "input_tokens": 12, "output_tokens": 40,
                 "cost_cny": 0.000126, "estimated": true },
  "created_at": "2026-10-01T09:00:00+00:00"
}

Streaming

Pass stream=True and you get server-sent events. The SDK turns them into typed objects, so you never parse SSE yourself:

EventSDK typeCarries
message.startMessageStartid, model, thread_id (null for stateless chat)
message.deltaMessageDeltadelta — one text fragment
message.doneMessageDonecontent and usage
errorraisesthe matching exception class
from tai_sdk.types import MessageDelta, MessageDone

stream = client.chat.stream(
    model="tfmf",
    messages=[{"role": "user", "content": "Count to five."}],
)

for event in stream:
    if isinstance(event, MessageDelta):
        print(event.delta, end="", flush=True)
    elif isinstance(event, MessageDone):
        print("\n%d tokens" % event.usage.output_tokens)

Breaking out of the loop early is safe: the HTTP response is closed for you, and any text already generated is still recorded and billed.

Assistants

An assistant stores a model plus system instructions, so you reuse a configuration instead of resending it. POST /api/v1/assistants

assistant = client.assistants.create(
    model="tfmf",                                  # required
    name="Tutor",
    instructions="You are a patient tutor. Answer briefly.",
    metadata={"team": "education"},                # optional, flat strings
)

print(assistant.id)          # asst_...
print(assistant.created_at)
MethodSDKEndpoint
Createclient.assistants.create(model=, name=, instructions=, metadata=)POST /api/v1/assistants
Retrieveclient.assistants.get(id)GET /api/v1/assistants/{id}
Updateclient.assistants.update(id, instructions=...)PATCH /api/v1/assistants/{id}
Listclient.assistants.list(limit=, after=)GET /api/v1/assistants
Deleteclient.assistants.delete(id)DELETE /api/v1/assistants/{id}

Deleting an assistant detaches it from any threads rather than deleting them. A thread keeps its own model, so it stays usable.

Threads

A thread is a stored conversation. You never resend history — that is the whole point. POST /api/v1/threads

thread = client.threads.create(
    assistant_id=assistant.id,     # or model="tfmf"
    title="LoRA questions",
)
print(thread.id)                   # thrd_...
MethodSDK
Createclient.threads.create(assistant_id=, model=, title=, metadata=)
Retrieveclient.threads.get(id)
Updateclient.threads.update(id, title=, assistant_id=, model=, metadata=)
Listclient.threads.list(limit=, after=)
Deleteclient.threads.delete(id)

Messages

One call appends your message and generates the reply, returning both so you do not need a second round trip. POST /api/v1/threads/{id}/messages

exchange = client.threads.messages.create(
    thread.id,
    content="What does rank mean in LoRA?",
    temperature=0.3,
    max_output_tokens=256,
)

print(exchange.user_message.content)
print(exchange.assistant_message.content)
print(exchange.usage.cost_cny)

Or stream it, exactly like chat:

for event in client.threads.messages.create(thread.id, content="Go on", stream=True):
    ...

Reading history back:

page = client.threads.messages.list(thread.id, limit=20, order="asc")
for message in page.data:
    print(message.role, message.content[:60])

Pagination

Every list endpoint takes limit (1–100, default 20) and after (an object id), returns newest first, and reports whether more exist:

page = client.assistants.list(limit=2)
print(len(page.data), page.has_more)

while page.has_more:
    page = client.assistants.list(limit=2, after=page.last_id)
    print(len(page.data), page.has_more)

Models

ModelDescriptionInput / Output per 1MStatus
GTC-2.5 mini
gtc-2.5-mini
Lightweight 64M base model¥0.5 / ¥1Live
TFMF
tfmf
General foundation model on Qwen 3.5-4B (INT8)¥0.5 / ¥3Live
Giggle
giggle
General companion model (TFMF + LoRA)¥1 / ¥5In training
Bro
bro
Companion model (TFMF + LoRA)¥1 / ¥5In training
Sweetie
sweetie
Companion model (TFMF + LoRA)¥1 / ¥5In training

Query the live list at runtime rather than hard-coding prices:

for model in client.models.list():          # iterable, or use .data
    print(model.id, model.live, model.context_window,
          model.input_cny_per_1m, model.output_cny_per_1m)

tfmf = client.models.retrieve("tfmf")      # raises ModelNotFoundError if unknown

Using a model that exists but is still training returns 409 model_not_available.

Usage

u = client.usage.retrieve()

print(u.requests_this_month, u.tokens_this_month)
print(u.cost_cny_this_month, u.balance_cny)
print(u.free_tokens)          # {"granted": 5000, "used": .., "remaining": ..}
print(u.by_model)             # per-model request, token and cost breakdown

Billing & Codes

The API is prepaid: every new account starts with 5,000 free tokens, then draws on its balance. Usage is billed per token at the prices above.

When the free tokens are gone and the balance is empty, generation stops with 402 insufficient_balance, and the response carries the real numbers:

{
  "code": "insufficient_balance",
  "message": "Your free tokens are used up and your balance is empty. ...",
  "balance_cny": 0.0,
  "free_tokens_remaining": 0
}

Redeem codes are the simplest way to top up. They are sold or handed out however is convenient, and redeemed on the Dashboard:

TAI-XXXX-XXXX-XXXX

Codes are not case sensitive and ignore spaces and dashes, so tai - xxxx - xxxx - xxxx works too.

Errors

Every failure returns a stable code next to a human-readable message. The SDK maps each code to a typed exception, so you never match on strings:

HTTPcodeSDK exceptionMeans
401missing_api_key / invalid_api_keyAuthenticationErrorNo key, or unknown / revoked.
402insufficient_balanceInsufficientBalanceErrorFree tokens used up and balance empty.
403account_disabledPermissionDeniedErrorAccount switched off.
404not_foundNotFoundErrorNo such object, or it belongs to another account.
404model_not_foundModelNotFoundErrorUnknown model id.
409model_not_availableModelNotAvailableErrorExists but still training.
422invalid_requestInvalidRequestErrorBad field; param names it.
429rate_limitedRateLimitErrorSlow down.
502backend_unavailableBackendUnavailableErrorThe model server failed or is unreachable.
from tai_sdk import TAI, InsufficientBalanceError, ModelNotAvailableError

try:
    client.chat.create(model="giggle", messages=[{"role": "user", "content": "hi"}])
except ModelNotAvailableError as exc:
    print(exc.code, exc.message)          # model_not_available ...

except InsufficientBalanceError as exc:
    print(exc.balance_cny)                # whatever the server reported

Transport problems raise APIConnectionError or APITimeoutError, and both name the base URL they tried, which is what you want when the API host moves.

Limits & Reliability

  • Context: 8,192 tokens per request. Prompt plus completion must fit.
  • Message size: 32,000 characters; 100 messages per request.
  • Metadata: 16 keys, flat string values.
  • Retries: the SDK retries connection errors and 429/5xx up to max_retries (default 2) with exponential backoff, honouring Retry-After. POSTs are only retried when the request provably never left the client, so a retry cannot double-charge.

Per-key rate limits are not switched on yet, and neither is automatic charging for overage — the platform is a research preview. Balance is enforced.

SDK Reference

NamespaceMethods
client.chatcreate(), stream()
client.assistantscreate(), get(), update(), list(), delete()
client.threadscreate(), get(), update(), list(), delete()
client.threads.messagescreate() (alias add()), list()
client.modelslist(), retrieve()
client.usageretrieve()

Constructor options:

TAI(
    api_key=None,            # default: TAI_API_KEY
    base_url=None,           # default: discovered (see Introduction)
    timeout=30.0,
    max_retries=2,
    discover=True,           # set False to skip the discovery document
    discovery_url=None,      # default: TAI_DISCOVERY_URL
    discovery_timeout=3.0,
)

Async Client

AsyncTAI exposes the same methods and response types; only the call sites change.

import asyncio
from tai_sdk import AsyncTAI

async def main():
    async with AsyncTAI() as client:
        reply = await client.chat.create(
            model="tfmf",
            messages=[{"role": "user", "content": "Hello!"}],
        )
        print(reply.content)

        # Note: no `await` here. The async client returns an async generator,
        # so it is consumed with `async for` directly.
        stream = client.chat.stream(
            model="tfmf",
            messages=[{"role": "user", "content": "Count to three."}],
        )
        async for event in stream:
            ...

asyncio.run(main())

Next Steps

  • Grab a key in the Dashboard and make your first call.
  • See Pricing for the full price list and a cost calculator.
  • Read the SDK README for the same reference as a single document.