Docs

Quick start

Point your base_url at TokenPP and call every listed model with a platform key. Bodies match the official protocols, so existing code usually only needs a new base_url and key.

Three steps

From sign-up to your first response: create a key, add credit, change the base_url.

  1. Create a platform key

    Sign in and create a key on the console API Keys page. Keys look like sk-tpp-…, are shown in full only once, and are listed by prefix afterwards. Store it somewhere safe right away.

  2. Add credit

    Billing is prepaid. Top-ups and redemption codes credit your balance immediately. With a zero balance, requests return 402.

  3. Point base_url at the platform

    OpenAI-compatible clients use the platform address plus /v1. Anthropic-compatible clients (Claude Code, Anthropic SDK) use the platform address alone and the SDK appends /v1/messages. The quick-start card on the console overview shows the actual address of your deployment and copies it in one click.

curl https://<platform-host>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-tpp-..." \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{ "role": "user", "content": "ping" }]
  }'

Common clients

Any client that lets you override base_url can be pointed here. These are the usual setups.

OpenAI SDK

Set base_url to the platform address plus /v1 and use a platform key. Works with the Python and Node SDKs and the official CLI.

Cursor

In Settings → Models, enable OpenAI API Key, paste a platform key, and set Override OpenAI Base URL to the platform address plus /v1.

Claude Code

Set ANTHROPIC_BASE_URL to the platform address without /v1 and ANTHROPIC_AUTH_TOKEN to a platform key.

Your own app

Any HTTP client that can change base_url and the Authorization header works. Keep the body in the official OpenAI or Anthropic shape.

Endpoints

The platform exposes both OpenAI and Anthropic compatible entry points and authenticates them the same way: Authorization: Bearer <platform key>. Anthropic clients may also use the x-api-key header.

EndpointProtocolNotes
POST/v1/chat/completionsOpenAIChat completions with SSE streaming
POST/v1/responsesOpenAIResponses API with SSE streaming
POST/v1/messagesAnthropicMessages API with SSE streaming
GET/v1/modelsOpenAILists sellable models with unit prices; models without a price are omitted

Streaming and timeouts

All three chat endpoints support SSE. Turn it on for long answers.

  • With stream: true the response is text/event-stream, frames arrive as data: {...}, and the stream ends with data: [DONE].
  • Frames are returned as they arrive. The whole response is not buffered; if the connection drops, it closes.
  • If no new data arrives for too long (90 seconds by default), the connection is closed as a timeout. Clients should handle an early close.

Billing and usage

You pay for the tokens actually used, with no subscription tiers.

  • Prices are quoted per 1M tokens in USD, with input, output and cache read/write priced separately. See the model pricing page.
  • Before a request is sent, the platform estimates a temporary hold from max_tokens (or max_output_tokens) and the output price, settles against the real usage afterwards, and releases the unused part of the hold.
  • Streaming chat requests automatically get stream_options.include_usage so the final frame carries usage, and that usage is what gets billed.
  • Every call is written to the usage log with model, token breakdown, charged amount, latency and status, all visible on the console usage page.
  • Once the balance runs out, requests return 402.

Error codes

Gateway errors follow the OpenAI error shape: error.message, error.type, error.code, error.param.

HTTPcodeMeaning
401invalid_api_keyKey missing, invalid or deleted
401key_disabledThis key is disabled
401key_expiredThis key has expired
403account_disabledThe account is disabled
402insufficient_balanceBalance too low to place the hold
402insufficient_quotaThis key has used up its quota
400invalid_bodyBody is not valid JSON, or model is missing
404model_not_foundModel does not exist or is not listed
429rate_limitedRPM, TPM or concurrency limit for this key was hit
503upstream_unavailableThe service is temporarily unavailable. Try again later.

FAQ

What exactly goes in base_url?

OpenAI-compatible clients use the platform address plus /v1; Anthropic-compatible clients use the platform address alone. The console overview copies the real address of your deployment for you.

Why do I get model not found?

The model is not listed, or has no price configured. Check the model pricing page for the slug.

What should I check on a 401?

Make sure the header is Authorization: Bearer <platform key> and that you are using a platform key starting with sk-tpp-. A disabled or expired key also returns 401.

My stream stopped halfway. What happened?

Usually the connection dropped, or no data arrived long enough to hit the idle timeout. Handle an early close and retry if needed. Usage already generated is still billed and logged.

How do I see what each call cost?

The console usage page lists token breakdown and charged amount per request, grouped by model and key.

Back to home