New Sign up free, 10 calls on us. Up to $1, no card needed.

Billing API

Pull usage and cost data for your own workspace into your own systems, without opening the console. A billing key is a read-only credential scoped to one workspace - it can read billing data, and nothing else.

⚠

A billing key can read every request logged in its workspace, including keys that have since been deleted or revoked, but it cannot call any model and cannot create, modify, or delete API keys. Using it anywhere else returns 401 wrong_key_kind.

Creating a billing key

  1. Open Console → Billing API keys (workspace owner only). Name it, optionally restrict it to an IP allow-list or set an expiry, then copy the key - it starts with sk-syn-bill- and is shown only once.
  2. Call the endpoints below with it as a Bearer token. Revoke it at any time from the same page - revoking immediately invalidates it, and nothing else in the workspace is affected.
  3. A key is tied to its creator: if they stop being the workspace owner, it stops working and answers 403 key_owner_not_authorized.

Authentication

Every call needs Authorization: Bearer sk-syn-bill-... against the base URL below. A billing key only works on these endpoints, and these endpoints only accept a billing key - any other combination returns 401 wrong_key_kind.

Authorization: Bearer sk-syn-bill-...

Base URL: https://synthorai.io/api/v1/billing

ℹ

Scope is the whole workspace: every inference key in it, including deleted ones - their historical rows stay queryable. api_key_id and model filters can only narrow the result, never widen it; every query is anchored on the key's own workspace.

ℹ

Every query runs on the read replica, never the primary. A process with no replica configured answers 503 replica_unavailable rather than silently falling back.

✓

Token fields: follow the OpenAI convention - prompt_tokens is total input tokens and already includes every cache read and cache write below it, and completion_tokens already includes reasoning_tokens. Neither is additive on top of the totals - they are a breakdown, not extra tokens.

The cache-read portion of prompt_tokens is cached_tokens; the cache-write portion is cache_write_tokens, split further into cache_write_5m_tokens and cache_write_1h_tokens on /records rows by TTL.

✓

Money fields: cost_usd is the amount actually charged. byok_list_price_usd is what a BYOK request would have cost at list price - 0 for a non-BYOK request. There is no list_price_usd anywhere in this API.

ℹ

"Error" here means any request that reached the API and was rejected - a 4xx such as 402 (insufficient balance), 403, or 429, as well as an upstream failure. Only our own internal infrastructure rejections are excluded, exactly as in the console. Rejections are recorded at most once per key, status code and minute on each server, so counts of 402, 403 and 429 are a lower bound; successful requests and charges are complete.

GET /usage - aggregated usage

GET /usage

Pre-aggregated rows bucketed by hour or day - built for invoices and dashboards. They come from the hourly rollup, plus the requests written since its last ingest, so recent buckets do not lag behind by a whole ingest cycle.

ParameterTypeDescription
start*stringStart of the range: an RFC 3339 timestamp or a date such as 2026-09-01.
endstringEnd of the range. Defaults to now.
granularitystringBucket size: hour or day. Defaults to day.
group_bystringDimensions to break rows out by: api_key, model, or both (any subset). Defaults to both.
api_key_idintegerFilter to one or more API key ids. Repeatable; narrows only, never widens.
modelstringFilter to one or more model ids. Repeatable.
limitintegerRows per page. Defaults to 1000, capped at 5000.
cursorstringOpaque pagination cursor, taken from the previous page's meta.next_cursor.

In each row, api_key_id and api_key_name appear only when group_by includes api_key, and model only when it includes model. avg_latency_ms can be null when a bucket has no latency samples.

Example Request

curl "https://synthorai.io/api/v1/billing/usage?start=2026-09-01&end=2026-09-24&granularity=day" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{
  "data": [
    {
      "bucket_start": "2026-09-23T00:00:00Z",
      "bucket_end": "2026-09-24T00:00:00Z",
      "api_key_id": 42,
      "api_key_name": "prod-backend",
      "model": "claude-sonnet-5",
      "requests": 1280,
      "success_requests": 1265,
      "error_requests": 15,
      "errors_4xx": 12,
      "errors_429": 2,
      "errors_5xx": 1,
      "prompt_tokens": 512000,
      "completion_tokens": 98000,
      "cached_tokens": 210000,
      "cache_write_tokens": 8000,
      "reasoning_tokens": 4000,
      "cost_usd": 4.1732,
      "byok_list_price_usd": 0,
      "byok_requests": 0,
      "avg_latency_ms": 812,
      "final": true
    }
  ],
  "meta": {
    "generated_at": "2026-09-24T08:00:00Z",
    "data_complete_until": "2026-09-24T06:00:00Z",
    "timezone": "UTC",
    "currency": "USD",
    "next_cursor": null,
    "has_more": false
  }
}

GET /records - per-request rows

GET /records

One row per request, meant for incremental sync into your own database. Sync by cursor, never by time, and dedupe incoming rows by record_id - see "Avoiding time gaps" below for why.

ParameterTypeDescription
cursorstringOpaque sync cursor (string), taken from the previous call's meta.next_cursor. Give exactly one of cursor or start; it has no range cap.
startstringStart of the range (RFC 3339 timestamp or date). Give exactly one of cursor or start; a start/end range is capped at 31 days, while cursor has no range cap.
endstringEnd of the range. Defaults to now.
api_key_idintegerFilter to one or more API key ids. Repeatable; narrows only, never widens.
modelstringFilter to one or more model ids. Repeatable.
limitintegerRows per page. Defaults to 500, capped at 1000.

Example Request

curl "https://synthorai.io/api/v1/billing/records?start=2026-09-24T00:00:00Z&limit=500" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{
  "data": [
    {
      "record_id": "rec_8f2a91c4",
      "request_id": "req_9f3c2b1a",
      "recorded_at": "2026-09-24T05:58:31Z",
      "completed_at": "2026-09-24T05:58:29Z",
      "api_key_id": 42,
      "api_key_name": "prod-backend",
      "model": "claude-sonnet-5",
      "status": "success",
      "status_code": 200,
      "prompt_tokens": 812,
      "completion_tokens": 194,
      "cached_tokens": 512,
      "cache_write_tokens": 64,
      "cache_write_5m_tokens": 64,
      "cache_write_1h_tokens": 0,
      "cost_usd": 0.00412,
      "byok_list_price_usd": 0,
      "is_byok": false,
      "is_stream": true,
      "duration_ms": 1340,
      "ttft_ms": 210
    }
  ],
  "meta": {
    "next_cursor": "eyJvIjoiODgxNDA5MCJ9",
    "has_more": false,
    "window_final": true,
    "safe_until": "2026-09-24T05:58:31Z",
    "latest_recorded_at": "2026-09-24T05:58:31Z",
    "data_complete_until": "2026-09-24T05:58:00Z"
  }
}
ParameterDescription
next_cursorCursor for the next call. Always present.
has_moreWhether calling again right now with this cursor would return more rows.
window_finaltrue only when you gave an end and it is at or before safe_until - the range you asked for is then guaranteed complete. Absent or false otherwise.
safe_untilThe newest recorded_at a caller can trust as settled; rows younger than this are still being held back.
latest_recorded_atThe most recent recorded_at across all rows, settled or not.
data_complete_untilSame value /freshness returns - how far the hourly rollup is complete, for cross-referencing against /usage.
⚠

Time basis: /usage buckets by completed_at (when a request finished). /records filters start/end by recorded_at (when the row was written, which can lag completion under load). To reconcile /records against /usage, bucket your synced rows by completed_at yourself, not recorded_at.

GET /records/export - streamed export

GET /records/export

Same filters as /records, streamed as newline-delimited JSON (application/x-ndjson) - built for a large backfill instead of a paged list. Only one export can stream per workspace at a time; a second one gets 429 too_many_concurrent_requests.

Example Request

curl -N "https://synthorai.io/api/v1/billing/records/export?cursor=eyJvIjoiODgxNDAzMiJ9" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{"record_id":"rec_8f2a91c4","request_id":"req_9f3c2b1a","model":"claude-sonnet-5","status":"success","cost_usd":0.00412}
{"record_id":"rec_8f2a91c5","request_id":"req_9f3c2b1b","model":"claude-sonnet-5","status":"success","cost_usd":0.00398}
{"_meta":{"next_cursor":"eyJvIjoiODgxNDA5MSJ9","rows":2,"complete":true,"window_final":true,"generated_at":"2026-09-24T06:00:02Z"}}

The last line of the stream is a _meta object:

ParameterDescription
next_cursorCursor for the next export call.
rowsNumber of records written in this stream.
completefalse means the stream stopped before exhausting the range - resume with cursor=next_cursor.
window_finalSame meaning as on /records: true only when the export's end is at or before safe_until.
generated_atWhen this line was written.
stopped_becausePresent when complete is false: row_cap, scan_cap (the stream walked its maximum id range without filling up - common with a narrow filter), time_cap, query_error, or window_not_final.

GET /summary - preliminary statistics and anomalies

GET /summary

Totals, top models and top keys by cost, and an error rate for the range, plus an anomalies list (high_error_rate, rate_limited, data_not_final, rollup_delayed), each carrying a severity of info or warning and a human-readable message. Numbers in an unfinished range can still move - see below.

ParameterTypeDescription
startstringStart of the range. Defaults to 30 days before end.
endstringEnd of the range. Defaults to now.
api_key_idintegerFilter to one or more API key ids. Repeatable; narrows only, never widens.
modelstringFilter to one or more model ids. Repeatable.

models_count and keys_count give the totals behind top_models and top_keys, which list at most the top 20 by cost. The range defaults to the 30 days before end when start is omitted.

Example Request

curl "https://synthorai.io/api/v1/billing/summary?start=2026-09-17&end=2026-09-24" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{
  "data": {
    "totals": { "requests": 48210, "cost_usd": 132.55, "error_rate": 0.021 },
    "models_count": 6,
    "keys_count": 3,
    "top_models": [ { "model": "claude-sonnet-5", "cost_usd": 88.10, "byok_list_price_usd": 0 } ],
    "top_keys": [ { "api_key_id": 42, "api_key_name": "prod-backend", "cost_usd": 61.30, "byok_list_price_usd": 0 } ],
    "anomalies": [
      { "type": "data_not_final", "severity": "info", "message": "The last 2 hours of this range are not yet final." }
    ]
  },
  "meta": {
    "generated_at": "2026-09-24T06:00:02Z",
    "data_complete_until": "2026-09-24T04:00:00Z"
  }
}

GET /freshness - how complete the data is

GET /freshness

Returns data_complete_until, latest_recorded_at, safe_until and pipeline_lag_seconds - poll it before treating a just-fetched range as settled. safe_until is the newest recorded_at a caller can trust as complete; rows younger than that may still be arriving.

Example Request

curl "https://synthorai.io/api/v1/billing/freshness" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{
  "data": {
    "data_complete_until": "2026-09-24T04:00:00Z",
    "latest_recorded_at": "2026-09-24T05:58:31Z",
    "safe_until": "2026-09-24T05:58:31Z",
    "pipeline_lag_seconds": 47
  },
  "meta": { "generated_at": "2026-09-24T06:00:02Z" }
}

GET /balance - what the workspace can spend now

GET /balance

Returns the workspace's current available balance - the same figure the console shows. Built for polling: every 30 to 60 seconds is plenty. It has its own rate limit, so polling it never uses up the budget of the data endpoints, and it keeps answering when the balance is zero.

Example Request

curl "https://synthorai.io/api/v1/billing/balance" \
  -H "Authorization: Bearer sk-syn-bill-..."

Example Response

{
  "data": {
    "available_usd": 1284.37,
    "debt_usd": 0,
    "source": "live",
    "voucher": null,
    "scheduled_credits": [
      { "amount_usd": 250, "release_at": "2026-10-01T00:00:00Z" }
    ],
    "scheduled_credit_total_usd": 250
  },
  "meta": { "generated_at": "2026-09-24T06:00:02Z", "currency": "USD" }
}
ParameterDescription
available_usdThe general balance that requests are charged against right now. Never below 0. It can briefly rise as in-flight requests settle; for spend, use /usage or /records.
debt_usdThe outstanding balance: usage beyond what was credited. 0 when nothing is owed. A top-up pays it first; only the rest becomes available_usd.
sourcelive: read from the real-time ledger. delayed: the ledger could not be reached, and the value comes from a database copy that can trail it by about 30 seconds.
voucherPromotional credit for image generation only (applies_to), on the listed model patterns and until expires_at; other requests draw on available_usd alone, and it is not part of available_usd. null when there is none; active is false once it has expired.
scheduled_creditsCredits already agreed that arrive in tranches (amount_usd, release_at); each is added to the balance within about 10 minutes of release_at, not before. scheduled_credit_total_usd is their sum.

Errors

Every error is {"error": {"type", "message", "hint"}}:

{
  "error": {
    "type": "range_too_large",
    "message": "the requested range exceeds this endpoint's cap",
    "hint": "narrow start/end to at most 31 days, or sync /records with a cursor instead"
  }
}
TypeHTTP statusMeaning
invalid_parameter400A query parameter is missing, malformed, or out of range.
range_too_large400The requested time range or id window is larger than the endpoint allows; the hint names the cap.
start_too_recent400A start/end window starts after safe_until, where its start boundary is not settled yet. Retry after the Retry-After header, or sync with a cursor.
wrong_key_kind401The credential is not a billing key, or a billing key was used outside these endpoints.
authentication_error401The key is missing, invalid, expired, or blocked by its IP allow-list.
key_owner_not_authorized403The key's creator is no longer the workspace owner. Billing keys are owner-only and stop working the moment ownership changes hands; the new owner must create their own.
rate_limited429More than 60 requests per minute on this workspace. Retry after the time in the Retry-After header.
too_many_concurrent_requests429More than 3 concurrent requests on this workspace, or a second export already streaming.
replica_unavailable503This process has no read replica configured; the API never falls back to the primary.
query_timeout503The query ran past the 10-second statement timeout or the 15-second deadline. Narrow the range and try again.
internal_error500An unexpected server error. Retry, and report it if it persists.

Rate limits and caps

GuardValue
Rate limit60 requests per minute per workspace, a sliding window shared across every instance.
Concurrency3 in-flight requests per workspace on each server, and one export stream per workspace on each server.
Balance polling/balance: 60 requests per minute and 2 in flight per workspace, counted separately from the other endpoints.
Range caps/usage: 31 days at hour granularity, 366 days at day. /summary: up to 366 days. /records and the export: up to 31 days.
Page caps/usage returns up to 5,000 rows per page, /records up to 1,000. The export streams pages of 2,000 and stops at 1,000,000 rows - resumable with cursor. It also stops after walking 2,000,000 record ids (stopped_because=scan_cap) and paces itself to the database; resume with cursor.
CachingEvery response carries Cache-Control: no-store.

Never returned: channel_id, upstream_request_id, usage_raw, other, content, ip_address, user_agent, cost_detail.

Avoiding time gaps

Requests are written to the log asynchronously after they finish (seconds normally, longer under backlog), and the hourly rollup that /usage and /summary read from ingests in batches. The most recent hours are always a little incomplete. /usage and /summary also add the requests written since the last ingest (meta.live_tail is true); when the rollup is too far behind for that, meta.live_tail is false and recent buckets are partial.

data_complete_until is the earlier of two points: the hourly rollup's watermark minus its current pipeline lag, or the oldest request that has not yet been rolled up - minus a 30-minute safety margin, floored to the hour.

ℹ

A final bucket is stable with one exception: the daily reconciliation job may still correct an hour within 48 hours if it finds drift. /records is always the source of truth for a request's exact numbers.

Every /usage row carries a final flag: true once its bucket ends at or before data_complete_until from /freshness. A final bucket never changes again.

  • Aggregates: upsert each row into your own store by (bucket_start, api_key_id, model), and re-fetch any bucket while it is still final: false. Never derive a delta by subtracting an old total from a new one.
  • Raw rows: sync by cursor, never by time. Keep next_cursor and resume from it on the next call, and dedupe incoming rows by record_id (a stable opaque id per row) in case a resumed page redelivers one. Rows younger than about 5 minutes are held back - safe_until is the newest recorded row's time minus 5 minutes - so a row still being committed can never slip past a cursor position already handed out, and nothing falls into a gap. To count requests, count distinct request_id.
  • Bounded window: if you query a start/end range instead of syncing by cursor, trust it as complete only once window_final is true (it is true only when you gave an end and that end is at or before safe_until). If window_final is false, call again later with the same range, or switch to syncing by cursor. A window must start at or before safe_until; a later start answers 400 start_too_recent with a Retry-After header.

Recommended integration

  • Nightly invoicing: call /usage once a day with granularity=day, and re-fetch any bucket your last run still saw as final: false.
  • Per-request detail: call /records hourly with cursor set to the last next_cursor you stored, keep paging while has_more is true, and dedupe incoming rows by record_id.
  • Dashboards: call /summary for the visible range instead of aggregating /records yourself - it already carries the anomalies list.

Need to mint or manage regular API keys programmatically instead of reading billing data? See Provisioning Keys.