New Sign up free, 10 calls on us. Up to $1, no card needed.

MCP Server

Connect any MCP-capable agent (Claude Code, Cursor, VS Code, Codex, the OpenAI Agents SDK and more) to Synthorai with one URL and your API key. Every model and modality, API key management and billing data become tools the agent can call. Each call is authenticated, limited and billed exactly like the matching REST request.

Endpointhttps://synthorai.io/v1/mcp
TransportStreamable HTTP, stateless (no sessions), JSON responses
AuthenticationAuthorization: Bearer <API key>
ℹ

The key you connect with decides which tools you get: an inference key (sk-syn-…) calls models, a provisioning key (sk-syn-prov-…) manages API keys, and a billing key (sk-syn-bill-…) reads balance and usage. Connect the server once per key kind if an agent needs more than one.

Quick start

  1. Create an API key in the console on the API Keys page. For an agent, give it a spend limit and a model allow-list.
  2. Add the server to your client with one of the snippets below, reading the key from an environment variable rather than pasting it into a file.
  3. Ask your agent, for example: “List the three cheapest models that support tools, then ask the fastest one to summarize this file, and tell me what it cost.”

Connect your client

All clients use the same endpoint and header. Set SYNTHORAI_API_KEY in your environment first.

Claude Code

Run /mcp inside Claude Code to check the connection. Add --scope user to make the server available in every project, or commit the .mcp.json form to share it with your team (the key stays in each developer's environment).

claude mcp add --transport http synthorai https://synthorai.io/v1/mcp \
  --header "Authorization: Bearer $SYNTHORAI_API_KEY"
{
  "mcpServers": {
    "synthorai": {
      "type": "http",
      "url": "https://synthorai.io/v1/mcp",
      "headers": { "Authorization": "Bearer ${SYNTHORAI_API_KEY}" }
    }
  }
}

Cursor

Put this in ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project).

{
  "mcpServers": {
    "synthorai": {
      "url": "https://synthorai.io/v1/mcp",
      "headers": { "Authorization": "Bearer ${env:SYNTHORAI_API_KEY}" }
    }
  }
}

VS Code (GitHub Copilot)

Put this in .vscode/mcp.json. VS Code asks for the key once and stores it securely.

{
  "inputs": [
    { "type": "promptString", "id": "synthorai-key", "description": "Synthorai API key", "password": true }
  ],
  "servers": {
    "synthorai": {
      "type": "http",
      "url": "https://synthorai.io/v1/mcp",
      "headers": { "Authorization": "Bearer ${input:synthorai-key}" }
    }
  }
}

Codex CLI

Put this in ~/.codex/config.toml, or run codex mcp add synthorai --url https://synthorai.io/v1/mcp --bearer-token-env-var SYNTHORAI_API_KEY.

[mcp_servers.synthorai]
url = "https://synthorai.io/v1/mcp"
bearer_token_env_var = "SYNTHORAI_API_KEY"

OpenAI Responses API

OpenAI's servers call the endpoint on your behalf, so leave the key's IP allow-list empty (or allow OpenAI's egress ranges). Use allowed_tools to expose only the tools the model needs.

import os
from openai import OpenAI

client = OpenAI()
resp = client.responses.create(
    model="gpt-4.1",
    tools=[{
        "type": "mcp",
        "server_label": "synthorai",
        "server_url": "https://synthorai.io/v1/mcp",
        "headers": {"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"},
        "allowed_tools": ["list_models", "chat_completion"],
        "require_approval": "never",
    }],
    input="Find the cheapest Synthorai chat model with tool support and ask it for a haiku about gateways.",
)
print(resp.output_text)

OpenAI Agents SDK

Works the same with any framework built on the MCP client SDKs (LangChain, LlamaIndex, Pydantic AI, Vercel AI SDK, Mastra …): point its Streamable HTTP transport at the endpoint with the header.

import asyncio, os
from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp

async def main():
    async with MCPServerStreamableHttp(
        name="synthorai",
        params={
            "url": "https://synthorai.io/v1/mcp",
            "headers": {"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"},
            "timeout": 120,
        },
        cache_tools_list=True,
    ) as synthorai:
        agent = Agent(name="assistant", instructions="Use the Synthorai tools.", mcp_servers=[synthorai])
        result = await Runner.run(agent, "List three chat models under $1 per million input tokens.")
        print(result.final_output)

asyncio.run(main())

MCP Python SDK

import asyncio, os, httpx2
from mcp import Client
from mcp.client.streamable_http import streamable_http_client

async def main():
    http = httpx2.AsyncClient(
        headers={"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"}, timeout=120)
    async with Client(streamable_http_client("https://synthorai.io/v1/mcp", http_client=http)) as mcp:
        tools = await mcp.list_tools()
        print([t.name for t in tools.tools])
        result = await mcp.call_tool("chat_completion",
                                     {"model": "deepseek-v4-flash", "prompt": "Say hi in five words"})
        print(result.content[0].text)

asyncio.run(main())

Claude Desktop

Custom connectors in Claude Desktop and claude.ai sign in with OAuth, which this server does not offer yet (see the roadmap below). Until then, Claude Desktop can connect through the mcp-remote bridge, which adds the header for you:

{
  "mcpServers": {
    "synthorai": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://synthorai.io/v1/mcp", "--header", "Authorization:${AUTH_HEADER}"],
      "env": { "AUTH_HEADER": "Bearer sk-syn-..." }
    }
  }
}

curl

The server is plain JSON-RPC over HTTP, so no SDK is needed. No handshake is required: every request stands on its own.

curl https://synthorai.io/v1/mcp \
  -H "Authorization: Bearer $SYNTHORAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
       "params":{"name":"chat_completion",
                 "arguments":{"model":"deepseek-v4-flash","prompt":"Say hi in five words"}}}'

Example response

{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "content": [
      { "type": "text", "text": "Hi there, great to meet!" },
      { "type": "text", "text": "[deepseek-v4-flash · finish_reason stop · 12 in / 7 out tokens · cost $0.0000031 · request_id 0199a7c2-…]" }
    ],
    "structuredContent": {
      "model": "deepseek-v4-flash",
      "content": "Hi there, great to meet!",
      "finish_reason": "stop",
      "usage": { "prompt_tokens": 12, "completion_tokens": 7, "total_tokens": 19 },
      "cost_usd": 0.0000031,
      "request_id": "0199a7c2-…"
    }
  }
}

Tools

tools/list returns only the tools your key can use: it is filtered by key kind and by the features enabled for your workspace (image and video generation are in limited preview). Tools that run a model cost exactly what the REST call costs; every other tool is free.

Inference key

ToolWhat it doesREST endpointNotes
list_models Models the key can call, cheapest first, with context length, modalities, capabilities and list prices. Filters: category, capability, input modality, search. GET /v1/models + /api/models read-only
get_model Full details of one model, including the effective price after any discount and whether the key can call it. GET /v1/models/{id} read-only
get_pricing The price list (USD after discounts): per-token, per-call, per-minute, per-second and tiered prices. GET /api/pricing read-only
chat_completion One chat completion on any model: a prompt or a full messages array (images, tool results), optional structured output, function tools, reasoning effort and fallback models. Returns text, tool calls, usage, cost and request_id. POST /v1/chat/completions billed
generate_image Generate or edit images; returned as image content, or as links with response_format=url. POST /v1/images/generations billed
list_image_models Image models with prices and supported inputs. GET /v1/images/models read-only
create_video Start a video job from a prompt or a first frame; can wait up to 60 s for completion. Billed when the job completes. POST /v1/videos billed
get_video Job status; when completed, links to the videos (valid about 24 hours). GET /v1/videos/{id} read-only
cancel_video Cancel a job that is still queued (a job that has started cannot be stopped). Cancelled and failed jobs are not billed. DELETE /v1/videos/{id} destructive
list_video_models Video models with resolutions, durations and prices. GET /v1/videos/models read-only
text_to_speech Speech from text, returned as audio content. POST /v1/audio/speech billed
transcribe_audio Text from audio given as a URL or base64 data (up to 25 MB). POST /v1/audio/transcriptions billed
create_embeddings Embedding vectors for one text or a batch of up to 2048. POST /v1/embeddings billed
get_key_info The key's spend limit, used and remaining amount, reset period, and spend today / this week / this month (USD). GET /v1/key read-only
get_generation The billing record of one request by request_id: cost, token breakdown, latency, status. GET /v1/generation read-only

Provisioning key

ToolWhat it doesREST endpointNotes
create_api_key Create an inference key with spend limit, reset period, model and IP allow-lists, expiry and metadata. The secret is in the result only. POST /api/provisioning/keys changes data
list_api_keys The workspace's inference keys with limits, spend and status; filter by metadata. GET /api/provisioning/keys read-only
get_api_key One key's settings, spend and status. GET /api/provisioning/keys/{id} read-only
update_api_key Change a key's settings, or disable / re-enable it. Only the fields given change. PATCH /api/provisioning/keys/{id} destructive
delete_api_key Delete a key permanently. DELETE /api/provisioning/keys/{id} destructive

Billing key

ToolWhat it doesREST endpointNotes
get_balance Spendable balance now (USD), with voucher and scheduled credits. GET /api/v1/billing/balance read-only
get_usage Spend and requests per hour or day, by key and/or model. GET /api/v1/billing/usage read-only
list_usage_records Per-request records with tokens, cost and status, paged by cursor. GET /api/v1/billing/records read-only
get_billing_summary Totals over a range with the top keys and models. GET /api/v1/billing/summary read-only
get_data_freshness How current the billing data is. GET /api/v1/billing/freshness read-only

list_models, get_model and get_pricing are offered to every key kind; for an inference key, list_models shows only the models that key can call.

What results look like

Every result carries a readable text block and the same data as JSON in structuredContent, so both chat clients and programs can use it. Images and audio come back as image and audio content, videos as links. Tools that run a model end with a receipt line (model, tokens, cost and request_id) which get_generation can look up later.

Permissions and safety

  • Same rules as the REST API. Each tool call is carried out by the REST endpoint it names, with your key: authentication, key kind, IP allow-list, model allow-list, spend limit, rate limits and billing all apply unchanged. MCP adds no permission and skips no check.
  • Least privilege by key kind. Inference keys cannot manage keys or read workspace billing; provisioning keys cannot call models; billing keys are read-only.
  • Approval hints. Every tool carries MCP annotations. Read-only tools are marked readOnlyHint; tools that spend money are not read-only; update_api_key, delete_api_key and cancel_video are marked destructiveHint. Clients use these to decide what to ask you before running.
  • Recommended agent key: a dedicated inference key with a spend limit that resets daily, an allowed_models list, and an IP allow-list when the agent runs on known hosts. Revoke it on its own without touching production keys.
  • Secrets and untrusted content. create_api_key returns the new key once: tell your agent where to store it. Treat model output and tool results as untrusted input (prompt injection) before letting an agent act on them.

Protocol details

ItemDetail
Protocol versions2026-07-28 (stateless, per-request _meta and Mcp-Method / Mcp-Name headers, server/discover) and 2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05 (initialize handshake). Both on the same URL; each request is served in the era it speaks.
SessionsNone. No Mcp-Session-Id is issued, so any request can reach any server replica and nothing needs to be re-initialized after a reconnect.
Methodsinitialize, ping, tools/list, tools/call; with 2026-07-28, server/discover instead of the handshake. GET and DELETE return 405; JSON-RPC batches are not accepted.
ErrorsAn unknown tool is a JSON-RPC error (-32602). Everything else (invalid arguments, an API refusal, a failed generation) is a tool result with isError: true and a message the model can act on. With 2026-07-28, envelope problems return HTTP 400 with -32020 (header mismatch) or -32022 (unsupported version, listing the supported ones).
Cachingtools/list is returned in a fixed order (prompt-cache friendly) with ttlMs of 10 minutes and cacheScope: "private", since the list depends on your key.
Rate limit600 MCP messages per minute per key (HTTP 429 with Retry-After). Each tool call also counts against the limits of the REST endpoint it calls.
Long callsA call can run up to 30 minutes (long reasoning). Video generation is asynchronous: create_video returns a job and get_video polls it.

Troubleshooting

SymptomCause and fix
401, or the client reports that authentication is neededThe key is missing, invalid, expired or disabled. Send it as Authorization: Bearer <key> and check that the environment variable is set where the client runs.
A tool you expect is not listedTools depend on the key kind (see the tables above) and on preview features: image and video tools appear only for workspaces admitted to those previews.
A tool result says HTTP 402The workspace balance is used up. Top up in the console; get_key_info shows the key's own limit.
A tool result says HTTP 403 for a modelThe key's allowed_models or your workspace's model access excludes it. list_models shows exactly what the key can call.
chat_completion returns no text with finish_reason lengthA reasoning model spent the whole max_tokens budget thinking. Raise max_tokens or lower reasoning_effort.
HTTP 429More than 600 messages a minute on one key, or the REST endpoint's own limit. Wait for Retry-After.

Coming next

  • OAuth sign-in, so claude.ai, Claude Desktop and ChatGPT connectors can connect without a pasted key, with short-lived, spend-capped credentials.
  • MCP gateway: the same endpoint also aggregating third-party MCP servers (GitHub, Slack, your own) under your key, with per-tool permissions, credentials kept server-side and one audit log.
  • Progress for long calls, streamed while a tool runs.

Related: API Keys · Provisioning Keys · Billing API · Chat Completions