MCP Server
Connect any MCP-capable agent (Claude Code, Cursor, VS Code, Codex, the OpenAI Agents SDK and more) to Synthorai with one URL and your API key. Every model and modality, API key management and billing data become tools the agent can call. Each call is authenticated, limited and billed exactly like the matching REST request.
| Endpoint | https://synthorai.io/v1/mcp |
| Transport | Streamable HTTP, stateless (no sessions), JSON responses |
| Authentication | Authorization: Bearer <API key> |
The key you connect with decides which tools you get: an inference key (sk-syn-…) calls models, a provisioning key (sk-syn-prov-…) manages API keys, and a billing key (sk-syn-bill-…) reads balance and usage. Connect the server once per key kind if an agent needs more than one.
Quick start
- Create an API key in the console on the API Keys page. For an agent, give it a spend limit and a model allow-list.
- Add the server to your client with one of the snippets below, reading the key from an environment variable rather than pasting it into a file.
- Ask your agent, for example: “List the three cheapest models that support tools, then ask the fastest one to summarize this file, and tell me what it cost.”
Connect your client
All clients use the same endpoint and header. Set SYNTHORAI_API_KEY in your environment first.
Claude Code
Run /mcp inside Claude Code to check the connection. Add --scope user to make the server available in every project, or commit the .mcp.json form to share it with your team (the key stays in each developer's environment).
claude mcp add --transport http synthorai https://synthorai.io/v1/mcp \
--header "Authorization: Bearer $SYNTHORAI_API_KEY"{
"mcpServers": {
"synthorai": {
"type": "http",
"url": "https://synthorai.io/v1/mcp",
"headers": { "Authorization": "Bearer ${SYNTHORAI_API_KEY}" }
}
}
} Cursor
Put this in ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project).
{
"mcpServers": {
"synthorai": {
"url": "https://synthorai.io/v1/mcp",
"headers": { "Authorization": "Bearer ${env:SYNTHORAI_API_KEY}" }
}
}
} VS Code (GitHub Copilot)
Put this in .vscode/mcp.json. VS Code asks for the key once and stores it securely.
{
"inputs": [
{ "type": "promptString", "id": "synthorai-key", "description": "Synthorai API key", "password": true }
],
"servers": {
"synthorai": {
"type": "http",
"url": "https://synthorai.io/v1/mcp",
"headers": { "Authorization": "Bearer ${input:synthorai-key}" }
}
}
} Codex CLI
Put this in ~/.codex/config.toml, or run codex mcp add synthorai --url https://synthorai.io/v1/mcp --bearer-token-env-var SYNTHORAI_API_KEY.
[mcp_servers.synthorai]
url = "https://synthorai.io/v1/mcp"
bearer_token_env_var = "SYNTHORAI_API_KEY" OpenAI Responses API
OpenAI's servers call the endpoint on your behalf, so leave the key's IP allow-list empty (or allow OpenAI's egress ranges). Use allowed_tools to expose only the tools the model needs.
import os
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-4.1",
tools=[{
"type": "mcp",
"server_label": "synthorai",
"server_url": "https://synthorai.io/v1/mcp",
"headers": {"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"},
"allowed_tools": ["list_models", "chat_completion"],
"require_approval": "never",
}],
input="Find the cheapest Synthorai chat model with tool support and ask it for a haiku about gateways.",
)
print(resp.output_text) OpenAI Agents SDK
Works the same with any framework built on the MCP client SDKs (LangChain, LlamaIndex, Pydantic AI, Vercel AI SDK, Mastra …): point its Streamable HTTP transport at the endpoint with the header.
import asyncio, os
from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp
async def main():
async with MCPServerStreamableHttp(
name="synthorai",
params={
"url": "https://synthorai.io/v1/mcp",
"headers": {"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"},
"timeout": 120,
},
cache_tools_list=True,
) as synthorai:
agent = Agent(name="assistant", instructions="Use the Synthorai tools.", mcp_servers=[synthorai])
result = await Runner.run(agent, "List three chat models under $1 per million input tokens.")
print(result.final_output)
asyncio.run(main()) MCP Python SDK
import asyncio, os, httpx2
from mcp import Client
from mcp.client.streamable_http import streamable_http_client
async def main():
http = httpx2.AsyncClient(
headers={"Authorization": f"Bearer {os.environ['SYNTHORAI_API_KEY']}"}, timeout=120)
async with Client(streamable_http_client("https://synthorai.io/v1/mcp", http_client=http)) as mcp:
tools = await mcp.list_tools()
print([t.name for t in tools.tools])
result = await mcp.call_tool("chat_completion",
{"model": "deepseek-v4-flash", "prompt": "Say hi in five words"})
print(result.content[0].text)
asyncio.run(main()) Claude Desktop
Custom connectors in Claude Desktop and claude.ai sign in with OAuth, which this server does not offer yet (see the roadmap below). Until then, Claude Desktop can connect through the mcp-remote bridge, which adds the header for you:
{
"mcpServers": {
"synthorai": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://synthorai.io/v1/mcp", "--header", "Authorization:${AUTH_HEADER}"],
"env": { "AUTH_HEADER": "Bearer sk-syn-..." }
}
}
} curl
The server is plain JSON-RPC over HTTP, so no SDK is needed. No handshake is required: every request stands on its own.
curl https://synthorai.io/v1/mcp \
-H "Authorization: Bearer $SYNTHORAI_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
"params":{"name":"chat_completion",
"arguments":{"model":"deepseek-v4-flash","prompt":"Say hi in five words"}}}' Example response
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"content": [
{ "type": "text", "text": "Hi there, great to meet!" },
{ "type": "text", "text": "[deepseek-v4-flash · finish_reason stop · 12 in / 7 out tokens · cost $0.0000031 · request_id 0199a7c2-…]" }
],
"structuredContent": {
"model": "deepseek-v4-flash",
"content": "Hi there, great to meet!",
"finish_reason": "stop",
"usage": { "prompt_tokens": 12, "completion_tokens": 7, "total_tokens": 19 },
"cost_usd": 0.0000031,
"request_id": "0199a7c2-…"
}
}
} Tools
tools/list returns only the tools your key can use: it is filtered by key kind and by the features enabled for your workspace (image and video generation are in limited preview). Tools that run a model cost exactly what the REST call costs; every other tool is free.
Inference key
| Tool | What it does | REST endpoint | Notes |
|---|---|---|---|
list_models | Models the key can call, cheapest first, with context length, modalities, capabilities and list prices. Filters: category, capability, input modality, search. | GET /v1/models + /api/models | read-only |
get_model | Full details of one model, including the effective price after any discount and whether the key can call it. | GET /v1/models/{id} | read-only |
get_pricing | The price list (USD after discounts): per-token, per-call, per-minute, per-second and tiered prices. | GET /api/pricing | read-only |
chat_completion | One chat completion on any model: a prompt or a full messages array (images, tool results), optional structured output, function tools, reasoning effort and fallback models. Returns text, tool calls, usage, cost and request_id. | POST /v1/chat/completions | billed |
generate_image | Generate or edit images; returned as image content, or as links with response_format=url. | POST /v1/images/generations | billed |
list_image_models | Image models with prices and supported inputs. | GET /v1/images/models | read-only |
create_video | Start a video job from a prompt or a first frame; can wait up to 60 s for completion. Billed when the job completes. | POST /v1/videos | billed |
get_video | Job status; when completed, links to the videos (valid about 24 hours). | GET /v1/videos/{id} | read-only |
cancel_video | Cancel a job that is still queued (a job that has started cannot be stopped). Cancelled and failed jobs are not billed. | DELETE /v1/videos/{id} | destructive |
list_video_models | Video models with resolutions, durations and prices. | GET /v1/videos/models | read-only |
text_to_speech | Speech from text, returned as audio content. | POST /v1/audio/speech | billed |
transcribe_audio | Text from audio given as a URL or base64 data (up to 25 MB). | POST /v1/audio/transcriptions | billed |
create_embeddings | Embedding vectors for one text or a batch of up to 2048. | POST /v1/embeddings | billed |
get_key_info | The key's spend limit, used and remaining amount, reset period, and spend today / this week / this month (USD). | GET /v1/key | read-only |
get_generation | The billing record of one request by request_id: cost, token breakdown, latency, status. | GET /v1/generation | read-only |
Provisioning key
| Tool | What it does | REST endpoint | Notes |
|---|---|---|---|
create_api_key | Create an inference key with spend limit, reset period, model and IP allow-lists, expiry and metadata. The secret is in the result only. | POST /api/provisioning/keys | changes data |
list_api_keys | The workspace's inference keys with limits, spend and status; filter by metadata. | GET /api/provisioning/keys | read-only |
get_api_key | One key's settings, spend and status. | GET /api/provisioning/keys/{id} | read-only |
update_api_key | Change a key's settings, or disable / re-enable it. Only the fields given change. | PATCH /api/provisioning/keys/{id} | destructive |
delete_api_key | Delete a key permanently. | DELETE /api/provisioning/keys/{id} | destructive |
Billing key
| Tool | What it does | REST endpoint | Notes |
|---|---|---|---|
get_balance | Spendable balance now (USD), with voucher and scheduled credits. | GET /api/v1/billing/balance | read-only |
get_usage | Spend and requests per hour or day, by key and/or model. | GET /api/v1/billing/usage | read-only |
list_usage_records | Per-request records with tokens, cost and status, paged by cursor. | GET /api/v1/billing/records | read-only |
get_billing_summary | Totals over a range with the top keys and models. | GET /api/v1/billing/summary | read-only |
get_data_freshness | How current the billing data is. | GET /api/v1/billing/freshness | read-only |
list_models, get_model and get_pricing are offered to every key kind; for an inference key, list_models shows only the models that key can call.
What results look like
Every result carries a readable text block and the same data as JSON in structuredContent, so both chat clients and programs can use it. Images and audio come back as image and audio content, videos as links. Tools that run a model end with a receipt line (model, tokens, cost and request_id) which get_generation can look up later.
Permissions and safety
- Same rules as the REST API. Each tool call is carried out by the REST endpoint it names, with your key: authentication, key kind, IP allow-list, model allow-list, spend limit, rate limits and billing all apply unchanged. MCP adds no permission and skips no check.
- Least privilege by key kind. Inference keys cannot manage keys or read workspace billing; provisioning keys cannot call models; billing keys are read-only.
- Approval hints. Every tool carries MCP annotations. Read-only tools are marked
readOnlyHint; tools that spend money are not read-only;update_api_key,delete_api_keyandcancel_videoare markeddestructiveHint. Clients use these to decide what to ask you before running. - Recommended agent key: a dedicated inference key with a spend limit that resets daily, an
allowed_modelslist, and an IP allow-list when the agent runs on known hosts. Revoke it on its own without touching production keys. - Secrets and untrusted content.
create_api_keyreturns the new key once: tell your agent where to store it. Treat model output and tool results as untrusted input (prompt injection) before letting an agent act on them.
Protocol details
| Item | Detail |
|---|---|
| Protocol versions | 2026-07-28 (stateless, per-request _meta and Mcp-Method / Mcp-Name headers, server/discover) and 2025-11-25, 2025-06-18, 2025-03-26, 2024-11-05 (initialize handshake). Both on the same URL; each request is served in the era it speaks. |
| Sessions | None. No Mcp-Session-Id is issued, so any request can reach any server replica and nothing needs to be re-initialized after a reconnect. |
| Methods | initialize, ping, tools/list, tools/call; with 2026-07-28, server/discover instead of the handshake. GET and DELETE return 405; JSON-RPC batches are not accepted. |
| Errors | An unknown tool is a JSON-RPC error (-32602). Everything else (invalid arguments, an API refusal, a failed generation) is a tool result with isError: true and a message the model can act on. With 2026-07-28, envelope problems return HTTP 400 with -32020 (header mismatch) or -32022 (unsupported version, listing the supported ones). |
| Caching | tools/list is returned in a fixed order (prompt-cache friendly) with ttlMs of 10 minutes and cacheScope: "private", since the list depends on your key. |
| Rate limit | 600 MCP messages per minute per key (HTTP 429 with Retry-After). Each tool call also counts against the limits of the REST endpoint it calls. |
| Long calls | A call can run up to 30 minutes (long reasoning). Video generation is asynchronous: create_video returns a job and get_video polls it. |
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| 401, or the client reports that authentication is needed | The key is missing, invalid, expired or disabled. Send it as Authorization: Bearer <key> and check that the environment variable is set where the client runs. |
| A tool you expect is not listed | Tools depend on the key kind (see the tables above) and on preview features: image and video tools appear only for workspaces admitted to those previews. |
| A tool result says HTTP 402 | The workspace balance is used up. Top up in the console; get_key_info shows the key's own limit. |
| A tool result says HTTP 403 for a model | The key's allowed_models or your workspace's model access excludes it. list_models shows exactly what the key can call. |
chat_completion returns no text with finish_reason length | A reasoning model spent the whole max_tokens budget thinking. Raise max_tokens or lower reasoning_effort. |
| HTTP 429 | More than 600 messages a minute on one key, or the REST endpoint's own limit. Wait for Retry-After. |
Coming next
- OAuth sign-in, so claude.ai, Claude Desktop and ChatGPT connectors can connect without a pasted key, with short-lived, spend-capped credentials.
- MCP gateway: the same endpoint also aggregating third-party MCP servers (GitHub, Slack, your own) under your key, with per-tool permissions, credentials kept server-side and one audit log.
- Progress for long calls, streamed while a tool runs.
Related: API Keys · Provisioning Keys · Billing API · Chat Completions