신규 무료 가입, 10회 호출 제공. 최대 $1, 카드 불필요.

Parameters by Upstream

Some request fields do not mean the same thing everywhere. This page states what the gateway does with each one on the way to each provider - passed through, translated into the provider's own equivalent, or dropped.

This is not a list of every accepted field; each endpoint page documents its own request body. Listed here are the fields where supported is not a yes or no, because the answer changes with the provider your model runs on.

reasoning_effort

Full guide, with values and cost impact →

호출 대상 업스트림 동작 설명
/v1/chat/completions OpenAI-compatible (GPT, GLM, DeepSeek, Qwen, Kimi…) 그대로 전달 보낸 그대로 공급사로 전달되며, 예외가 하나 있습니다. 모델의 카탈로그 항목이 허용하는 추론 강도 단계를 선언한 경우, 가장 낮은 단계보다 낮은 값은 그 단계로 올리고 가장 높은 단계보다 높은 값은 낮춥니다(GPT-6.1 Sol은 추론을 끌 수 없으므로 none과 minimal은 low로 실행됩니다). 그 밖에는 게이트웨이가 값을 검증하지 않으며 공급사가 이를 따른다고 보장하지도 않습니다. 흔히 minimal, low, medium, high를 쓰지만 허용 범위는 공급사가 정합니다(OpenAI는 대부분의 모델에서 none도, GLM은 max까지 받습니다).
/v1/chat/completions Google Vertex - Gemini 3.x 변환 Gemini의 thinkingLevel로 변환됩니다. none과 minimal은 모두 minimal로, low/medium/high는 그대로 대응됩니다. 요청에 명시적인 google.thinking_config가 있으면 그쪽이 우선합니다.
/v1/chat/completions Google Vertex - Gemini 2.5 폐기 Gemini 2.5는 thinkingBudget만 이해하고 thinkingLevel은 거부하므로, 게이트웨이는 둘 다 보내지 않고 모델의 기본 사고 동작을 유지합니다. 2.5를 제어하려면 google.thinking_config의 thinking_budget을 사용하세요.
/v1/chat/completions Anthropic (Claude) 변환 Anthropic의 thinking으로 변환됩니다. none과 minimal은 사고를 끄고({"type":"disabled"}), low/medium/high/xhigh/max는 각각 2,048 / 8,192 / 16,384 / 32,768 토큰 예산이 됩니다 - 반대 방향에서 쓰는 버킷 매핑의 역매핑입니다. 단계에서 파생된 예산은 max_tokens 미만으로 제한되며, max_tokens가 Anthropic의 최소 1,024 토큰조차 담지 못하면 아예 보내지 않습니다. 따라서 단계를 지정해도 잘 동작하던 요청이 업스트림 400이 되는 일은 없습니다.
/v1/messages OpenAI-compatible (GPT, GLM, DeepSeek, Qwen, Kimi…) 변환 확장 사고가 켜져 있고 예산이 0보다 클 때 thinking.budget_tokens에서 구간별로 도출됩니다: 2,048 이하는 low, 8,192 이하는 medium, 그 이상은 high. 예산이 0이거나 없으면 아무것도 보내지 않습니다. 카탈로그 항목에 허용 추론 강도 단계를 선언한 모델(GPT-6.1 Sol)은 대신 요청의 의도를 따릅니다. thinking이 disabled 또는 between_tools이면 output_config.effort가 함께 지정되어 있어도 가장 낮은 단계(low)로 실행하고, 그 밖에는 output_config.effort를 그대로 추론 강도로 보냅니다.
/v1/responses OpenAI-compatible (GPT, GLM, DeepSeek, Qwen, Kimi…) 변환 Responses API의 reasoning.effort 객체에서 읽어 평면 reasoning_effort 필드로 전송합니다.

thinking / thinking_budget / enable_thinking

호출 대상 업스트림 동작 설명
/v1/chat/completions Anthropic (Claude) 변환 같은 다이얼에 대한 다른 벤더의 이름들이며, 게이트웨이가 정규화합니다. thinking:{"type":"disabled"}, enable_thinking:false, thinking_budget:0은 모두 사고를 끕니다. thinking_budget:N과 thinking:{"type":"enabled","budget_tokens":N}의 예산은 그대로 전달되며, 범위를 벗어난 값은 조용히 조정되지 않고 업스트림이 거부합니다. thinking:{"type":"adaptive"}는 업스트림에 맡깁니다.
/v1/chat/completions Google Vertex - Gemini 3.x 변환 thinkingConfig로 정규화됩니다. 사고 끄기는 thinkingLevel:"minimal"이 되고, 명시적 예산은 thinkingBudget이 됩니다. 요청에 명시적인 google.thinking_config가 있으면 그쪽이 우선합니다.
/v1/chat/completions Google Vertex - Gemini 2.5 변환 thinkingConfig.thinkingBudget으로 정규화됩니다 - 사고 끄기는 예산 0을 보내고, 명시적 예산은 그대로 전달됩니다. 2.5는 thinkingLevel을 거부하므로 여기서는 쓰지 않습니다. 요청에 명시적인 google.thinking_config가 있으면 그쪽이 우선합니다.

Every row is transcribed from the converter that implements it and pinned by a Go test, so this table cannot drift from the running gateway without a build failing.