🎁 Novo Cadastre-se grátis, 10 chamadas por nossa conta. Até US$ 1, sem cartão.

MiniMax M3 vs Qwen3.7 Plus

vs

Qual usar, quando — veredito curado, não uma tabela de benchmark

Esses dois são próximos no papel: tanto o minimax-m3 quanto o qwen3.7-plus oferecem contexto de 1000000 tokens, entrada de texto/imagem/vídeo com saída de texto, e os mesmos sinalizadores de chat, código, raciocínio, ferramentas e contexto longo, com pensamento desativável, e ambos foram lançados em 2026-06-01. O que os separa é a tabela de preços e o teto de saída: o qwen3.7-plus cobra cerca de 1.33x o minimax-m3 na entrada ($0.4 contra $0.3), na saída ($1.6 contra $1.2) e nas leituras de cache ($0.08 contra $0.06), enquanto o minimax-m3 permite 524288 tokens de saída máxima contra 65536, ou seja 8x a folga. Escolha o minimax-m3 para execuções mais baratas e gerações únicas muito longas; escolha o qwen3.7-plus se preferir a Alibaba como fornecedora.

Preços

MiniMax M3 Qwen3.7 Plus Δ
Entrada / 1M tokens $0.3 $0.4 0.75×
Saída / 1M tokens $1.2 $1.6 0.75×
Leitura de cache / 1M tokens $0.06 $0.08 0.75×
Escrita de cache sem cobrança separada 1.25x

Tarifas do catálogo em tempo real no momento do build; a página de cada modelo contém o cartão atual.

Onde eles se posicionam — preço de entrada por 1M tokens entre todos os 63 modelos de chat nesta unidade de cobrança (escala logarítmica)

MiniMax M3 · $0.3 Qwen3.7 Plus · $0.4
$0.05 · Qwen3 VL Flash $30 · GPT-5.4 Pro

Capacidades

MiniMax M3 Qwen3.7 Plus
Uso de ferramentas sim sim
Controle de raciocínio configurável configurável
Saída estruturada sim
Cache de prompt implícito (automático) implícito + explícito
Tempo de vida do cache não publicado explicit: 5m, reset on hit
Prefixo mínimo em cache 512 tokens 1024 tokens

Especificações

MiniMax M3 Qwen3.7 Plus
Modalidades de entrada texto imagem vídeo texto imagem vídeo
Modalidades de saída texto texto
Lançamento 2026-06-01 2026-06-01
Janela de contexto 1M 1M
Saída máxima 524K 66K
Parâmetro de raciocínio
  • thinking.type
  • reasoning_split
  • enable_thinking
  • thinking_budget
  • preserve_thinking
Valores aceitos
thinking.type
  • adaptive
  • disabled
reasoning_split
  • boolean
enable_thinking
  • true
  • false
thinking_budget
  • in tokens
preserve_thinking
  • true
  • false
Padrão adaptive: thinking on, with the model deciding when extra reasoning helps

on

the Qwen3.7 Plus series is hybrid thinking with thinking enabled by default, and preserve_thinking is off

As especificações são transcritas da documentação de cada fornecedor; uma linha que um fornecedor não publica é omitida em vez de ser inferida. Fontes completas: MiniMax M3 · Qwen3.7 Plus

Um prompt, ambos os modelos — medidos pelo gateway

PROMPT Explain to a CFO, in exactly three sentences, why our LLM bill doubled after we enabled reasoning. VERIFICAR exatamente 3 frases

MiniMax M3 passou · 3 sentences

Reasoning-enabled models deliberately produce extensive internal "chain-of-thought" tokens before answering, often consuming 5–10x more tokens per query than a standard model. Because LLM providers bill per token for both that reasoning trace and the final response, your cost-per-query rises proportionally—and that multiplier is the primary driver of the doubled bill. The right response isn't to disable reasoning, but to route only genuinely complex queries to it while keeping routine traffic on the cheaper standard tier, which typically recovers 60–80% of the cost without sacrificing quality where it matters.

saída 665 tok (+548 pensamento) latência 5.7 s

Qwen3.7 Plus passou · 3 sentences

Enabling reasoning requires the model to generate extensive internal "thinking" tokens to process complex logic before delivering a final answer. Because our vendor charges us per token consumed, this hidden computational overhead effectively doubles our output volume for the exact same number of user requests. Consequently, our costs have doubled because we are now paying for the model's increased processing time to achieve higher-quality, more accurate results.

saída 1428 tok (+1345 pensamento) latência 25.2 s

Cumprimento da instrução (exatamente três frases — contáveis), ajuste ao público (registro de CFO), e a lacuna de cobrança do pensamento oculto que os medidores de tokens abaixo expõem.

PROMPT This function is supposed to return the longest run of consecutive equal items, but callers report it is off by one on single-element inputs and misses runs that end at the last index. Fix it and explain the bug in one paragraph. VERIFICAR a correção passa nos testes

MiniMax M3 passou · 8/8 cases

The bug is in the loop bound: `range(1, len(xs))` iterates from index 1 up to but not including `len(xs) - 1`... wait, actually it goes up to `len(xs) - 1`, so the last index is included. The real problems are two: 1. **Starting `best` and `cur` at 0** and iterating from index 1 means a single-element list (`len(xs) == 1`) never enters the loop, so the function returns `0` instead of `1`. The fix is to either seed `best = cur = 1` when `xs` is non-empty, or include index 0 in the loop logic. 2. **Missing the final run**: the function only updates `best` *inside* the loop when it finds a continuation. If the longest run ends at the very last element, the loop body never gets to compare it after the increment — actually it does, but only if the run is longer than the previous `best`. […]

saída 464 tok latência 7.9 s

Qwen3.7 Plus passou · 8/8 cases

## Fixed Code ```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` ## Bug Explanation The function has two related bugs rooted in how it counts run lengths. First, `cur` is initialized to `0` and only incremented on matches, so it counts the number of *transitions* between equal elements rather than the total elements in the run—a run of N identical items only reaches `cur = N − 1`, undercounting by one. This also causes the single-element case to return `0`, since the loop never executes and `cur` never increments. […]

saída 2317 tok (+2005 pensamento) latência 40.5 s

Se a correção é de fato certa (executável), a densidade da explicação, e a eficiência em tokens numa tarefa delimitada.

PROMPT Extract a JSON object with fields {company, quarter, revenue_usd_m, growth_pct, guidance} from this text. Use null for anything not stated; add no other fields. VERIFICAR JSON válido, esquema exato

MiniMax M3 passou · 5/5 fields, guidance null

```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null } ```

saída 2212 tok (+2161 pensamento) latência 23.3 s

Qwen3.7 Plus passou · 5/5 fields, guidance null

{ "company": "Meridian Robotics", "quarter": "September", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null }

saída 1807 tok (+1750 pensamento) latência 31.6 s

Obediência ao esquema (sem campos inventados), pressão de alucinação (guidance é explicitamente retida), e diferenças no caminho de saída estruturada.

PROMPT Write a 120-word product announcement for an API gateway feature that caches prompts across providers. Forbidden words: "seamless", "unlock", "game-changer", "revolutionize", "empower". Exactly one sentence must be a question. VERIFICAR 120 palavras, 0 palavras proibidas

MiniMax M3 passou · 120 words, 0 banned, 1 question

Introducing PromptCache, our new API gateway feature that caches prompts across multiple LLM providers. Developers can store prompt-completion pairs centrally, reducing redundant inference calls and lowering costs without sacrificing quality. How will your workflow change when repeated requests resolve instantly from a shared cache layer? The gateway intercepts outgoing requests, checks for matching prompt fingerprints, and serves previously generated completions when available, falling back to the original provider on misses. Compatible with OpenAI, Anthropic, Cohere, and custom endpoints, PromptCache integrates with existing routing rules and supports TTL, versioning, and per-tenant namespaces. Teams now gain predictable latency, reduced token spend, and improved throughput during traffic spikes today. […]

saída 3256 tok (+2892 pensamento) latência 21.5 s

Qwen3.7 Plus passou · 120 words, 0 banned, 1 question

Introducing PromptCache, our new API gateway feature designed to store and reuse prompts across multiple artificial intelligence providers. Have you ever noticed how repeatedly sending identical queries drains your budget and increases latency? PromptCache solves this by intelligently storing responses at the gateway level. When your application requests the same prompt from a different provider, the system instantly returns the cached result. This drastically reduces API costs and accelerates response times for your users. You can configure custom expiration times and set specific fallback rules for each vendor. You must stop paying twice for the exact same computation. Please upgrade your entire infrastructure today and experience much faster and cheaper integrations without changing a single line of your application code.

saída 4453 tok (+4312 pensamento) latência 76.8 s

Obediência às restrições (orçamento de palavras, lista de palavras proibidas, a única pergunta), impressão digital de estilo, e controle de comprimento.

Alterne entre eles com uma linha

Ambos os IDs estão em todas as abas abaixo — o par de linhas destacado é a única edição. Mesmo endpoint, mesma chave, mesmo formato de requisição.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="minimax-m3",
    # model="qwen3.7-plus",  # descomente esta linha, comente a linha acima
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

Obter uma chave de API →

FAQ

Qual é mais barato, MiniMax M3 ou Qwen3.7 Plus?

MiniMax M3 é mais barato em entrada / 1m tokens ($0.3 vs $0.4, com 1.3× de diferença). Outras linhas podem apontar para o outro lado — a tabela acima traz o quadro completo, e o custo real depende do seu mix.

Posso fazer um teste A/B de MiniMax M3 contra Qwen3.7 Plus sem duas integrações?

Sim. Ambos são servidos pelo mesmo endpoint compatível com OpenAI com uma única chave de API — a troca é uma alteração de uma linha na string do modelo, de modo que você pode rotear uma fração do tráfego para cada um e comparar as faturas diretamente.

MiniMax M3 e Qwen3.7 Plus suportam prompt caching?

Sim — ambos cobram leituras em cache abaixo da sua taxa de entrada, então cargas de trabalho com warm-prefix custam menos do que as taxas listadas sugerem. As linhas exatas de leitura em cache estão na tabela de preços acima.

Comparações relacionadas

De nossos estudos medidos