New Sign up free, 10 calls on us. Up to $1, no card needed.

MiniMax M3 vs Qwen3.7 Plus

MiniMax M3 has been retired from our catalogue. Its figures below are the last published rates; calls to it are no longer served, while the other model in this comparison is.

vs

Which one, when

These two are close on paper: both minimax-m3 and qwen3.7-plus offer a 1000000-token context, text/image/video input with text output, and the same chat, code, reasoning, tools and long-context flags, with thinking that can be disabled, and both released 2026-06-01. The separation is rate card and output ceiling: qwen3.7-plus bills about 1.33x minimax-m3 on input ($0.4 vs $0.3), output ($1.6 vs $1.2) and cache reads ($0.08 vs $0.06), while minimax-m3 allows 524288 max output tokens against 65536, or 8x the room. Pick minimax-m3 for cheaper runs and very long single generations; pick qwen3.7-plus if you prefer Alibaba as the vendor.

Benchmarks

MiniMax M3: the vendor has not published benchmark scores.

Above averageNo peer higherQwen3.7 Plus11 / 244 / 24
MiniMax M3 Qwen3.7 Plus other models measured peer average ★ no peer scored higher
SWE-Bench Pro
N/A
55.8%
OSWorld 2.0 partial
N/A
21.5%
JobBench
N/A
27.6%
GPQA Diamond
N/A
90.3%
ERQA
N/A
69.8%
Agents' Last Exam Pass
N/A
13.2%
LVBench
N/A
76.2%

Vendor-published: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google MiniMax Moonshot OpenAI Tencent Z.ai

Pricing

MiniMax M3 Qwen3.7 Plus Δ
Input / 1M tokens $0.3 $0.4 0.75×
Output / 1M tokens $1.2 $1.6 0.75×
Cache read / 1M tokens $0.06 $0.08 0.75×
Cache write no separate charge 1.25x -

Rates from the live catalogue at build time; each model page carries the current rate card.

Where they sit · input price per 1M tokens across all 76 chat models on this billing unit (log scale)

MiniMax M3 · $0.3 Qwen3.7 Plus · $0.4
$0.05 · Qwen3 VL Flash $30 · GPT-5.4 Pro

Capabilities

MiniMax M3 Qwen3.7 Plus
Tool calling yes yes
Thinking control configurable configurable
Structured output - yes
Prompt caching implicit (automatic) implicit + explicit
Cache lifetime not published explicit: 5m, reset on hit
Minimum cached prefix 512 tokens 1024 tokens

Specs

MiniMax M3 Qwen3.7 Plus
Input modalities text image video text image video
Output modalities text text
Released 2026-06-01 2026-06-01
Context window 1M 1M
Max output 524K 66K
Thinking parameter
  • thinking.type
  • reasoning_split
  • enable_thinking
  • thinking_budget
  • preserve_thinking
Accepted values
thinking.type
  • adaptive
  • disabled
reasoning_split
  • boolean
enable_thinking
  • true
  • false
thinking_budget
  • in tokens
preserve_thinking
  • true
  • false
Default adaptive: thinking on, with the model deciding when extra reasoning helps

on

the Qwen3.7 Plus series is hybrid thinking with thinking enabled by default, and preserve_thinking is off

Specs are transcribed from each vendor’s documentation; a row a vendor does not publish is left out rather than inferred. Full sources: MiniMax M3 · Qwen3.7 Plus

Switch between them with one line

Both ids are in every tab below; the highlighted pair of lines is the only edit. Same endpoint, same key, same request shape.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="minimax-m3",
    # model="qwen3.7-plus",  # uncomment this line, comment the one above
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

Get your API key →

FAQ

Which is cheaper, MiniMax M3 or Qwen3.7 Plus?

MiniMax M3 is cheaper on the "Input / 1M tokens" row ($0.3 vs $0.4, 1.3× apart). Other rows may point the other way; the table above carries the full rate card, and real cost depends on your mix.

Can I A/B test MiniMax M3 against Qwen3.7 Plus without two integrations?

Yes. Both are served through the same OpenAI-compatible endpoint with one API key. Switching is a one-line change to the model id, so you can route a fraction of traffic to each and compare bills directly.

Do MiniMax M3 and Qwen3.7 Plus support prompt caching?

Yes. Both bill cache reads below their input rate, so warm-prefix workloads cost less than the list rates suggest. The exact cache-read rows are in the pricing table above.

Related comparisons

From our measured studies