新人 免费注册,送 10 次调用,最高 $1,无需绑卡。

Qwen3 TTS Instruct Flash vs BytePlus Seed TTS 2.0

vs

什么时候选哪个

qwen3-tts-instruct-flash 每百万字符便宜 2.6 倍($11.5 对 $30),两者也都接受自然语言的音色指令,但请求长度差别很大:600 个字符对 4096。seed-tts-2.0 还在描述式控制之外提供数值型语音参数,并按场景分组公布音色。短句要控成本选 qwen3-tts-instruct-flash,长段落和更细的音色控制选 seed-tts-2.0。

价格

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0 Δ
每 1M 字符 $11.5 $30 0.38×

价格取自构建时的实时目录,最新价格见各模型页面。

两个模型所处的位置:全部 7 个按同一单位计费的文字转语音模型的每 1M 字符价格分布(对数刻度)

能力

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0
流式输出 是 是
SSML undocumented unsupported
计费单位 character character

规格

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0
输入模态 文本 文本
输出模态 音频 音频
发布日期 - 2025-10-16
请求限制 600 characters -
音色

System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella

the Qwen-TTS voice list pairs each voice with the exact model ids that accept it.

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

语言

Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy

naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese).

English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean
声音控制
  • instructions steers speed, emotion and style in natural language, up to 1,600 tokens and documented for Chinese and English only
  • optimize_instructions (default false) rewrites the instruction into a directive better suited to synthesis and has no effect when instructions is empty
  • voice cloning and voice design are both listed as unsupported for this model id
context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context.
限制

HTTP non-real-time speech synthesis API

the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours

streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series)

output audio is free.

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

规格照录自各供应商的文档;供应商没有公布的项目,对应的行直接省略,不做推断。 完整来源: Qwen3 TTS Instruct Flash · BytePlus Seed TTS 2.0

改一行代码就能在两个模型之间切换

下方每个标签页里都有两个模型 ID,高亮的那两行是唯一要改的地方。端点不变,API key 不变,请求结构也不变。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="qwen3-tts-instruct-flash",
    # model="seed-tts-2.0",  # 取消注释此行,注释上一行
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

获取 API key →

常见问题

Qwen3 TTS Instruct Flash 和 BytePlus Seed TTS 2.0 哪个更便宜?

按「每 1M 字符」算,Qwen3 TTS Instruct Flash 更便宜($11.5 对 $30,相差 2.6×)。其他计费项的结论可能相反,完整价格见上方表格,实际成本取决于你的用量构成。

不用分别集成两次,就能对 Qwen3 TTS Instruct Flash 和 BytePlus Seed TTS 2.0 做 A/B 测试吗?

可以。两个模型走同一个 OpenAI 兼容端点,用同一个 API key,切换时只要改一行里的模型名,所以可以给两个模型各分一部分流量,直接对比账单。

文字转语音如何计费?

按输入文本的字符数计费,单次请求的字符上限见规格表。不管用哪个模型,长文本都得拆成多个请求。

相关对比

我们的实测研究