新人 免费注册,送 10 次调用,最高 $1,无需绑卡。

Google TTS Chirp 3 HD vs BytePlus Seed TTS 2.0

vs

什么时候选哪个

根据产品参数,这两者是可互换的:google-tts-chirp3-hd 和 seed-tts-2.0 均收取每百万输入 token $30,都具备 4096 token 的上下文窗口,且都在相同的语音能力下接受文本输入并返回音频。由于没有费率、上下文或模态差异需要权衡,选择的决定因素在于供应商偏好——google-tts-chirp3-hd 属于 Google,seed-tts-2.0 属于 ByteDance——以及你的听众在 A/B 测试中更喜欢哪种语音输出。两者的预算和提示词长度规划可以完全相同。

价格

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0 Δ
每 1M 字符 $30 $30 =

价格取自构建时的实时目录,最新价格见各模型页面。

两个模型所处的位置:全部 7 个按同一单位计费的文字转语音模型的每 1M 字符价格分布(对数刻度)

能力

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
流式输出 是 是
SSML preview unsupported
计费单位 character character

规格

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
输入模态 文本 文本
输出模态 音频 音频
发布日期 2025-03 2025-10-16
请求限制 5000 bytes -
音色 Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages.

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

语言 ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean
声音控制
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context.
限制

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

规格照录自各供应商的文档;供应商没有公布的项目,对应的行直接省略,不做推断。 完整来源: Google TTS Chirp 3 HD · BytePlus Seed TTS 2.0

改一行代码就能在两个模型之间切换

下方每个标签页里都有两个模型 ID,高亮的那两行是唯一要改的地方。端点不变,API key 不变,请求结构也不变。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="seed-tts-2.0",  # 取消注释此行,注释上一行
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

获取 API key →

常见问题

Google TTS Chirp 3 HD 和 BytePlus Seed TTS 2.0 哪个更便宜?

两个模型的「每 1M 字符」价格相同($30),所以这一项上价格分不出高下,请看下方的规格与能力。

不用分别集成两次,就能对 Google TTS Chirp 3 HD 和 BytePlus Seed TTS 2.0 做 A/B 测试吗?

可以。两个模型走同一个 OpenAI 兼容端点,用同一个 API key,切换时只要改一行里的模型名,所以可以给两个模型各分一部分流量,直接对比账单。

文字转语音如何计费?

按输入文本的字符数计费,单次请求的字符上限见规格表。不管用哪个模型,长文本都得拆成多个请求。

相关对比

我们的实测研究