新人 免费注册,送 10 次调用,最高 $1,无需绑卡。

Google TTS Neural2 vs Google TTS Standard

vs

什么时候选哪个

google-tts-standard 便宜 4 倍,每百万字符 $4 对 $16,并且是这条产品线里语言区跨度最广的一个;neural2 是质量更高的合成,但语言区列表短得多,只有 17 个。两者都不支持流式,每次请求都接受 5000 字节,都支持 SSML 和数值型语音参数——而且 standard 对多字节字符只计一次,这对中日韩文案很关键。要覆盖面和成本选 standard,语言区已被覆盖且看重质量时选 neural2。

价格

Google TTS Neural2 Google TTS Standard Δ
每 1M 字符 $16 $4 4×

价格取自构建时的实时目录,最新价格见各模型页面。

两个模型所处的位置:全部 7 个按同一单位计费的文字转语音模型的每 1M 字符价格分布(对数刻度)

能力

Google TTS Neural2 Google TTS Standard
流式输出 否 否
SSML supported supported
计费单位 character character

规格

Google TTS Neural2 Google TTS Standard
输入模态 文本 文本
输出模态 音频 音频
发布日期 2022-06-27 2018-03-27
请求限制 5000 bytes 5000 bytes
音色 Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.

Voice ids follow a locale-plus-letter pattern (en-US-Standard-A, cmn-CN-Standard-A). Google's comparison table lists Standard as cost efficient, generally available, controllable via SSML and not streaming-capable

the docs attribute the voices to parametric text-to-speech passed through vocoders.

语言 A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. The widest locale span of any Cloud TTS voice type. Standard voice ids are published for af-ZA, ar-XA, bg-BG, bn-IN, ca-ES, cmn-CN, cmn-TW, cs-CZ, da-DK, de-DE, el-GR, en-AU, en-GB, en-IN, en-US, es-ES, es-US, et-EE, eu-ES, fi-FI, fil-PH, fr-CA, fr-FR, gl-ES, gu-IN, he-IL, hi-IN, hu-HU, id-ID, is-IS, it-IT, ja-JP, kn-IN, ko-KR, lt-LT, lv-LV, ml-IN, mr-IN, ms-MY, nb-NO, nl-BE, nl-NL, pa-IN, pl-PL, pt-BR, pt-PT, ro-RO, ru-RU, sk-SK, sr-RS, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN, vi-VN and yue-HK.
声音控制
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
限制

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Content limit of 5,000 total bytes per synthesize request (a single character is multiple bytes in some locales). Output LINEAR16 (returned with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis

Long Audio Synthesis (Preview) covers up to 1 million bytes of input asynchronously. Billed per character including spaces and newlines, and all SSML tags except <mark> count

for Standard and WaveNet a multi-byte character is charged once.

规格照录自各供应商的文档;供应商没有公布的项目,对应的行直接省略,不做推断。 完整来源: Google TTS Neural2 · Google TTS Standard

改一行代码就能在两个模型之间切换

下方每个标签页里都有两个模型 ID,高亮的那两行是唯一要改的地方。端点不变,API key 不变,请求结构也不变。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-neural2",
    # model="google-tts-standard",  # 取消注释此行,注释上一行
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

获取 API key →

常见问题

Google TTS Neural2 和 Google TTS Standard 哪个更便宜?

按「每 1M 字符」算,Google TTS Standard 更便宜($4 对 $16,相差 4.0×)。其他计费项的结论可能相反,完整价格见上方表格,实际成本取决于你的用量构成。

不用分别集成两次,就能对 Google TTS Neural2 和 Google TTS Standard 做 A/B 测试吗?

可以。两个模型走同一个 OpenAI 兼容端点,用同一个 API key,切换时只要改一行里的模型名,所以可以给两个模型各分一部分流量,直接对比账单。

文字转语音如何计费?

按输入文本的字符数计费,单次请求的字符上限见规格表。不管用哪个模型,长文本都得拆成多个请求。

相关对比

我们的实测研究