新帳號 免費註冊,送 10 次呼叫,最高 $1,免綁卡。

Google TTS Neural2 vs Google TTS Standard

vs

什麼情況選哪一個

google-tts-standard 便宜 4 倍,每百萬字元 $4 對 $16,並且是這條產品線裡語言區跨度最廣的一個;neural2 是品質更高的合成,但語言區列表短得多,只有 17 個。兩者都不支援串流,每次請求都接受 5000 位元組,都支援 SSML 和數值型語音參數——而且 standard 對多位元組字元只計一次,這對中日韓文案很關鍵。要覆蓋面和成本選 standard,語言區已被涵蓋且看重品質時選 neural2。

定價

Google TTS Neural2 Google TTS Standard Δ
每 1M 字元 $16 $4 4×

費率取自網站建置時的即時目錄;各模型頁面都列有最新的價目。

兩者的相對位置:每 1M 字元的價格,涵蓋同一計費單位下全部 7 個文字轉語音模型(對數尺度)

功能

Google TTS Neural2 Google TTS Standard
串流 否 否
SSML supported supported
計費單位 character character

規格

Google TTS Neural2 Google TTS Standard
輸入模態 文字 文字
輸出模態 音訊 音訊
發布日期 2022-06-27 2018-03-27
請求限制 5000 bytes 5000 bytes
聲音 Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.

Voice ids follow a locale-plus-letter pattern (en-US-Standard-A, cmn-CN-Standard-A). Google's comparison table lists Standard as cost efficient, generally available, controllable via SSML and not streaming-capable

the docs attribute the voices to parametric text-to-speech passed through vocoders.

語言 A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. The widest locale span of any Cloud TTS voice type. Standard voice ids are published for af-ZA, ar-XA, bg-BG, bn-IN, ca-ES, cmn-CN, cmn-TW, cs-CZ, da-DK, de-DE, el-GR, en-AU, en-GB, en-IN, en-US, es-ES, es-US, et-EE, eu-ES, fi-FI, fil-PH, fr-CA, fr-FR, gl-ES, gu-IN, he-IL, hi-IN, hu-HU, id-ID, is-IS, it-IT, ja-JP, kn-IN, ko-KR, lt-LT, lv-LV, ml-IN, mr-IN, ms-MY, nb-NO, nl-BE, nl-NL, pa-IN, pl-PL, pt-BR, pt-PT, ro-RO, ru-RU, sk-SK, sr-RS, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN, vi-VN and yue-HK.
聲音控制
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
限制

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Content limit of 5,000 total bytes per synthesize request (a single character is multiple bytes in some locales). Output LINEAR16 (returned with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis

Long Audio Synthesis (Preview) covers up to 1 million bytes of input asynchronously. Billed per character including spaces and newlines, and all SSML tags except <mark> count

for Standard and WaveNet a multi-byte character is charged once.

規格摘錄自各供應商的文件;供應商沒有公布的項目就直接略過,不自行推測。 完整來源: Google TTS Neural2 · Google TTS Standard

改一行程式碼就能在兩者之間切換

下面每個頁籤都列了這兩個模型 ID,要改的只有醒目標示的那兩行。端點、金鑰和請求格式都不變。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-neural2",
    # model="google-tts-standard",  # 取消這一行的註解,並把上一行註解掉
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

取得 API 金鑰 →

常見問題

Google TTS Neural2 和 Google TTS Standard 哪個比較便宜?

以「每 1M 字元」來看,Google TTS Standard 比較便宜($4 對 $16,相差 4.0×)。其他項目的結果可能相反,完整價目請看上表;實際成本要看你的用量組合。

可以只串接一次,就對 Google TTS Neural2 和 Google TTS Standard 做 A/B 測試嗎?

可以。兩個模型都走同一個 OpenAI 相容端點,用的也是同一把 API 金鑰,切換時只要改一行裡的模型名稱字串。你可以把一部分流量分別導到兩邊,再直接比較帳單。

文字轉語音是如何計費的?

依輸入文字的字元數計費,每次請求的字元上限列在規格表裡。兩個模型都一樣,稿子太長就得拆成多次請求。

相關比較