新規 無料登録で、呼び出し 10 回分(最大 $1)をお試しいただけます。カード登録は不要です。

Google TTS Neural2 vs Google TTS Standard

vs

どんなときに、どちらを選ぶか

google-tts-standard は 4 分の 1 の 100万文字あたり $4 対 $16 で、この系列で最も広いロケール範囲を持ちます。neural2 は品質の高い合成ですが、ロケールは 17 とはるかに短い一覧です。どちらもストリーミングせず、1 リクエスト 5000 バイトを受け付け、SSML と数値パラメータによる音声制御に対応します。さらに standard はマルチバイト文字を 1 文字として数えるので、CJK の原稿では効いてきます。到達範囲とコストなら standard、対象ロケールが揃っていて品質が要るなら neural2 を選んでください。

料金

Google TTS Neural2 Google TTS Standard Δ
100 万文字あたり $16 $4 4×

ビルド時点の現行カタログの料金です。最新の料金表は各モデルのページに掲載しています。

両モデルの位置:100 万文字あたりの料金(この課金単位の音声合成モデル全 7 件、対数スケール)

対応機能

Google TTS Neural2 Google TTS Standard
ストリーミング なし なし
SSML supported supported
課金単位 character character

仕様

Google TTS Neural2 Google TTS Standard
入力モダリティ テキスト テキスト
出力モダリティ 音声 音声
リリース 2022-06-27 2018-03-27
リクエストあたりの上限 5000 bytes 5000 bytes
ボイス Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.

Voice ids follow a locale-plus-letter pattern (en-US-Standard-A, cmn-CN-Standard-A). Google's comparison table lists Standard as cost efficient, generally available, controllable via SSML and not streaming-capable

the docs attribute the voices to parametric text-to-speech passed through vocoders.

言語 A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. The widest locale span of any Cloud TTS voice type. Standard voice ids are published for af-ZA, ar-XA, bg-BG, bn-IN, ca-ES, cmn-CN, cmn-TW, cs-CZ, da-DK, de-DE, el-GR, en-AU, en-GB, en-IN, en-US, es-ES, es-US, et-EE, eu-ES, fi-FI, fil-PH, fr-CA, fr-FR, gl-ES, gu-IN, he-IL, hi-IN, hu-HU, id-ID, is-IS, it-IT, ja-JP, kn-IN, ko-KR, lt-LT, lv-LV, ml-IN, mr-IN, ms-MY, nb-NO, nl-BE, nl-NL, pa-IN, pl-PL, pt-BR, pt-PT, ro-RO, ru-RU, sk-SK, sr-RS, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN, vi-VN and yue-HK.
ボイス調整
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
制限

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Content limit of 5,000 total bytes per synthesize request (a single character is multiple bytes in some locales). Output LINEAR16 (returned with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis

Long Audio Synthesis (Preview) covers up to 1 million bytes of input asynchronously. Billed per character including spaces and newlines, and all SSML tags except <mark> count

for Standard and WaveNet a multi-byte character is charged once.

仕様は各ベンダーのドキュメントから転記しています。ベンダーが公表していない項目は、推測で埋めずに省いています。 出典の一覧: Google TTS Neural2 · Google TTS Standard

1 行の変更で切り替え

以下のどのタブにも両方のモデル ID が入っています。書き換えるのはハイライトされた 2 行だけで、エンドポイント、キー、リクエスト形式はすべて同じです。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-neural2",
    # model="google-tts-standard",  # この行のコメントを外し、上の行をコメントアウト
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API キーを取得 →

よくある質問

Google TTS Neural2 と Google TTS Standard ではどちらが安いですか?

100 万文字あたりは Google TTS Standard のほうが安くなります($4 対 $16、4.0 倍の差)。ほかの行では逆になることもあります。料金の全体は上の表に掲載しており、実際のコストは利用の内訳によって変わります。

Google TTS Neural2 と Google TTS Standard の A/B テストは、実装を 2 つ用意せずにできますか?

はい。どちらも同じ OpenAI 互換エンドポイントから 1 つの API キーで利用できます。切り替えはモデル名の文字列を 1 行変えるだけなので、トラフィックの一部をそれぞれに振り分けて、請求額を直接比較できます。

音声合成(Text-to-Speech)の課金方法は?

入力テキストの文字数に応じて課金されます。1 リクエストあたりの文字数上限は仕様表に掲載しています。どちらのモデルでも、長い原稿は複数のリクエストに分けて送る必要があります。

関連する比較

当社の実測レポートより