Novo Cadastre-se grátis, 10 chamadas por nossa conta. Até US$ 1, sem cartão.

Google TTS Neural2 vs Google TTS Standard

vs

Qual usar e quando

O google-tts-standard é 4x mais barato, $4 por milhão de caracteres contra $16, e tem a maior amplitude de locais da linha; o neural2 é a síntese de maior qualidade numa lista bem mais curta de 17 locais. Nenhum faz streaming, os dois admitem 5000 bytes por requisição, os dois suportam SSML e parâmetros numéricos de voz - e o standard conta um caractere multibyte uma única vez, o que importa em texto CJK. Escolha o standard por alcance e custo, o neural2 onde seu local esteja coberto e a qualidade importe.

Preços

Google TTS Neural2 Google TTS Standard Δ
Por 1M caracteres $16 $4 4×

Tarifas do catálogo em tempo real, lidas no momento do build; a página de cada modelo traz a tabela de preços atual.

Onde cada um fica: preço por 1M caracteres entre os 7 modelos de síntese de voz com esta unidade de cobrança (escala logarítmica)

Capacidades

Google TTS Neural2 Google TTS Standard
Streaming não não
SSML supported supported
Unidade de cobrança character character

Especificações

Google TTS Neural2 Google TTS Standard
Modalidades de entrada texto texto
Modalidades de saída áudio áudio
Lançamento 2022-06-27 2018-03-27
Limite por requisição 5000 bytes 5000 bytes
Vozes Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.

Voice ids follow a locale-plus-letter pattern (en-US-Standard-A, cmn-CN-Standard-A). Google's comparison table lists Standard as cost efficient, generally available, controllable via SSML and not streaming-capable

the docs attribute the voices to parametric text-to-speech passed through vocoders.

Idiomas A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. The widest locale span of any Cloud TTS voice type. Standard voice ids are published for af-ZA, ar-XA, bg-BG, bn-IN, ca-ES, cmn-CN, cmn-TW, cs-CZ, da-DK, de-DE, el-GR, en-AU, en-GB, en-IN, en-US, es-ES, es-US, et-EE, eu-ES, fi-FI, fil-PH, fr-CA, fr-FR, gl-ES, gu-IN, he-IL, hi-IN, hu-HU, id-ID, is-IS, it-IT, ja-JP, kn-IN, ko-KR, lt-LT, lv-LV, ml-IN, mr-IN, ms-MY, nb-NO, nl-BE, nl-NL, pa-IN, pl-PL, pt-BR, pt-PT, ro-RO, ru-RU, sk-SK, sr-RS, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN, vi-VN and yue-HK.
Controle de voz
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
Limites

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Content limit of 5,000 total bytes per synthesize request (a single character is multiple bytes in some locales). Output LINEAR16 (returned with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis

Long Audio Synthesis (Preview) covers up to 1 million bytes of input asynchronously. Billed per character including spaces and newlines, and all SSML tags except <mark> count

for Standard and WaveNet a multi-byte character is charged once.

As especificações são transcritas da documentação de cada provedor; quando um provedor não publica um dado, a linha é omitida, e não inferida. Fontes completas: Google TTS Neural2 · Google TTS Standard

Troque de um para o outro mudando uma linha

Os dois ids estão em todas as abas abaixo; o par de linhas destacado é a única alteração. O endpoint, a chave e o formato da requisição são os mesmos.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-neural2",
    # model="google-tts-standard",  # descomente esta linha, comente a linha acima
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

Obtenha sua chave de API →

Perguntas frequentes

Qual é mais barato, Google TTS Neural2 ou Google TTS Standard?

Google TTS Standard sai mais barato na linha “Por 1M caracteres” ($4 contra $16, 4.0× de diferença). Outras linhas podem pender para o outro lado: a tabela acima mostra todos os preços, e o custo real depende do seu mix de uso.

Posso fazer um teste A/B de Google TTS Neural2 contra Google TTS Standard sem duas integrações?

Sim. Os dois são servidos pelo mesmo endpoint compatível com OpenAI, com uma única chave de API. Para trocar, basta mudar uma linha, a string do modelo; assim você pode direcionar uma parte do tráfego para cada um e comparar as contas diretamente.

Como o text-to-speech é cobrado?

Por caractere do texto de entrada, com um limite de caracteres por requisição indicado na tabela de especificações. Nos dois modelos, textos longos precisam ser divididos em várias requisições.

Comparações relacionadas

Dos nossos estudos com medições