Neu Kostenlos registrieren und 10 Aufrufe gratis nutzen. Bis zu 1 $, ohne Kreditkarte.

Google TTS Neural2 vs Google TTS Standard

vs

Welches Modell wofür

google-tts-standard ist 4x günstiger, $4 pro Million Zeichen gegenüber $16, und trägt die breiteste Locale-Spanne der Reihe; neural2 ist die höherwertige Wiedergabe auf einer viel kürzeren Liste von 17 Locales. Keines von beiden streamt, beide nehmen 5000 Bytes pro Anfrage, beide unterstützen SSML und numerische Stimmparameter - und standard zählt ein Mehrbyte-Zeichen nur einmal, was für CJK-Texte zählt. Nehmen Sie standard für Reichweite und Kosten, neural2, wo seine Locale abgedeckt ist und Qualität zählt.

Preise

Google TTS Neural2 Google TTS Standard Δ
Pro 1M Zeichen $16 $4 4×

Preise aus dem Live-Katalog, Stand des letzten Builds. Die aktuellen Preise stehen auf der jeweiligen Modellseite.

Einordnung: Preis pro 1M Zeichen aller 7 Text-to-Speech-Modelle mit dieser Abrechnungseinheit (logarithmische Skala)

Fähigkeiten

Google TTS Neural2 Google TTS Standard
Streaming nein nein
SSML supported supported
Abrechnungseinheit character character

Spezifikationen

Google TTS Neural2 Google TTS Standard
Eingabemodalitäten Text Text
Ausgabemodalitäten Audio Audio
Veröffentlicht 2022-06-27 2018-03-27
Limit pro Anfrage 5000 bytes 5000 bytes
Stimmen Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.

Voice ids follow a locale-plus-letter pattern (en-US-Standard-A, cmn-CN-Standard-A). Google's comparison table lists Standard as cost efficient, generally available, controllable via SSML and not streaming-capable

the docs attribute the voices to parametric text-to-speech passed through vocoders.

Sprachen A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. The widest locale span of any Cloud TTS voice type. Standard voice ids are published for af-ZA, ar-XA, bg-BG, bn-IN, ca-ES, cmn-CN, cmn-TW, cs-CZ, da-DK, de-DE, el-GR, en-AU, en-GB, en-IN, en-US, es-ES, es-US, et-EE, eu-ES, fi-FI, fil-PH, fr-CA, fr-FR, gl-ES, gu-IN, he-IL, hi-IN, hu-HU, id-ID, is-IS, it-IT, ja-JP, kn-IN, ko-KR, lt-LT, lv-LV, ml-IN, mr-IN, ms-MY, nb-NO, nl-BE, nl-NL, pa-IN, pl-PL, pt-BR, pt-PT, ro-RO, ru-RU, sk-SK, sr-RS, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN, vi-VN and yue-HK.
Stimmsteuerung
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
Limits

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Content limit of 5,000 total bytes per synthesize request (a single character is multiple bytes in some locales). Output LINEAR16 (returned with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis

Long Audio Synthesis (Preview) covers up to 1 million bytes of input asynchronously. Billed per character including spaces and newlines, and all SSML tags except <mark> count

for Standard and WaveNet a multi-byte character is charged once.

Die Spezifikationen stammen aus der Dokumentation des jeweiligen Anbieters. Veröffentlicht ein Anbieter eine Angabe nicht, lassen wir die Zeile weg, statt sie zu schätzen. Vollständige Quellen: Google TTS Neural2 · Google TTS Standard

Eine Zeile genügt für den Wechsel

Beide IDs stehen in jedem Tab unten. Die beiden hervorgehobenen Zeilen sind die einzige Änderung. Gleicher Endpunkt, gleicher Schlüssel, gleiches Anfrageformat.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-neural2",
    # model="google-tts-standard",  # diese Zeile einkommentieren, die darüberliegende auskommentieren
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API-Schlüssel erstellen →

FAQ

Welches Modell ist günstiger: Google TTS Neural2 oder Google TTS Standard?

Google TTS Standard ist bei Pro 1M Zeichen günstiger ($4 vs. $16, Faktor 4.0). Bei anderen Zeilen kann es umgekehrt sein. Die Tabelle oben zeigt alle Preise, und die tatsächlichen Kosten hängen von Ihrem Mix ab.

Kann ich Google TTS Neural2 gegen Google TTS Standard A/B-testen, ohne zweimal zu integrieren?

Ja. Beide laufen über denselben OpenAI-kompatiblen Endpunkt mit einem API-Schlüssel. Für den Wechsel ändern Sie nur eine Zeile, den Modellnamen. So können Sie einen Teil des Traffics an jedes Modell schicken und die Kosten direkt vergleichen.

Wie wird Text-to-Speech abgerechnet?

Pro Zeichen des Eingabetexts, mit einer Obergrenze für die Zeichenzahl pro Anfrage, die in der Spezifikationstabelle steht. Lange Texte müssen Sie bei beiden Modellen auf mehrere Anfragen aufteilen.

Verwandte Vergleiche

Aus unseren Studien mit eigenen Messungen