🎁 Neu Kostenlos registrieren, 10 Aufrufe gratis. Bis zu 1 $, ohne Karte.

Google TTS Chirp 3 HD vs Google TTS Neural2

vs

Welches und wann — kuratiertes Fazit, keine Benchmark-Tabelle

google-tts-neural2 ist 1.9x günstiger ($16 pro Million Zeichen gegenüber $30) und unterstützt SSML regulär, während google-tts-chirp3-hd es als Vorschau führt — aber neural2 streamt nicht und veröffentlicht eine deutlich kürzere Locale-Liste, 17 Locales mit benannten Stimm-IDs. Beide nehmen 5000 Bytes pro Anfrage und bieten numerische Stimmparameter. Nehmen Sie google-tts-neural2 für SSML-getriebene Stapelarbeit in einer abgedeckten Locale, google-tts-chirp3-hd für Streaming und Locale-Breite.

Preise

Google TTS Chirp 3 HD Google TTS Neural2 Δ
Pro 1M Zeichen $30 $16 1.9×

Preise aus dem Live-Katalog zum Zeitpunkt des Builds; jede Modellseite enthält die aktuelle Übersicht.

Wo sie stehen — Preis pro 1M Zeichen über alle 7 Text-to-Speech-Modelle mit dieser Abrechnungseinheit (logarithmische Skala)

Fähigkeiten

Google TTS Chirp 3 HD Google TTS Neural2
Streaming ja nein
SSML preview supported
Abrechnungseinheit character character

Spezifikationen

Google TTS Chirp 3 HD Google TTS Neural2
Input-Modalitäten Text Text
Ausgabemodalitäten Audio Audio
Request-Limit 5000 bytes 5000 bytes
Stimmen Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages. Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own.
Sprachen ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints.
Sprachsteuerung
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
  • AudioConfig controls: speakingRate 0.25 to 2.0 (1.0 native)
  • pitch -20.0 to 20.0 semitones
  • volumeGainDb -96.0 to 16.0
  • effectsProfileId device profiles for wearable, handset, headphone, small and medium Bluetooth speaker, large home entertainment, large automotive and telephony playback
  • SSML tags include speak, break, say-as, sub, mark, prosody, emphasis, phoneme, voice and lang
Limits

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW

optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted

the multi-byte-counts-once note applies to Standard and WaveNet only.

Die Spezifikationen sind aus der Dokumentation der jeweiligen Anbieter übernommen; eine Zeile, die ein Anbieter nicht veröffentlicht, wird weggelassen und nicht abgeleitet. Vollständige Quellen: Google TTS Chirp 3 HD · Google TTS Neural2

Mit einer Zeile zwischen ihnen wechseln

Beide IDs befinden sich in jedem Tab unten — das hervorgehobene Zeilenpaar ist die einzige Änderung. Gleicher Endpunkt, gleicher Schlüssel, gleiche Request-Struktur.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="google-tts-neural2",  # diese Zeile einkommentieren, die darüberliegende auskommentieren
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API-Schlüssel holen →

FAQ

Welches ist günstiger, Google TTS Chirp 3 HD oder Google TTS Neural2?

Google TTS Neural2 ist günstiger bei pro 1m zeichen ($16 vs. $30, 1.9× Unterschied). Andere Zeilen können in die andere Richtung deuten — die obige Tabelle enthält alle Daten, und die tatsächlichen Kosten hängen von Ihrem Mix ab.

Kann ich Google TTS Chirp 3 HD gegen Google TTS Neural2 ohne zwei Integrationen A/B-testen?

Ja. Beide werden über denselben OpenAI-kompatiblen Endpunkt mit einem API-Schlüssel bereitgestellt — der Wechsel ist eine einzeilige Änderung des Modell-Strings, sodass Sie einen Bruchteil des Traffics an jedes Modell leiten und die Rechnungen direkt vergleichen können.

Wie wird Text-to-Speech abgerechnet?

Pro Zeichen des Eingabetextes, mit einer Zeichenobergrenze pro Anfrage, die in der Spezifikationstabelle angegeben ist. Lange Skripte müssen bei beiden Modellen über mehrere Anfragen hinweg aufgeteilt werden.

Verwandte Vergleiche