🎁 New Sign up free, 10 calls on us. Up to $1, no card needed.

Google TTS Chirp 3 HD vs BytePlus Seed TTS 2.0

vs

Which one, when — curated verdict, not a benchmark table

On the catalogue facts these two are interchangeable: google-tts-chirp3-hd and seed-tts-2.0 both bill $30 per million input tokens, both carry a 4096-token context window, and both take text in and return audio under the same speech capability. With no rate, context, or modality gap to trade off, the choice comes down to vendor preference — Google for google-tts-chirp3-hd, ByteDance for seed-tts-2.0 — and to whichever voice output your listeners actually prefer in an A/B test. Budget and prompt-length planning can be identical for either.

Pricing

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0 Δ
Per 1M characters $30 $30 =

Rates from the live catalog at build time; each model page carries the current card.

Where they sit — price per 1M characters across all 7 text-to-speech models on this billing unit (log scale)

Capabilities

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
Streaming yes yes
SSML preview unsupported
Billing unit character character

Specs

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
Input modalities text text
Output modalities audio audio
Request limit 5000 bytes
Voices Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages.

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

Languages ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean
Voice control
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context.
Limits

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

Specs are transcribed from each vendor’s documentation; a row a vendor does not publish is left out rather than inferred. Full sources: Google TTS Chirp 3 HD · BytePlus Seed TTS 2.0

Switch between them with one line

Both ids are in every tab below — the highlighted pair of lines is the only edit. Same endpoint, same key, same request shape.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="seed-tts-2.0",  # uncomment this line, comment the one above
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

Get an API key →

FAQ

Which is cheaper, Google TTS Chirp 3 HD or BytePlus Seed TTS 2.0?

They list the same per 1m characters ($30), so price does not decide this one — see the specs and capabilities below.

Can I A/B test Google TTS Chirp 3 HD against BytePlus Seed TTS 2.0 without two integrations?

Yes. Both are served through the same OpenAI-compatible endpoint with one API key — switching is a one-line model-string change, so you can route a fraction of traffic to each and compare bills directly.

How is text-to-speech billed?

Per character of input text, with a per-request character ceiling shown in the spec table. Long scripts must be chunked across requests on either model.

Related comparisons