🎁 신규 무료 가입, 10회 호출 제공. 최대 $1, 카드 불필요.

Google TTS Chirp 3 HD vs BytePlus Seed TTS 2.0

vs

언제 어떤 모델을 사용할까 — 벤치마크 표가 아닌 선별된 평가

카탈로그 정보에 따르면 이 둘은 상호 교환이 가능합니다: google-tts-chirp3-hd와 seed-tts-2.0 모두 백만 입력 토큰당 $30를 청구하며, 둘 다 4096-token 컨텍스트 창을 지원하고, 동일한 speech 기능 하에서 텍스트를 입력받아 오디오를 반환합니다. 타협해야 할 요금, 컨텍스트 또는 모달리티 차이가 없으므로, 선택은 벤더 선호도(google-tts-chirp3-hd의 경우 Google, seed-tts-2.0의 경우 ByteDance)와 A/B 테스트에서 청취자가 실제로 선호하는 음성 출력에 달려 있습니다. 예산 및 프롬프트 길이 계획은 두 모델에 대해 동일하게 설정할 수 있습니다.

가격

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0 Δ
1M 문자당 $30 $30 =

빌드 시점의 라이브 카탈로그 요금입니다. 각 모델 페이지에 현재 요금표가 표시됩니다.

현재 위치 — 이 청구 단위를 사용하는 모든 7개의 텍스트 음성 변환 모델 전체의 1M 문자당 가격 (로그 스케일)

기능

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
스트리밍
SSML preview unsupported
과금 단위 character character

사양

Google TTS Chirp 3 HD BytePlus Seed TTS 2.0
입력 모달리티 텍스트 텍스트
출력 모달리티 오디오 오디오
요청 제한 5000 bytes
음성 Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages.

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

언어 ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean
음성 제어
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context.
제한

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

사양은 각 공급업체의 문서를 그대로 기록한 것입니다. 공급업체가 공개하지 않은 항목은 추론하지 않고 제외했습니다. 전체 출처: Google TTS Chirp 3 HD · BytePlus Seed TTS 2.0

코드 한 줄로 모델 전환

두 id는 아래의 모든 탭에 있습니다 — 강조 표시된 두 줄이 유일한 수정 사항입니다. 동일한 엔드포인트, 동일한 키, 동일한 요청 형태입니다.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="seed-tts-2.0",  # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API 키 발급받기 →

FAQ

Google TTS Chirp 3 HD와(과) BytePlus Seed TTS 2.0 중 어느 것이 더 저렴한가요?

동일한 1m 문자당($30)을(를) 나열하므로 가격만으로는 결정할 수 없습니다 — 아래의 사양과 기능을 확인하세요.

두 번의 연동 과정 없이 Google TTS Chirp 3 HD와(과) BytePlus Seed TTS 2.0를 A/B 테스트할 수 있나요?

네. 둘 다 하나의 API 키를 사용하여 동일한 OpenAI 호환 엔드포인트를 통해 제공됩니다 — 모델 문자열을 한 줄만 변경하면 전환되므로, 트래픽의 일부를 각각 라우팅하여 요금을 직접 비교할 수 있습니다.

텍스트 음성 변환은 어떻게 과금되나요?

입력 텍스트의 글자당 청구되며, 사양 표에 요청당 글자 수 한도가 표시되어 있습니다. 긴 스크립트는 어느 모델에서든 여러 요청으로 분할해야 합니다.

관련 비교