신규 무료로 가입하고 10회 호출해 보세요. 최대 $1, 카드 등록 불필요.

Google TTS Chirp 3 HD vs TTS-1 HD

vs

언제 어떤 모델을 쓸까

둘 다 100만 자당 $30이고 둘 다 스트리밍하므로 갈리는 지점은 제어와 길이입니다. google-tts-chirp3-hd는 요청당 5000바이트를 받고 수치형 음성 파라미터를 제공하며 SSML은 프리뷰입니다. tts-1-hd는 4096자를 받고 음성 제어가 전혀 없습니다. 지원 범위는 둘 다 넓지만 tts-1-hd의 목록은 열거된 로케일이 아니라 Whisper의 언어 집합을 따릅니다. 전달 방식을 다듬어야 하면 google-tts-chirp3-hd, 기본 목소리로 충분하면 tts-1-hd를 고르세요.

가격

Google TTS Chirp 3 HD TTS-1 HD Δ
1M 문자당 $30 $30 =

빌드 시점에 실시간 카탈로그에서 가져온 요금입니다. 현재 요금은 각 모델 페이지에서 확인할 수 있습니다.

가격 위치: 이 과금 단위를 쓰는 음성 합성 모델 7개 전체의 1M 문자당 가격(로그 스케일)

기능

Google TTS Chirp 3 HD TTS-1 HD
스트리밍 지원 지원
SSML preview undocumented
과금 단위 character character

사양

Google TTS Chirp 3 HD TTS-1 HD
입력 모달리티 텍스트 텍스트
출력 모달리티 오디오 오디오
출시일 2025-03 2023-11-06
요청 제한 5000 bytes 4096 characters
보이스 Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages. alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer, a smaller set than the full 13-voice TTS list, and voices are currently optimized for English.
언어 ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. Generally follows the Whisper model’s language support: Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese and Welsh, despite voices being optimized for English
음성 제어
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
none
제한

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

Speech generation endpoint (v1/audio/speech) only, with Chat Completions, Responses, Realtime, Batch and Fine-tuning all listed as not supported

input text is capped at 4096 characters

output mp3 (default), opus, aac, flac, wav or pcm (raw 24 kHz 16-bit signed little-endian samples)

realtime playback via chunked transfer encoding

사양은 각 공급자의 문서에서 그대로 옮겨 적었습니다. 공급자가 공개하지 않은 항목은 추정해 채우지 않고 뺐습니다. 전체 출처: Google TTS Chirp 3 HD · TTS-1 HD

코드 한 줄로 모델 전환

아래 탭마다 두 모델 id가 모두 들어 있습니다. 바꿀 곳은 강조된 두 줄뿐이고, 엔드포인트, 키, 요청 형식은 그대로입니다.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="tts-1-hd",  # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API 키 받기 →

자주 묻는 질문

Google TTS Chirp 3 HD vs TTS-1 HD, 어느 쪽이 더 저렴한가요?

두 모델의 1M 문자당 요금이 같습니다($30). 가격만으로는 가릴 수 없으니 아래 사양과 기능을 확인하세요.

연동을 두 번 하지 않고도 Google TTS Chirp 3 HD vs TTS-1 HD A/B 테스트를 할 수 있나요?

네. 두 모델 모두 API 키 하나로 같은 OpenAI 호환 엔드포인트에서 호출합니다. 모델 문자열 한 줄만 바꾸면 전환되므로, 트래픽 일부를 각 모델로 보내 청구 금액을 직접 비교할 수 있습니다.

음성 합성(TTS)은 어떻게 과금되나요?

입력 텍스트의 글자 수만큼 과금되며, 요청당 글자 수 상한은 사양 표에 있습니다. 긴 스크립트는 어느 모델이든 여러 요청으로 나눠 보내야 합니다.

관련 비교

직접 측정한 자료