신규 무료로 가입하고 10회 호출해 보세요. 최대 $1, 카드 등록 불필요.

Google TTS Chirp 3 HD vs Qwen3 TTS Instruct Flash

vs

언제 어떤 모델을 쓸까

qwen3-tts-instruct-flash는 100만 자당 2.6배 저렴하고($11.5 대 $30) 평이한 말로 목소리를 지시할 수 있습니다. 다만 요청당 600자만 받고 google-tts-chirp3-hd는 5000바이트이므로 긴 원고는 쪼개야 합니다. 구글 쪽 모델은 로케일 목록이 더 넓고 수치형 음성 파라미터를 제공하며 SSML은 프리뷰입니다. 짧은 대사에 연출을 얹으려면 qwen3-tts-instruct-flash, 긴 텍스트와 로케일 폭이면 google-tts-chirp3-hd를 고르세요.

가격

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash Δ
1M 문자당 $30 $11.5 2.6×

빌드 시점에 실시간 카탈로그에서 가져온 요금입니다. 현재 요금은 각 모델 페이지에서 확인할 수 있습니다.

가격 위치: 이 과금 단위를 쓰는 음성 합성 모델 7개 전체의 1M 문자당 가격(로그 스케일)

기능

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash
스트리밍 지원 지원
SSML preview undocumented
과금 단위 character character

사양

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash
입력 모달리티 텍스트 텍스트
출력 모달리티 오디오 오디오
출시일 2025-03 -
요청 제한 5000 bytes 600 characters
보이스 Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages.

System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella

the Qwen-TTS voice list pairs each voice with the exact model ids that accept it.

언어 ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN.

Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy

naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese).

음성 제어
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
  • instructions steers speed, emotion and style in natural language, up to 1,600 tokens and documented for Chinese and English only
  • optimize_instructions (default false) rewrites the instruction into a directive better suited to synthesis and has no effect when instructions is empty
  • voice cloning and voice design are both listed as unsupported for this model id
제한

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

HTTP non-real-time speech synthesis API

the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours

streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series)

output audio is free.

사양은 각 공급자의 문서에서 그대로 옮겨 적었습니다. 공급자가 공개하지 않은 항목은 추정해 채우지 않고 뺐습니다. 전체 출처: Google TTS Chirp 3 HD · Qwen3 TTS Instruct Flash

코드 한 줄로 모델 전환

아래 탭마다 두 모델 id가 모두 들어 있습니다. 바꿀 곳은 강조된 두 줄뿐이고, 엔드포인트, 키, 요청 형식은 그대로입니다.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="qwen3-tts-instruct-flash",  # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API 키 받기 →

자주 묻는 질문

Google TTS Chirp 3 HD vs Qwen3 TTS Instruct Flash, 어느 쪽이 더 저렴한가요?

1M 문자당 기준으로는 Qwen3 TTS Instruct Flash 쪽이 더 저렴합니다($11.5 대 $30, 2.6배 차이). 다른 항목에서는 반대일 수 있습니다. 전체 요금은 위 표에 있으며, 실제 비용은 사용 패턴에 따라 달라집니다.

연동을 두 번 하지 않고도 Google TTS Chirp 3 HD vs Qwen3 TTS Instruct Flash A/B 테스트를 할 수 있나요?

네. 두 모델 모두 API 키 하나로 같은 OpenAI 호환 엔드포인트에서 호출합니다. 모델 문자열 한 줄만 바꾸면 전환되므로, 트래픽 일부를 각 모델로 보내 청구 금액을 직접 비교할 수 있습니다.

음성 합성(TTS)은 어떻게 과금되나요?

입력 텍스트의 글자 수만큼 과금되며, 요청당 글자 수 상한은 사양 표에 있습니다. 긴 스크립트는 어느 모델이든 여러 요청으로 나눠 보내야 합니다.

관련 비교

직접 측정한 자료