신규 무료로 가입하고 10회 호출해 보세요. 최대 $1, 카드 등록 불필요.

BytePlus Seed TTS 2.0 vs TTS-1 HD

vs

언제 어떤 모델을 쓸까

제공된 정보에 따르면 이 둘은 상호 교환이 가능합니다: seed-tts-2.0과 tts-1-hd 모두 텍스트를 입력받아 오디오를 반환하고, speech 기능을 지원하며, 4096-token 컨텍스트로 제한되고, 입력에 대해 백만 토큰당 $30를 청구합니다. 요금표와 단위가 정확히 일치하므로, 어느 쪽이든 비용이나 컨텍스트에 대한 논쟁의 여지가 없습니다 - 벤더 선호도, 제품에 적합한 음성 출력, 또는 이미 연동된 제공업체를 기준으로 선택하십시오. 자체 스크립트의 짧은 샘플을 각각에 실행해 보고 적절하게 들리는 것을 선택하십시오.

가격

BytePlus Seed TTS 2.0 TTS-1 HD Δ
1M 문자당 $30 $30 =

빌드 시점에 실시간 카탈로그에서 가져온 요금입니다. 현재 요금은 각 모델 페이지에서 확인할 수 있습니다.

가격 위치: 이 과금 단위를 쓰는 음성 합성 모델 7개 전체의 1M 문자당 가격(로그 스케일)

기능

BytePlus Seed TTS 2.0 TTS-1 HD
스트리밍 지원 지원
SSML unsupported undocumented
과금 단위 character character

사양

BytePlus Seed TTS 2.0 TTS-1 HD
입력 모달리티 텍스트 텍스트
출력 모달리티 오디오 오디오
출시일 2025-10-16 2023-11-06
요청 제한 - 4096 characters
보이스

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer, a smaller set than the full 13-voice TTS list, and voices are currently optimized for English.
언어 English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean Generally follows the Whisper model’s language support: Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese and Welsh, despite voices being optimized for English
음성 제어 context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context. none
제한

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

Speech generation endpoint (v1/audio/speech) only, with Chat Completions, Responses, Realtime, Batch and Fine-tuning all listed as not supported

input text is capped at 4096 characters

output mp3 (default), opus, aac, flac, wav or pcm (raw 24 kHz 16-bit signed little-endian samples)

realtime playback via chunked transfer encoding

사양은 각 공급자의 문서에서 그대로 옮겨 적었습니다. 공급자가 공개하지 않은 항목은 추정해 채우지 않고 뺐습니다. 전체 출처: BytePlus Seed TTS 2.0 · TTS-1 HD

코드 한 줄로 모델 전환

아래 탭마다 두 모델 id가 모두 들어 있습니다. 바꿀 곳은 강조된 두 줄뿐이고, 엔드포인트, 키, 요청 형식은 그대로입니다.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="seed-tts-2.0",
    # model="tts-1-hd",  # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API 키 받기 →

자주 묻는 질문

BytePlus Seed TTS 2.0 vs TTS-1 HD, 어느 쪽이 더 저렴한가요?

두 모델의 1M 문자당 요금이 같습니다($30). 가격만으로는 가릴 수 없으니 아래 사양과 기능을 확인하세요.

연동을 두 번 하지 않고도 BytePlus Seed TTS 2.0 vs TTS-1 HD A/B 테스트를 할 수 있나요?

네. 두 모델 모두 API 키 하나로 같은 OpenAI 호환 엔드포인트에서 호출합니다. 모델 문자열 한 줄만 바꾸면 전환되므로, 트래픽 일부를 각 모델로 보내 청구 금액을 직접 비교할 수 있습니다.

음성 합성(TTS)은 어떻게 과금되나요?

입력 텍스트의 글자 수만큼 과금되며, 요청당 글자 수 상한은 사양 표에 있습니다. 긴 스크립트는 어느 모델이든 여러 요청으로 나눠 보내야 합니다.

관련 비교

직접 측정한 자료