🎁 新人 免费注册,送 10 次调用,最高 $1,免绑卡。

BytePlus Seed TTS 2.0 vs TTS-1 HD

vs

何时使用哪一个 — 经过整理的结论,而非基准测试表格

根据所提供的事实,这两者是可互换的:seed-tts-2.0 和 tts-1-hd 均接受文本输入并返回音频,两者都具备语音能力,上下文上限均为 4096 token,且对输入均收取每百万 token $30。由于价格表和单位完全一致,任何一方都没有成本或上下文优势方面的理由——请根据供应商偏好、适合你产品的语音输出,或者你已经集成了哪家提供商来做选择。在两者中分别运行一段你的短脚本样本,并保留听起来合适的那个。

定价

BytePlus Seed TTS 2.0 TTS-1 HD Δ
每 1M 字符 $30 $30 =

费率取自构建时的实时目录;每个模型页面均附有当前的费率卡。

它们的位置 — 以该计费单位计费的所有 7 个 文本转语音 模型的 每 1M 字符的价格(对数刻度)

能力

BytePlus Seed TTS 2.0 TTS-1 HD
流式传输
SSML unsupported undocumented
计费单位 character character

规格

BytePlus Seed TTS 2.0 TTS-1 HD
输入模态 文本 文本
输出模态 音频 音频
请求限制 4096 characters
声音

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer, a smaller set than the full 13-voice TTS list, and voices are currently optimized for English.
语言 English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean Generally follows the Whisper model’s language support: Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese and Welsh, despite voices being optimized for English
语音控制 context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context. none
限制

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

Speech generation endpoint (v1/audio/speech) only, with Chat Completions, Responses, Realtime, Batch and Fine-tuning all listed as not supported

input text is capped at 4096 characters

output mp3 (default), opus, aac, flac, wav or pcm (raw 24 kHz 16-bit signed little-endian samples)

realtime playback via chunked transfer encoding

规格转录自各供应商的文档;若供应商未发布某项数据,则直接省略该行,而非进行推断。 完整来源: BytePlus Seed TTS 2.0 · TTS-1 HD

只需一行代码即可在它们之间切换

两个 ID 都包含在下方的每个选项卡中 — 高亮显示的两行是唯一的修改。相同的端点,相同的密钥,相同的请求结构。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="seed-tts-2.0",
    # model="tts-1-hd",  # 取消注释此行,注释上一行
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

获取 API 密钥 →

常见问题

BytePlus Seed TTS 2.0 和 TTS-1 HD 哪个更便宜?

它们列出的 每 1m 字符 相同($30),因此价格不是这一项的决定因素——请参考下文的规格与能力。

我可以在不进行两次集成的情况下,对 BytePlus Seed TTS 2.0 和 TTS-1 HD 进行 A/B 测试吗?

可以。两者均通过同一个兼容 OpenAI 的端点提供服务,并使用同一把 API 密钥——切换只需更改一行模型字符串,因此你可以将一部分流量路由到各个模型并直接比较账单。

文本转语音如何计费?

按输入文本的字符数计费,每次请求的字符上限如规格表所示。无论使用哪种模型,长文本都必须在多个请求中分块。

相关对比