BytePlus Seed TTS 2.0 vs TTS-1 HD
什么时候选哪个
根据所提供的事实,这两者是可互换的:seed-tts-2.0 和 tts-1-hd 均接受文本输入并返回音频,两者都具备语音能力,上下文上限均为 4096 token,且对输入均收取每百万 token $30。由于价格表和单位完全一致,任何一方都没有成本或上下文优势方面的理由——请根据供应商偏好、适合你产品的语音输出,或者你已经集成了哪家提供商来做选择。在两者中分别运行一段你的短脚本样本,并保留听起来合适的那个。
价格
| BytePlus Seed TTS 2.0 | TTS-1 HD | Δ | |
|---|---|---|---|
| 每 1M 字符 | $30 | $30 | = |
价格取自构建时的实时目录,最新价格见各模型页面。
两个模型所处的位置:全部 7 个按同一单位计费的文字转语音模型的每 1M 字符价格分布(对数刻度)
能力
| BytePlus Seed TTS 2.0 | TTS-1 HD | |
|---|---|---|
| 流式输出 | 是 | 是 |
| SSML | unsupported | undocumented |
| 计费单位 | character | character |
规格
| BytePlus Seed TTS 2.0 | TTS-1 HD | |
|---|---|---|
| 输入模态 | 文本 | 文本 |
| 输出模态 | 音频 | 音频 |
| 发布日期 | 2025-10-16 | 2023-11-06 |
| 请求限制 | - | 4096 characters |
| 音色 | TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12]. | alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer, a smaller set than the full 13-voice TTS list, and voices are currently optimized for English. |
| 语言 | English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean | Generally follows the Whisper model’s language support: Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese and Welsh, despite voices being optimized for English |
| 声音控制 | context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context. | none |
| 限制 | Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming) SSML is not supported | Speech generation endpoint (v1/audio/speech) only, with Chat Completions, Responses, Realtime, Batch and Fine-tuning all listed as not supported input text is capped at 4096 characters output mp3 (default), opus, aac, flac, wav or pcm (raw 24 kHz 16-bit signed little-endian samples) realtime playback via chunked transfer encoding |
规格照录自各供应商的文档;供应商没有公布的项目,对应的行直接省略,不做推断。 完整来源: BytePlus Seed TTS 2.0 · TTS-1 HD
改一行代码就能在两个模型之间切换
下方每个标签页里都有两个模型 ID,高亮的那两行是唯一要改的地方。端点不变,API key 不变,请求结构也不变。
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="seed-tts-2.0",
# model="tts-1-hd", # 取消注释此行,注释上一行
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "seed-tts-2.0",
// model: "tts-1-hd", // 取消注释此行,注释上一行
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="seed-tts-2.0" \
# -F model="tts-1-hd" \ # 取消注释此行,注释上一行
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "seed-tts-2.0",
// Model: "tts-1-hd", // 取消注释此行,注释上一行
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("seed-tts-2.0")
// .model("tts-1-hd") // 取消注释此行,注释上一行
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());常见问题
BytePlus Seed TTS 2.0 和 TTS-1 HD 哪个更便宜?
两个模型的「每 1M 字符」价格相同($30),所以这一项上价格分不出高下,请看下方的规格与能力。
不用分别集成两次,就能对 BytePlus Seed TTS 2.0 和 TTS-1 HD 做 A/B 测试吗?
可以。两个模型走同一个 OpenAI 兼容端点,用同一个 API key,切换时只要改一行里的模型名,所以可以给两个模型各分一部分流量,直接对比账单。
文字转语音如何计费?
按输入文本的字符数计费,单次请求的字符上限见规格表。不管用哪个模型,长文本都得拆成多个请求。