Qwen3 TTS Instruct Flash vs BytePlus Seed TTS 2.0
什麼情況選哪一個
qwen3-tts-instruct-flash 每百萬字元便宜 2.6 倍($11.5 對 $30),兩者也都接受自然語言的音色指令,但請求長度差別很大:600 個字元對 4096。seed-tts-2.0 還在描述式控制之外提供數值型語音參數,並按場景分組公布音色。短句要控成本選 qwen3-tts-instruct-flash,長段落和更細的音色控制選 seed-tts-2.0。
定價
| Qwen3 TTS Instruct Flash | BytePlus Seed TTS 2.0 | Δ | |
|---|---|---|---|
| 每 1M 字元 | $11.5 | $30 | 0.38× |
費率取自網站建置時的即時目錄;各模型頁面都列有最新的價目。
兩者的相對位置:每 1M 字元的價格,涵蓋同一計費單位下全部 7 個文字轉語音模型(對數尺度)
功能
| Qwen3 TTS Instruct Flash | BytePlus Seed TTS 2.0 | |
|---|---|---|
| 串流 | 是 | 是 |
| SSML | undocumented | unsupported |
| 計費單位 | character | character |
規格
| Qwen3 TTS Instruct Flash | BytePlus Seed TTS 2.0 | |
|---|---|---|
| 輸入模態 | 文字 | 文字 |
| 輸出模態 | 音訊 | 音訊 |
| 發布日期 | - | 2025-10-16 |
| 請求限制 | 600 characters | - |
| 聲音 | System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella the Qwen-TTS voice list pairs each voice with the exact model ids that accept it. | TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12]. |
| 語言 | Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese). | English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean |
| 聲音控制 |
| context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context. |
| 限制 | HTTP non-real-time speech synthesis API the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series) output audio is free. | Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming) SSML is not supported |
規格摘錄自各供應商的文件;供應商沒有公布的項目就直接略過,不自行推測。 完整來源: Qwen3 TTS Instruct Flash · BytePlus Seed TTS 2.0
改一行程式碼就能在兩者之間切換
下面每個頁籤都列了這兩個模型 ID,要改的只有醒目標示的那兩行。端點、金鑰和請求格式都不變。
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="qwen3-tts-instruct-flash",
# model="seed-tts-2.0", # 取消這一行的註解,並把上一行註解掉
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "qwen3-tts-instruct-flash",
// model: "seed-tts-2.0", // 取消這一行的註解,並把上一行註解掉
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="qwen3-tts-instruct-flash" \
# -F model="seed-tts-2.0" \ # 取消這一行的註解,並把上一行註解掉
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "qwen3-tts-instruct-flash",
// Model: "seed-tts-2.0", // 取消這一行的註解,並把上一行註解掉
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("qwen3-tts-instruct-flash")
// .model("seed-tts-2.0") // 取消這一行的註解,並把上一行註解掉
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());常見問題
Qwen3 TTS Instruct Flash 和 BytePlus Seed TTS 2.0 哪個比較便宜?
以「每 1M 字元」來看,Qwen3 TTS Instruct Flash 比較便宜($11.5 對 $30,相差 2.6×)。其他項目的結果可能相反,完整價目請看上表;實際成本要看你的用量組合。
可以只串接一次,就對 Qwen3 TTS Instruct Flash 和 BytePlus Seed TTS 2.0 做 A/B 測試嗎?
可以。兩個模型都走同一個 OpenAI 相容端點,用的也是同一把 API 金鑰,切換時只要改一行裡的模型名稱字串。你可以把一部分流量分別導到兩邊,再直接比較帳單。
文字轉語音是如何計費的?
依輸入文字的字元數計費,每次請求的字元上限列在規格表裡。兩個模型都一樣,稿子太長就得拆成多次請求。