新規 無料登録で、呼び出し 10 回分(最大 $1)をお試しいただけます。カード登録は不要です。

Qwen3 TTS Instruct Flash vs BytePlus Seed TTS 2.0

vs

どんなときに、どちらを選ぶか

qwen3-tts-instruct-flash は 100万文字あたり 2.6 分の 1($11.5 対 $30)で、どちらも自然な言葉での声の指示を受け付けますが、リクエストの上限が大きく違います — 600 文字 対 4096。seed-tts-2.0 は説明による制御に加えて数値パラメータも備え、音声をシーン別に分類して公開しています。短い台詞のコストなら qwen3-tts-instruct-flash、長い文章とより細かい声の制御なら seed-tts-2.0 を選んでください。

料金

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0 Δ
100 万文字あたり $11.5 $30 0.38×

ビルド時点の現行カタログの料金です。最新の料金表は各モデルのページに掲載しています。

両モデルの位置:100 万文字あたりの料金(この課金単位の音声合成モデル全 7 件、対数スケール)

対応機能

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0
ストリーミング あり あり
SSML undocumented unsupported
課金単位 character character

仕様

Qwen3 TTS Instruct Flash BytePlus Seed TTS 2.0
入力モダリティ テキスト テキスト
出力モダリティ 音声 音声
リリース - 2025-10-16
リクエストあたりの上限 600 characters -
ボイス

System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella

the Qwen-TTS voice list pairs each voice with the exact model ids that accept it.

TTS 2.0 voices carry *_uranus_bigtts speaker IDs and are listed by scenario (General, Entertainment, Education, Dubbing, AudioBook, RolePlay, CustomerService) on the Voice List page

per-voice emotion via audio_params.emotion with emotion_scale 1 to 5 (default 4), speech_rate and loudness_rate in [-50, 100], and post_process.pitch in [-12, 12].

言語

Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy

naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese).

English, Chinese, Japanese, German, French, Mexican Spanish, Indonesian Bahasa, Brazilian Portuguese, Italian and Korean
ボイス調整
  • instructions steers speed, emotion and style in natural language, up to 1,600 tokens and documented for Chinese and English only
  • optimize_instructions (default false) rewrites the instruction into a directive better suited to synthesis and has no effect when instructions is empty
  • voice cloning and voice design are both listed as unsupported for this model id
context_texts and section_id are TTS 2.0 only: a plain-language instruction steers rate, emotion, volume and style (only the first list value takes effect, and its text is not billed), and section_id links up to 30 rounds or 10 minutes of earlier synthesis as historical context.
制限

HTTP non-real-time speech synthesis API

the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours

streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series)

output audio is free.

Integration via uni-directional streaming HTTP, uni- or bi-directional streaming WebSocket and online SDK

sampling rate 24K/16K/8K on the uni-directional streaming and non-streaming interfaces, 48K/24K/16K/8K on the bi-directional interface

output PCM, OGG_OPUS or MP3 (WAV is accepted but returns multiple headers when streaming)

SSML is not supported

仕様は各ベンダーのドキュメントから転記しています。ベンダーが公表していない項目は、推測で埋めずに省いています。 出典の一覧: Qwen3 TTS Instruct Flash · BytePlus Seed TTS 2.0

1 行の変更で切り替え

以下のどのタブにも両方のモデル ID が入っています。書き換えるのはハイライトされた 2 行だけで、エンドポイント、キー、リクエスト形式はすべて同じです。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="qwen3-tts-instruct-flash",
    # model="seed-tts-2.0",  # この行のコメントを外し、上の行をコメントアウト
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

API キーを取得 →

よくある質問

Qwen3 TTS Instruct Flash と BytePlus Seed TTS 2.0 ではどちらが安いですか?

100 万文字あたりは Qwen3 TTS Instruct Flash のほうが安くなります($11.5 対 $30、2.6 倍の差)。ほかの行では逆になることもあります。料金の全体は上の表に掲載しており、実際のコストは利用の内訳によって変わります。

Qwen3 TTS Instruct Flash と BytePlus Seed TTS 2.0 の A/B テストは、実装を 2 つ用意せずにできますか?

はい。どちらも同じ OpenAI 互換エンドポイントから 1 つの API キーで利用できます。切り替えはモデル名の文字列を 1 行変えるだけなので、トラフィックの一部をそれぞれに振り分けて、請求額を直接比較できます。

音声合成(Text-to-Speech)の課金方法は?

入力テキストの文字数に応じて課金されます。1 リクエストあたりの文字数上限は仕様表に掲載しています。どちらのモデルでも、長い原稿は複数のリクエストに分けて送る必要があります。

関連する比較

当社の実測レポートより