新帳號 免費註冊,送 10 次呼叫,最高 $1,免綁卡。

Google TTS Chirp 3 HD vs Qwen3 TTS Instruct Flash

vs

什麼情況選哪一個

qwen3-tts-instruct-flash 每百萬字元便宜 2.6 倍($11.5 對 $30),並且能用自然語言指揮音色,但它每次請求只接受 600 個字元,而 google-tts-chirp3-hd 是 5000 位元組,所以長文案必須切分。Google 這一款涵蓋更寬的語言區列表,提供數值型語音參數,SSML 處於預覽。要為短句描述交付方式選 qwen3-tts-instruct-flash,長文字和語言區廣度選 google-tts-chirp3-hd。

定價

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash Δ
每 1M 字元 $30 $11.5 2.6×

費率取自網站建置時的即時目錄;各模型頁面都列有最新的價目。

兩者的相對位置:每 1M 字元的價格,涵蓋同一計費單位下全部 7 個文字轉語音模型(對數尺度)

功能

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash
串流 是 是
SSML preview undocumented
計費單位 character character

規格

Google TTS Chirp 3 HD Qwen3 TTS Instruct Flash
輸入模態 文字 文字
輸出模態 音訊 音訊
發布日期 2025-03 -
請求限制 5000 bytes 600 characters
聲音 Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages.

System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella

the Qwen-TTS voice list pairs each voice with the exact model ids that accept it.

語言 ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN.

Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy

naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese).

聲音控制
  • Voice controls are Preview: pace via speaking_rate 0.25 to 2.0
  • pause tags [pause short], [pause long] and [pause] accepted only in the markup input field, never in text, and the model may disregard tags placed unnaturally
  • custom pronunciations in IPA or X-SAMPA
  • SSML support is Preview and synchronous-only, unsupported for streaming requests, with unlisted tags ignored and say-as interpret-as=expletive or bleep not supported
  • instructions steers speed, emotion and style in natural language, up to 1,600 tokens and documented for Chinese and English only
  • optimize_instructions (default false) rewrites the instruction into a directive better suited to synthesis and has no effect when instructions is empty
  • voice cloning and voice design are both listed as unsupported for this model id
限制

Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16

streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted.

HTTP non-real-time speech synthesis API

the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours

streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series)

output audio is free.

規格摘錄自各供應商的文件;供應商沒有公布的項目就直接略過,不自行推測。 完整來源: Google TTS Chirp 3 HD · Qwen3 TTS Instruct Flash

改一行程式碼就能在兩者之間切換

下面每個頁籤都列了這兩個模型 ID,要改的只有醒目標示的那兩行。端點、金鑰和請求格式都不變。

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.audio.transcriptions.create(
    model="google-tts-chirp3-hd",
    # model="qwen3-tts-instruct-flash",  # 取消這一行的註解,並把上一行註解掉
    file=open("meeting.mp3", "rb"),
    language="en",
)
print(resp.text)

取得 API 金鑰 →

常見問題

Google TTS Chirp 3 HD 和 Qwen3 TTS Instruct Flash 哪個比較便宜?

以「每 1M 字元」來看,Qwen3 TTS Instruct Flash 比較便宜($11.5 對 $30,相差 2.6×)。其他項目的結果可能相反,完整價目請看上表;實際成本要看你的用量組合。

可以只串接一次,就對 Google TTS Chirp 3 HD 和 Qwen3 TTS Instruct Flash 做 A/B 測試嗎?

可以。兩個模型都走同一個 OpenAI 相容端點,用的也是同一把 API 金鑰,切換時只要改一行裡的模型名稱字串。你可以把一部分流量分別導到兩邊,再直接比較帳單。

文字轉語音是如何計費的?

依輸入文字的字元數計費,每次請求的字元上限列在規格表裡。兩個模型都一樣,稿子太長就得拆成多次請求。

相關比較