Google TTS Chirp 3 HD vs Google TTS Neural2
什麼情況選哪一個
google-tts-neural2 便宜 1.9 倍(每百萬字元 $16 對 $30)並且直接支援 SSML,而 google-tts-chirp3-hd 把 SSML 列為預覽——但 neural2 不支援串流,公布的語言區列表也短得多,只有 17 個帶命名語音 id 的語言區。兩者每次請求都接受 5000 位元組,也都提供數值型語音參數。在已涵蓋的語言區裡做 SSML 驅動的批次合成選 google-tts-neural2,要串流和語言區廣度選 google-tts-chirp3-hd。
定價
| Google TTS Chirp 3 HD | Google TTS Neural2 | Δ | |
|---|---|---|---|
| 每 1M 字元 | $30 | $16 | 1.9× |
費率取自網站建置時的即時目錄;各模型頁面都列有最新的價目。
兩者的相對位置:每 1M 字元的價格,涵蓋同一計費單位下全部 7 個文字轉語音模型(對數尺度)
功能
| Google TTS Chirp 3 HD | Google TTS Neural2 | |
|---|---|---|
| 串流 | 是 | 否 |
| SSML | preview | supported |
| 計費單位 | character | character |
規格
| Google TTS Chirp 3 HD | Google TTS Neural2 | |
|---|---|---|
| 輸入模態 | 文字 | 文字 |
| 輸出模態 | 音訊 | 音訊 |
| 發布日期 | 2025-03 | 2022-06-27 |
| 請求限制 | 5000 bytes | 5000 bytes |
| 聲音 | Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages. | Voice ids follow the locale-plus-letter pattern (en-US-Neural2-F, ja-JP-Neural2-B). Google's comparison table lists Neural2 as general purpose, generally available, controllable via SSML and not streaming-capable, and the docs state the voices are based on the same technology used to create a Custom Voice, letting anyone use Custom Voice technology without training their own. |
| 語言 | ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. | A much shorter locale list than Standard. Neural2 voice ids are published for da-DK, de-DE, en-AU, en-GB, en-IN, en-US, es-ES, es-US, fr-CA, fr-FR, hi-IN, it-IT, ja-JP, ko-KR, pt-BR, th-TH and vi-VN. Google notes Neural2 voices are available on global and single-region endpoints. |
| 聲音控制 |
|
|
| 限制 | Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16 streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted. | Content limit of 5,000 total bytes per synthesize request. Output LINEAR16 (with a WAV header), MP3 at 32 kbps, OGG_OPUS, or G.711 MULAW and ALAW optional sampleRateHertz resamples and fails the request if the rate is unsupported for the encoding. Not offered over streaming synthesis. Billed per character including spaces and newlines, with all SSML tags except <mark> counted the multi-byte-counts-once note applies to Standard and WaveNet only. |
規格摘錄自各供應商的文件;供應商沒有公布的項目就直接略過,不自行推測。 完整來源: Google TTS Chirp 3 HD · Google TTS Neural2
改一行程式碼就能在兩者之間切換
下面每個頁籤都列了這兩個模型 ID,要改的只有醒目標示的那兩行。端點、金鑰和請求格式都不變。
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="google-tts-chirp3-hd",
# model="google-tts-neural2", # 取消這一行的註解,並把上一行註解掉
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "google-tts-chirp3-hd",
// model: "google-tts-neural2", // 取消這一行的註解,並把上一行註解掉
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="google-tts-chirp3-hd" \
# -F model="google-tts-neural2" \ # 取消這一行的註解,並把上一行註解掉
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "google-tts-chirp3-hd",
// Model: "google-tts-neural2", // 取消這一行的註解,並把上一行註解掉
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("google-tts-chirp3-hd")
// .model("google-tts-neural2") // 取消這一行的註解,並把上一行註解掉
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());常見問題
Google TTS Chirp 3 HD 和 Google TTS Neural2 哪個比較便宜?
以「每 1M 字元」來看,Google TTS Neural2 比較便宜($16 對 $30,相差 1.9×)。其他項目的結果可能相反,完整價目請看上表;實際成本要看你的用量組合。
可以只串接一次,就對 Google TTS Chirp 3 HD 和 Google TTS Neural2 做 A/B 測試嗎?
可以。兩個模型都走同一個 OpenAI 相容端點,用的也是同一把 API 金鑰,切換時只要改一行裡的模型名稱字串。你可以把一部分流量分別導到兩邊,再直接比較帳單。
文字轉語音是如何計費的?
依輸入文字的字元數計費,每次請求的字元上限列在規格表裡。兩個模型都一樣,稿子太長就得拆成多次請求。