Google TTS Chirp 3 HD vs Qwen3 TTS Instruct Flash
Qual usar e quando
O qwen3-tts-instruct-flash é 2.6x mais barato por milhão de caracteres ($11.5 contra $30) e aceita direção de voz em linguagem simples, mas só admite 600 caracteres por requisição contra 5000 bytes no google-tts-chirp3-hd, então textos longos precisam ser divididos. O modelo do Google cobre a lista de locais mais ampla e expõe parâmetros numéricos de voz, com SSML em prévia. Escolha o qwen3-tts-instruct-flash para falas curtas com uma intenção descrita, o google-tts-chirp3-hd para texto longo e amplitude de locais.
Preços
| Google TTS Chirp 3 HD | Qwen3 TTS Instruct Flash | Δ | |
|---|---|---|---|
| Por 1M caracteres | $30 | $11.5 | 2.6× |
Tarifas do catálogo em tempo real, lidas no momento do build; a página de cada modelo traz a tabela de preços atual.
Onde cada um fica: preço por 1M caracteres entre os 7 modelos de síntese de voz com esta unidade de cobrança (escala logarítmica)
Capacidades
| Google TTS Chirp 3 HD | Qwen3 TTS Instruct Flash | |
|---|---|---|
| Streaming | sim | sim |
| SSML | preview | undocumented |
| Unidade de cobrança | character | character |
Especificações
| Google TTS Chirp 3 HD | Qwen3 TTS Instruct Flash | |
|---|---|---|
| Modalidades de entrada | texto | texto |
| Modalidades de saída | áudio | áudio |
| Lançamento | 2025-03 | - |
| Limite por requisição | 5000 bytes | 600 characters |
| Vozes | Voices are named after stars and published with a gender: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr and Zubenelgenubi. A locale prefix forms the id, e.g. en-US-Chirp3-HD-Charon. Google describes the set as 30 distinct styles across many languages. | System voices published with one-line personas, among them Cherry (a sunny, positive, friendly and natural young woman), Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish (a designer who cannot pronounce retroflex sounds), Bella, Eldric Sage, Mia, Mochi, Bellona, Vincent, Bunny, Neil, Elias, Arthur, Nini, Seren, Pip and Stella the Qwen-TTS voice list pairs each voice with the exact model ids that accept it. |
| Idiomas | ar-XA, bn-IN, bg-BG, yue-HK, hr-HR, cs-CZ, da-DK, nl-BE, nl-NL, en-AU, en-IN, en-GB, en-US, et-EE, fi-FI, fr-CA, fr-FR, de-DE, el-GR, gu-IN, he-IL, hi-IN, hu-HU, id-ID, it-IT, ja-JP, kn-IN, ko-KR, lv-LV, lt-LT, ml-IN, cmn-CN, mr-IN, nb-NO, pl-PL, pt-BR, pa-IN, ro-RO, ru-RU, sr-RS, sk-SK, sl-SI, es-ES, es-US, sw-KE, sv-SE, ta-IN, te-IN, th-TH, tr-TR, uk-UA, ur-IN and vi-VN. | Chinese (Mandarin), English, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. language_type defaults to Auto for mixed-language or undetermined input, which Alibaba documents as not guaranteeing accuracy naming a single language is documented to significantly improve synthesis quality. Unlike the Qwen3-TTS-Flash series, the Instruct series lists no Chinese dialect voices (Beijing, Shanghainese, Sichuan, Nanjing, Shaanxi, Hokkien, Tianjin, Cantonese). |
| Controle de voz |
|
|
| Limites | Content limit of 5,000 total bytes per synthesize request. Default response format LINEAR16 streaming supports ALAW, MULAW, OGG_OPUS and PCM, batch adds MP3, so MP3 is not available on the streaming path. The only voice type Google's comparison table marks as streaming-capable. Generally available in the global, us and eu endpoints plus asia-southeast1, europe-west2 and asia-northeast1, but out of scope for regionalization and data residency. Billed per character including spaces and newlines, with all SSML tags except <mark> counted. | HTTP non-real-time speech synthesis API the -realtime suffix marks the WebSocket sibling. Input text capped at 600 characters, multilingual mixed input allowed. Non-streaming returns an audio file URL valid for 24 hours streaming returns Base64-encoded PCM in chunks with the URL only in the final packet, played back in the official samples as 24 kHz mono 16-bit audio. Billed by input text characters, reported as usage.characters (input_tokens and output_tokens are always 0 on the Qwen3-TTS series) output audio is free. |
As especificações são transcritas da documentação de cada provedor; quando um provedor não publica um dado, a linha é omitida, e não inferida. Fontes completas: Google TTS Chirp 3 HD · Qwen3 TTS Instruct Flash
Troque de um para o outro mudando uma linha
Os dois ids estão em todas as abas abaixo; o par de linhas destacado é a única alteração. O endpoint, a chave e o formato da requisição são os mesmos.
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="google-tts-chirp3-hd",
# model="qwen3-tts-instruct-flash", # descomente esta linha, comente a linha acima
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "google-tts-chirp3-hd",
// model: "qwen3-tts-instruct-flash", // descomente esta linha, comente a linha acima
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="google-tts-chirp3-hd" \
# -F model="qwen3-tts-instruct-flash" \ # descomente esta linha, comente a linha acima
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "google-tts-chirp3-hd",
// Model: "qwen3-tts-instruct-flash", // descomente esta linha, comente a linha acima
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("google-tts-chirp3-hd")
// .model("qwen3-tts-instruct-flash") // descomente esta linha, comente a linha acima
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());Perguntas frequentes
Qual é mais barato, Google TTS Chirp 3 HD ou Qwen3 TTS Instruct Flash?
Qwen3 TTS Instruct Flash sai mais barato na linha “Por 1M caracteres” ($11.5 contra $30, 2.6× de diferença). Outras linhas podem pender para o outro lado: a tabela acima mostra todos os preços, e o custo real depende do seu mix de uso.
Posso fazer um teste A/B de Google TTS Chirp 3 HD contra Qwen3 TTS Instruct Flash sem duas integrações?
Sim. Os dois são servidos pelo mesmo endpoint compatível com OpenAI, com uma única chave de API. Para trocar, basta mudar uma linha, a string do modelo; assim você pode direcionar uma parte do tráfego para cada um e comparar as contas diretamente.
Como o text-to-speech é cobrado?
Por caractere do texto de entrada, com um limite de caracteres por requisição indicado na tabela de especificações. Nos dois modelos, textos longos precisam ser divididos em várias requisições.