Chirp 3 vs Fun-ASR Flash
언제 어떤 모델을 쓸까
둘 다 오디오 분당 과금이고 fun-asr-flash가 7.6배 저렴합니다($0.0021 대 $0.016). 다만 더 좁은 도구입니다: 동기 인식만, 파일당 5분 또는 2GB, 30개 이상 언어, 화자 분리 없음. chirp-3은 GA 29개에 프리뷰 82개 로케일을 다루고 배치는 최대 한 시간, 실시간 스트림도 처리하며 14개 언어에서 분리를 합니다. 짧은 클립을 대량으로 처리하려면 fun-asr-flash를, 분리·긴 파일·넓은 로케일이 필요하면 chirp-3을 고르세요.
가격
| Chirp 3 | Fun-ASR Flash | Δ | |
|---|---|---|---|
| 오디오 분당 | $0.016 | $0.0021 | 7.6× |
빌드 시점에 실시간 카탈로그에서 가져온 요금입니다. 현재 요금은 각 모델 페이지에서 확인할 수 있습니다.
가격 위치: 이 과금 단위를 쓰는 음성 인식 모델 12개 전체의 오디오 분당 가격(로그 스케일)
기능
| Chirp 3 | Fun-ASR Flash | |
|---|---|---|
| 화자 분리 | 지원 | 미지원 |
| 스트리밍 | 지원 | 지원 |
| 타임스탬프 | 지원 | - |
사양
| Chirp 3 | Fun-ASR Flash | |
|---|---|---|
| 입력 모달리티 | 오디오 | 오디오 |
| 출력 모달리티 | 텍스트 | 텍스트 |
| 출시일 | 2025-10-13 | 2026-06 |
| 제한 | Auto-detected audio decoding sync Recognize <1 min, BatchRecognize 1 min - 1 hr (<=20 min with word timestamps), StreamingRecognize for real-time speaker diarization in BatchRecognize and Recognize (14 languages) utterance-level timestamps (StreamingRecognize only), word-level timestamps listed as unsupported language-agnostic transcription 29 GA + 82 preview locales | Synchronous fast recognition, <=5 min / <=2GB per audio context injection for domain terms 30+ languages no diarization |
| 언어 | 29 GA + 82 Preview locales (111 total) across StreamingRecognize, Recognize and BatchRecognize diarization covers 14 of them | Multilingual with dialects, the same 30-language list as the Fun-ASR main versions: Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin plus regional accents), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak |
사양은 각 공급자의 문서에서 그대로 옮겨 적었습니다. 공급자가 공개하지 않은 항목은 추정해 채우지 않고 뺐습니다. 전체 출처: Chirp 3 · Fun-ASR Flash
코드 한 줄로 모델 전환
아래 탭마다 두 모델 id가 모두 들어 있습니다. 바꿀 곳은 강조된 두 줄뿐이고, 엔드포인트, 키, 요청 형식은 그대로입니다.
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="chirp-3",
# model="fun-asr-flash", # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "chirp-3",
// model: "fun-asr-flash", // 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="chirp-3" \
# -F model="fun-asr-flash" \ # 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "chirp-3",
// Model: "fun-asr-flash", // 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("chirp-3")
// .model("fun-asr-flash") // 이 줄의 주석을 해제하고, 윗줄을 주석 처리하세요
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());자주 묻는 질문
Chirp 3 vs Fun-ASR Flash, 어느 쪽이 더 저렴한가요?
오디오 분당 기준으로는 Fun-ASR Flash 쪽이 더 저렴합니다($0.0021 대 $0.016, 7.6배 차이). 다른 항목에서는 반대일 수 있습니다. 전체 요금은 위 표에 있으며, 실제 비용은 사용 패턴에 따라 달라집니다.
연동을 두 번 하지 않고도 Chirp 3 vs Fun-ASR Flash A/B 테스트를 할 수 있나요?
네. 두 모델 모두 API 키 하나로 같은 OpenAI 호환 엔드포인트에서 호출합니다. 모델 문자열 한 줄만 바꾸면 전환되므로, 트래픽 일부를 각 모델로 보내 청구 금액을 직접 비교할 수 있습니다.
Chirp 3, Fun-ASR Flash 모두 화자 분리를 지원하나요?
위 기능 표에 모델별로 정리되어 있으며, 내용은 각 공급자 문서에서 그대로 가져왔습니다. 화자 분리, 스트리밍, 타임스탬프는 모델마다 지원 여부가 달라 항목을 따로 나눴습니다.