Fun-ASR is Alibaba Cloud Model Studio's recording-file speech recognition model for offline transcription of pre-recorded audio.
- Input
- audio
- Output
- text
- Price
- $0.0021/min
Price in context
Where the price sits among 11 comparable models
The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.
Specs & limits
Audio
| Languages | Multilingual with dialects, 30 languages: Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin plus regional accents), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak |
|---|---|
| Audio limits |
|
| Speaker diarization | Yes |
| Streaming transcription | No |
| Timestamps | Yes |
Model
| Modalities | audio → text |
|---|
- Singing recognition added 2025-12
- open-sourced Fun-ASR-Nano is a distinct model from these served weights
Use Fun-ASR in 30 seconds
OpenAI-compatible: swap the base_url, keep your SDK. POST /v1/audio/transcriptions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.audio.transcriptions.create(
model="fun-asr",
file=open("meeting.mp3", "rb"),
language="en",
)
print(resp.text)import OpenAI from "openai";
import fs from "node:fs";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.audio.transcriptions.create({
model: "fun-asr",
file: fs.createReadStream("meeting.mp3"),
});
console.log(resp.text);curl https://synthorai.io/v1/audio/transcriptions \
-H "Authorization: Bearer sk-syn-..." \
-F model="fun-asr" \
-F file=@meeting.mp3package main
import (
"context"
"fmt"
"os"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
f, _ := os.Open("meeting.mp3")
resp, _ := client.Audio.Transcriptions.New(context.TODO(), openai.AudioTranscriptionNewParams{
Model: "fun-asr",
File: f,
})
fmt.Println(resp.Text)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.audio.transcriptions.*;
import java.nio.file.Paths;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
Transcription resp = client.audio().transcriptions().create(
TranscriptionCreateParams.builder()
.model("fun-asr")
.file(Paths.get("meeting.mp3"))
.build()).asTranscription();
System.out.println(resp.text());About Fun-ASR
- It accepts files up to 12 hours and 2 GB, handles single-file and batch jobs, and recognizes Chinese (Mandarin plus many regional dialects such as Cantonese, Wu, and Hokkien), English, Japanese, Korean, and dozens of other languages.
- The official docs highlight hotword customization for domain terminology and speaker diarization for multi-speaker recordings.
- Both come with practical detail worth knowing up front: a hotword table holds up to 500 entries and suits terminology that changes rarely, diarization is off until you enable it and lets you declare an expected speaker count, and Alibaba recommends keeping diarized audio under about two hours or the job may time out.
- Transcription is asynchronous (submit the file, receive a task id, poll for the result), and sentence- and word-level timestamps in milliseconds come back as standard, alongside sensitive-word filtering and per-channel selection for multi-track recordings.
- Emotion labelling is not part of this model; that belongs to the Qwen3-ASR line.
- Language hints are accepted but the docs note only the first value in the list is read.
- Synthorai serves Fun-ASR through its OpenAI-compatible endpoint, so existing transcription clients work without modification.
FAQ
Is the Fun-ASR API free to try?
Yes: new accounts get 10 trial calls and up to $1 in free credit, no card required. That's enough to try Fun-ASR against your real workload before adding a payment method.
What is Fun-ASR best at?
Accepts files up to 12 hours; hotword customization for domain terminology; speaker diarization for multi-speaker recordings. See the About section for the full picture from the vendor's own release notes.
How much does Fun-ASR cost?
Fun-ASR costs $0.0021 per minute of audio transcribed on Synthorai: pay-as-you-go, no platform markup, no subscription.
Which languages does Fun-ASR support?
Fun-ASR supports Multilingual with dialects, 30 languages: Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin plus regional accents), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak. On Synthorai you call it through POST /v1/audio/transcriptions, the OpenAI transcription API shape.
How do I get access to Fun-ASR?
Point your existing OpenAI SDK at base_url="https://synthorai.io/v1", set model="fun-asr", and you're done. One API key covers every model on the gateway.
Related models
Compare
Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.