New Sign up free, 10 calls on us. Up to $1, no card needed.

Veo 3.1

Released 2025-10

videoVision

Veo 3.1 is Google's current flagship video generation model, the latest of the Veo line.

Input
text image
Output
video
Price
$0.4/s

Benchmarks

Veo 3.1: 12 published, but no benchmark it shares with enough other models to compare.

Veo 3.1 other models measured peer average no peer scored higher
SeedVideoBench 2.0 Text-to-Video (1-5) Motion · 2.39-2.73
2.73

Vendor-published: ByteDance

Price in context

Where the price sits among 5 comparable models

Per second$0.4/s
$0.1 · Sora 2 Veo 2 · $0.5

The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.

Specs & limits

Video

Resolution 720p / 1080p A fixed set of options, not a range.
Clip length 4 / 6 / 8 s A fixed set of options, not a range.
Frame rate 24 fps
Aspect ratios 16:9 · 9:16
Video inputs text-to-video · first-frame & first/last-frame image-to-video (up to 2 guide images)
Native audio Yes, generates sound
Format .mp4

Model

Modalities text + image → video
  • Google's Veo 3.1 video model, the current flagship of the Veo line: text-to-video, first-frame image-to-video, and first-and-last-frame interpolation (up to two guide images), all with native synchronized audio (dialogue, ambient sound, music)
  • 720p/1080p, 24 fps, 4/6/8 s, 16:9 and 9:16
  • a generateAudio switch selects sound-on vs silent output. Veo 3.1 sharpens audio-visual sync and reference-image adherence over Veo 3. Synthorai serves it via the async /v1/videos job API. Create a task, then poll it (or hold the request open with Prefer: wait) until the signed video URL is ready. Billing is per output second at separate sound-on and silent rates

per Google official docs ↗

One prompt, measured through the gateway

PROMPT A paper boat drifts across a rain puddle at dusk. The camera pushes in slowly as a single drop lands beside it and rings spread outward across the reflection.

Veo 3.1

Model returned 1280×720 Length 8s latency 63 s

One prompt, one request per model, no retries and no cherry-picking - the first result each model returned. Neither size nor duration was pinned: each model used its own default, because a request shaped to fit all of them would flatter none. Files here are re-encoded for the web, so judge composition and prompt adherence, not compression.

Use Veo 3.1 in 30 seconds

Asynchronous job API: create a task and poll it, or hold the request open with Prefer: wait. POST /v1/videos

import time

import requests

BASE = "https://synthorai.io/v1"
HEADERS = {"Authorization": "Bearer sk-syn-..."}

# 1) create the video generation job (POST /v1/videos)
task = requests.post(
    f"{BASE}/videos",
    headers=HEADERS,
    json={
        "model": "veo-3.1-generate-001",
        "prompt": "a watercolor lighthouse at dawn, waves rolling in, camera pulling back",
        "resolution": "720p",
        "duration": 5,
    },
).json()

# 2) poll the job until it reaches a terminal state
while task["status"] in ("queued", "in_progress"):
    time.sleep(5)
    task = requests.get(f"{BASE}/videos/{task['id']}", headers=HEADERS).json()

# 3) completed → signed video URL (valid ~24h - download and store it promptly)
if task["status"] == "completed":
    print(task["data"][0]["url"])
else:
    print(task["error"])

About Veo 3.1

  • It generates .mp4 clips from a text prompt, from a first-frame image, or as a first-and-last-frame interpolation between two guide images, all with native synchronized audio: dialogue with lip-sync, ambient sound, and music.
  • Output runs at 720p, 1080p, or 4K, 24 fps, 4, 6, or 8 seconds long, in 16:9 or 9:16, and a generateAudio switch selects sound-on or silent output.
  • Beyond Veo 3 it adds up to three asset reference images to carry a character, product, or place across shots, with reference-driven clips fixed at 8 seconds, and video extension that takes an existing clip of 1 to 30 seconds and continues it by 7 seconds per call, so longer sequences are built by chaining rather than by asking for a longer generation.
  • Google also raised the request quota fivefold for this generation.
  • Requests accept up to four videos per call, a seed, a negative prompt, camera-motion selection, and optimized or lossless compression, with Content Credentials embedded in the output.
  • Synthorai serves it through the asynchronous /v1/videos job API and bills per output second, with separate rates for sound-on and silent generations.
  • Create a task, then poll it (or hold the request open with Prefer: wait) until the signed video URL is ready.

FAQ

Is the Veo 3.1 API free to try?

Yes: new accounts get 10 trial calls and up to $1 in free credit, no card required. That's enough to try Veo 3.1 against your real workload before adding a payment method.

What is Veo 3.1 best at?

Current Veo flagship with native audio; first/last-frame image-to-video, up to two guides; sound-on or silent output at separate rates. See the About section for the full picture from the vendor's own release notes.

How much does Veo 3.1 cost?

Veo 3.1 costs $0.4 per second of generated video on Synthorai. Only successful generations are billed: pay-as-you-go, no platform markup, no subscription.

How do I generate videos with the Veo 3.1 API?

POST to /v1/videos on Synthorai with model="veo-3.1-generate-001" to create an asynchronous job, then poll the task (or hold the request open with a Prefer: wait header) until the video URL is ready. No vendor SDK is needed.

How do I get access to Veo 3.1?

Point your existing OpenAI SDK at base_url="https://synthorai.io/v1", set model="veo-3.1-generate-001", and you're done. One API key covers every model on the gateway.

Related models

Compare

Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.

Get your API key Compare your cost →