Back to guides
Video·June 18, 2026·Updated October 1, 2026·7 min read

Sora API — generate AI video at an OpenAI-like endpoint (currently with Veo 3.1)

AI video through the Sora API: Sora is not activated yet; the active text-to-video model is Google Veo 3.1 at the same OpenAI-like endpoint. Includes runnable examples and prices from $0.18 per clip.

Sora is OpenAI's text-to-video model, and “the Sora API” is the way teams generate video programmatically rather than through the consumer app. At Kunavo, text-to-video currently runs through Google Veo 3 at an OpenAI-style video endpoint—Sora access is on the roadmap, and because the endpoint is model-independent, switching to Sora later will be a one-word change. This guide shows the workflow with Veo 3 so every example runs immediately.

What is the Sora API?

Sora (Sora 2 and Sora 2 Pro) turns a text prompt—or a still image—into a short video clip, with synchronized audio for Sora 2. The API format lets you script generation into a pipeline: marketing shots, product animations, B-roll, and storyboard previews. The format is consistent across modern video models: send a prompt and parameters, and receive a hosted video URL.

Video at Kunavo today: Veo 3

Kunavo provides video generation through a single OpenAI-style endpoint, /v1/video/generations. Sora has not yet been activated in the catalog; the active text-to-video model is Google Veo 3, which generates cinematic clips with native audio. The examples below use veo-3—when Sora arrives, the only change will be the model field.

Text-to-video quickstart

A POST request with your Kunavo key. Generation can take a few minutes, so the synchronous call keeps the connection open until the clip is ready:

text_to_video.py
import requests

resp = requests.post(
    "https://api.kunavo.com/v1/video/generations",
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={
        "model": "veo-3",
        "prompt": "a cinematic dolly-in on a red origami crane unfolding, soft light",
        "aspect_ratio": "16:9",
    },
    timeout=600,  # die Generierung kann Minuten dauern
)
print(resp.json()["data"][0]["url"])

duration (seconds) sets the clip length on the models billed per second: the Seedance models and Wan 2.7. The Veo models ignore it; every Veo clip is 8 seconds, billed per video.

Image-to-video

To animate a still image, pass image_url (an https URL or a file you uploaded to /v1/files) together with the prompt. For controlled motion, you can pass a first and last frame using image_urls and image_mode: "frame". Full examples are in the Video Docs.

Asynchronous task lifecycle

In production, you should not keep a 10-minute connection open. Submit a task to /v1/videos, receive a task ID immediately, and then poll GET /v1/videos/{id} until it is complete. Result URLs are permanent.

async_submit.py
# Produktion: Task einreichen, dann pollen — keine langlebige Verbindung.
task = requests.post(
    "https://api.kunavo.com/v1/videos",
    headers={
        "Authorization": f"Bearer {API_KEY}",
        "Idempotency-Key": "my-task-uuid",   # innerhalb von ~24h retry-sicher
    },
    json={"model": "veo-3", "prompt": "...", "aspect_ratio": "16:9"},
    timeout=60,
).json()
# dann GET /v1/videos/{task["id"]} pollen, bis er fertig ist

The Video Docs cover the complete polling loop, idempotency keys, and webhook delivery.

Prices

Veo 3 is billed per video (here, per 8-second clip in 720p), approximately 50–70% below Google's list price. Higher resolutions cost more—the complete tier table is on the pricing page.

ModelFrom (720p / 8s)Google list priceYour savings
veo-3-lite$0.18$0.45~60%
veo-3 (Fast)$0.36$1.20~70%
veo-3-quality$1.60$3.20~50%

Video prompting tips

  • Describe the setting, not just the subject. Camera movement (dolly, pan, push-in), lens characteristics, lighting, and pacing matter more than adjectives.
  • Set the aspect ratio explicitly—16:9 for landscape, 9:16 for portrait/social.
  • Every Veo clip is 8 seconds long. Plan each shot to fit, and stitch multiple generations for longer sequences.
  • Use a reference frame (image-to-video) when a particular character or product needs to remain consistent.

Frequently asked questions

Can I use the Sora API at Kunavo today?

Sora has not yet been activated in the Kunavo catalog. Text-to-video currently runs through Google Veo 3 at the same OpenAI-style /v1/video/generations endpoint. Because the endpoint is model-independent, switching to Sora later will be a one-word change to the model field. Sora access is on the roadmap.

What is the Sora API?

Sora is OpenAI's text-to-video model (Sora 2 and Sora 2 Pro). The Sora API generates short video clips—from a text prompt or a still image, with synchronized audio for Sora 2—programmatically rather than through the consumer app.

How much does video generation cost at Kunavo?

Veo 3 is billed per video, approximately 50–70% below Google's list price: Veo 3 Lite from $0.18, Veo 3 Fast from $0.36, and Veo 3 Quality from $1.60 per 8-second clip in 720p. Higher resolutions cost more.

Does Kunavo support image-to-video?

Yes. Pass image_url (an https URL or an uploaded file) to /v1/video/generations to animate a still image. You can also pass a first and last frame for controlled transitions.

How long does generation take?

Minutes. In production, use the asynchronous /v1/videos task API and poll it; see the Video Docs. Per-model rates are also listed in the Gemini API pricing guide.

Do I need a separate key for this?

No—the same Kunavo key covers video, text, and image. Instructions for creating one are in Create a Google Gemini API Key; the key created there works unchanged with /v1/video/generations.