Wan 3.0 Is on Genfire: Alibaba's 30-Second AI Video Model With Native Audio (and a Prime Tier)
Wan 3.0 renders 2 to 30 seconds at 1080p with synchronized audio, takes 20 references, and has a faster Prime tier. What it does and how to run it.
What Wan 3.0 Is
Wan 3.0 is Alibaba's current video generation model, and it is live on Genfire in two tiers. It renders 2 to 30 seconds of video at up to 1080p, with synchronized audio produced in the same pass as the picture and switched on by default. It works from a prompt, from a start frame with an optional end frame, or from a pool of reference images, clips, and audio takes.
Two numbers there are unusual. The two-second floor is lower than almost anything else in the catalog — most models make you pay for five seconds whether the shot needs them or not. And 1080p is the default, not the upgrade: you drop to 720p or 480p for a cheaper draft rather than climbing for a finished one.
Checked against the shipped registry entries, the full spec:
- Duration: any whole second from 2 to 30
- Resolution: 480p, 720p, or 1080p — 1080p is the default, and there is no 4K
- Aspect ratios: adaptive, 16:9, 4:3, 1:1, 3:4, 9:16
- Modes: text-to-video, image-to-video with a start and optional end frame, and reference-to-video
- References: up to 10 images, 5 clips, and 5 audio takes
- Audio: native, synchronized, included at no extra cost, and switchable off
- Prompt length: up to 5,000 characters
- Tiers: Wan 3.0 and Wan 3.0 Prime — same inputs, Prime faster and higher fidelity at 1.4x the price
The spec sheet and prompt ideas live at /wan-3. This post is about how the pieces fit together and when to reach for each one.
The Two Tiers
There are no speed tiers on Wan 3.0 in the usual sense — no Fast, no Mini, no Lite. There are two entries in the picker, Wan 3.0 and Wan 3.0 Prime, and they accept identical inputs: the same durations, resolutions, aspect ratios, reference pools, audio toggle, and 5,000-character prompt. The difference is the render — Prime is the faster, higher-fidelity pass, at 1.4x the price per second.
That makes the choice simpler than a Fast tier would. Both are the same model:
- Standard Wan 3.0 for volume — social cuts, variants, anything where you generate six and keep two. It already outputs 1080p with sound.
- Wan 3.0 Prime when the shot is the deliverable: the hero clip in an ad, a close-up where skin and fabric have to hold up, a 30-second single take you do not want to re-roll.
A good pattern: find the shot on standard Wan 3.0 at 480p, then re-run the winning prompt and references on Prime at 1080p. Same inputs, so it is a model swap, not a rebuild.
Three Ways In
Wan 3.0 is one model with three endpoints behind it, and the mode follows the inputs you supply.
Text-to-video is prompt only — set a duration, a resolution, and an aspect ratio, or set the ratio to adaptive and let the output take its shape from the content rather than a fixed frame.
Image-to-video animates from a start frame, and if you also give it an end frame the clip travels from one still to the other. That second frame turns a generation into a transition: a product on a shelf becomes a product in a hand. With 2 seconds available, it is also how you get short connective inserts that would otherwise be five seconds of padding.
Reference-to-video has the most control, and gets its own section below.
There is a fourth input path no other video model on Genfire has: `file_url` and `web_url`. Hand the reference endpoint a document or a public webpage and Wan 3.0 reads it and builds the shot from its contents. Supplying either one switches on the model's deep-reasoning pass automatically — otherwise the source would be accepted and quietly ignored — and the Genfire API rejects both fields on every other model rather than dropping an input you have paid for. Available through the API, the MCP server, and the Wan 3.0 node in the workflow editor; pages behind a login cannot be read.
The Reference Pool
A reference-to-video request takes up to 10 images, 5 clips, and 5 audio takes — 20 files. You address each one in the prompt by name: @Image1, @Video1, @Audio1. Genfire rewrites those into the bare positional form Alibaba's endpoint actually reads ("the subject in Image 1 walks past Video 1"), so the wire format never becomes your problem.
Each modality does a different job:
- Images hold identity — a face from several angles, a product, a location, a logo.
- Clips carry motion: a camera move to follow, a piece of choreography, a rhythm.
- Audio gives a character a voice or a scene its music. Audio cannot be the only reference; pair it with at least one image or clip and say in the prompt who speaks with which take.
Two limits matter before you attach anything. The video pool carries a 15-second total budget across all five clips, and the audio pool carries its own 15 seconds — these are conditioning references, not source footage. Clips should run at 16 fps or better. You do not have to pre-trim an over-long clip: Genfire takes a time window per clip (reference_video_trims on the API, --ref-video-trim on the CLI, a trim control on the slot in the studio) and cuts exactly that window, sound included, before submitting.
What Changed From WAN 2.5
WAN 2.5 has not gone anywhere, and Wan 3.0 does not replace it — they do different jobs.
WAN 2.5 on Genfire is the character-animation tier. Animate Move drives a still image with the motion of a reference video; Animate Replace swaps the subject in existing footage for your reference character. That is puppeteering: the reference video supplies the movement, your image supplies who is doing it. For mascots, virtual influencers, and dance content, that is direction rather than dice — and Wan 3.0 does not do it.
Wan 3.0 drops those modes and becomes a general-purpose generator with much longer reach: 2 to 30 seconds in one pass at up to 1080p with sound, a start-and-end-frame image mode, and a reference pool WAN 2.5 never had.
So: puppeteer a still from a reference clip → WAN 2.5. Generate a finished shot with ambience and a consistent cast → Wan 3.0. Both sit in the same WAN group and bill from the same credit balance.
Wan 3.0 vs Seedance 2.5 vs Flux 3
Three audio-native models on Genfire overlap in the same territory. These figures come from the model registry.
| Wan 3.0 | Seedance 2.5 | Flux 3 | |
|---|---|---|---|
| Duration | 2–30s | 4–30s | 5–20s |
| Resolution | 480p · 720p · 1080p (default 1080p) | 480p · 720p · 1080p | 720p · 1080p |
| Native audio | Yes, included | Yes | Yes |
| Aspect ratios | adaptive · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | auto · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | auto · 21:9 · 2:1 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 |
| Start + end frame | Yes | Yes | Yes |
| Reference images / clips / audio | 10 / 5 / 5 | 30 / 10 / 10 | None |
| Pinned keyframes | No | No | Up to 10 |
| Document or webpage input | Yes | No | No |
| Second tier | Prime (1.4x, faster + sharper) | None | Draft (~3.5x cheaper, promotable) |
| 4K | No | No | No |
How to read it:
- Wan 3.0 has the 2-second floor and 1080p as the baseline, and is the only one that will read a document or a webpage. Reach for it when you want a lot of finished 1080p-with-sound.
- [Seedance 2.5](/blog/seedance-2-5-live-on-genfire-30-second-ai-video) has the deepest reference pools in the catalog, plus 21:9. Reach for it when a scene is built from many references that all have to hold.
- [Flux 3](/blog/flux-3-black-forest-labs-ai-video-with-audio-keyframes) pins images to exact frame positions and extends an existing clip, with a cheap Draft pass you can promote to 1080p. Reach for it when you care about what appears when.
None of the three has a 4K tier. If 4K is the deliverable, the Kling and Seedance 2.0 lines are where it lives.
Writing Prompts for Wan 3.0
Say how long the shot is inside the prompt. The range is wide enough that "a woman walks through a night market" means something very different at 3 seconds and at 25. Write the beats in order and give them rough weight — a walk-through, then a lift over the crowd, then a hold on the lanterns — so the model knows how to spread the time you bought.
Describe the sound. Audio is on by default and costs nothing extra, so it is being generated whether you thought about it or not. "Sizzling griddles and vendor calls in the mix, music dropping away as the camera lifts" is the difference between a soundtrack and a soundscape. For silence, switch audio off rather than writing "no sound" and hoping.
Give every reference a job. A reference the prompt never mentions is one the model has no reason to bind. Name each and say what it is for: "the woman in @Image1 sits at the piano in @Image2 and plays the melody in @Audio1 — slow dolly from her hands to her face." That sentence binds an identity, binds a location, and tells the model the audio take is a performance rather than ambience.
One control worth knowing: Wan 3.0 runs a prompt-expansion pass by default, elaborating your prompt before it generates, which adds 20 to 60 seconds of latency. It is a toggle on the workflow editor's Wan 3.0 node — turn it off when your prompt is precise and you want it taken literally.
How to Run It
Every surface hits the same endpoints with the same limits.
In the browser studio
Open the AI video generator and pick Wan 3.0 or Wan 3.0 Prime from the WAN group. Switch between text, image, and reference modes in the same panel; attach a start frame and an optional end frame, or drop reference images, clips, and audio into their slots and cite them with the @Image1 chips the studio inserts. Duration is a whole-second picker from 2 to 30, and the studio quotes the credit cost before you generate.
Through the REST API
POST /v1/videos/generations with model video.wan_3 or video.wan_3_prime. The mode follows the fields: prompt alone is text-to-video, image_url (plus optional end_image_url) is image-to-video, and any of reference_image_urls, reference_video_urls, reference_audio_urls, file_url, or web_url is reference-to-video.
curl -X POST https://api.genfire.ai/v1/videos/generations \
-H "Authorization: Bearer $GENFIRE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "video.wan_3_prime",
"prompt": "The woman in @Image1 walks through the night market, rain still on the pavement — handheld through steam and neon, vendor calls and sizzling griddles in the mix, then the camera lifts over the crowd as lanterns sway.",
"reference_image_urls": ["https://.../woman.jpg"],
"reference_video_urls": ["https://.../market-walk.mp4"],
"reference_video_trims": [{ "index": 0, "start": 3, "end": 11 }],
"duration": 20,
"resolution": "1080p",
"aspect_ratio": "16:9",
"generate_audio": true
}'The run comes back queued — poll /v1/runs/{runId} or register a webhook, and call POST /v1/models/estimate-cost first for an exact quote. Keys and docs are at /developers.
Through the MCP server
Point Claude, ChatGPT, Cursor, or any MCP client at mcp.genfire.ai and call genfire_generate_video with model: "video.wan_3". It takes the same fields, enforces the 10/5/5 pools, and exposes file_url and web_url — which makes "turn this spec sheet into a 15-second product video" a single tool call.
From the CLI
npm i -g @genfire/cli && genfire auth login
genfire generate video "Two-second insert: a match strikes in close-up, flares, and the flame steadies. Macro, shallow focus, the scrape and hiss audible." \
-m video.wan_3 -d 2 -r 1080p -a 16:9 -o match.mp4-i/--image and --end-image set the start and landing frames; --ref-image, --ref-video, and --ref-audio fill the pools and take URLs or local paths (local files upload automatically); --ref-video-trim 0:3-11 cuts a window out of a reference clip; --no-audio renders it silent.
Pricing
Genfire is pay-as-you-go. Credit packs start at $19 for 1,000 credits, purchased credits last 12 months, every purchase includes a commercial license, and optional monthly plans start at $29 a month. Wan 3.0 bills per second of video; lower resolutions cost fewer credits per second, audio is included rather than surcharged, and Prime is 1.4x standard. The studio, the API's estimate endpoint, and genfire cost video all quote the exact number before a run starts. Current rates are on the pricing page.
Where It Fits
Wan 3.0 is the model to default to when you want finished 1080p with sound and a lot of it — including the short clips other models will not sell you. Keep WAN 2.5 for Animate Move and Animate Replace. Look at Seedance 2.5 when a scene needs more references than Wan's pools hold, and at Flux 3 for pinned keyframes or a clip extended.
Open the video studio and pick Wan 3.0, or read the full spec at /wan-3. Credit packs are on the pricing page.