Hailuo 03 on Genfire: 2K Video, H3 Max, H3 Max Turbo, Director Sessions and Trained Models
MiniMax's third-generation video model in four shapes: Hailuo 03 at 2K with trained models, H3 Max, H3 Max Turbo at half the price, and live Director sessions.
Four Ways to Run MiniMax H3
MiniMax's third-generation video model arrives on Genfire as four different things, and picking the right one is most of the work:
- Hailuo 03 — the top tier. 2K output, the full reference set, and the only model on the platform that runs models you trained yourself.
- H3 Max — a post-trained variant with stronger prompt adherence and higher throughput. Tops out at 1080p, keeps references, and accepts far longer prompts.
- H3 Max Turbo — the same three resolutions from text or a first frame at half the H3 Max price per second, with no reference mode.
- H3 Max Director — not a queued render at all. A live session that generates faster than real time and takes new directions mid-scene.
The three queued tiers share a spec sheet: 5 to 15 seconds, any whole second; 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 (image-to-video takes its framing from the first frame instead); text-to-video, image-to-video with an optional last frame, and reference-to-video on everything but Turbo. None of them generates an audio track — the family renders picture, so sound comes from the audio studio or the editor.
The full spec sheet lives at /hailuo-03. This post is about which tier to reach for, what trained models actually do, and how a Director session differs from everything else on Genfire.
Hailuo 03: 2K, Fifteen Reference Files, Trained Models
Hailuo 03 is the version with the most inputs and the highest ceiling. It renders at 768p or 2K from three kinds of brief:
- 1A text prompt — six aspect ratios, 5 to 15 seconds.
- 2A first frame, with an optional last frame. The end frame lives on the image-to-video path, so Genfire only offers it once a start frame is attached — otherwise the keyframe would be quietly dropped from a generation you had already paid for.
- 3A reference set — up to 9 reference images, 3 reference clips, and 3 audio takes, each clip 2 to 15 seconds.
The three pools do different jobs. Images lock a face, a product, or a location. Video supplies a camera move or a piece of choreography to follow. Audio gives a character a consistent voice to perform to — "the woman in Image 1 speaks with the voice in Audio 1" — and can never be the only reference; it needs at least one image or clip alongside it.
One quirk before you write: Hailuo binds references by their plain positional names — Image 1, Video 1, Audio 1, with a space and no @. Genfire's surfaces emit the house @Image1 chip form and normalize it on the way out, so both spellings work from the studio; hand-written API prompts should use the plain form. Hailuo 03's prompt field is also comparatively short at around 2,000 characters, where H3 Max and Turbo take vastly longer ones.
H3 Max and H3 Max Turbo
These are post-trained variants of the same model, tuned for a different trade.
H3 Max gives up the 2K ceiling for 480p, 768p, and 1080p — where 1080p bills twice the 768p rate — and in exchange follows prompts more closely and turns clips around faster. It keeps the whole reference mode (the same 9 / 3 / 3 pools, the same 2–15 second clip window), keeps the first-to-last keyframe, and takes prompts long enough to hold an entire treatment.
H3 Max Turbo is the drafting tier: identical resolutions, durations, aspect ratios, and keyframe support at half the H3 Max price per second. What it does not have is a reference endpoint — no reference mode, no trained models — so the moment a face or a product has to stay consistent, step up to H3 Max.
Which tier to reach for
| Hailuo 03 | H3 Max | H3 Max Turbo | |
|---|---|---|---|
| Resolutions | 768p · 2K | 480p · 768p · 1080p | 480p · 768p · 1080p |
| Duration | 5–15s | 5–15s | 5–15s |
| Aspect ratios | 21:9 → 9:16 | 21:9 → 9:16 | 21:9 → 9:16 |
| Text-to-video | Yes | Yes | Yes |
| First frame + optional last frame | Yes | Yes | Yes |
| Reference-to-video (9 img / 3 clip / 3 audio) | Yes | Yes | No |
| Runs your trained models | Up to 3 | No | No |
| Prompt length | ~2,000 characters | Very long | Very long |
| Per-second price | Highest of the three | Middle | Half of H3 Max |
| Reach for it when | You need 2K, or a trained model | You are finishing at 1080p, or working with references | You are finding the shot |
The working pattern: Turbo to find the shot, H3 Max to finish it at 1080p, Hailuo 03 when the deliverable needs 2K or a trained model. All three take the same prompt shape, durations, and ratios, so moving between them is one click in the studio.
Trained Models: Teach It Your Own Footage
This is the part no other video model on Genfire does. Under Assets → Trained models you train a small adapter on your own clips and then apply it to Hailuo 03 generations. You pick what you are teaching, and the kind decides how the finished model runs:
- Style — a look, a grade, an animation language, a camera vocabulary. Runs later as image-to-video, start frame optional. Feed it clips that share the look; a variety of subjects helps.
- Subject — a person, a product, a mascot, a character. Runs later as reference-to-video, so live references can ride along on top. Feed it the same subject from many angles, and give each clip up to four reference stills to teach the binding. Subject models are the only kind you can warm-start later with more clips, and the kind Influencer Studio uses.
- Keyframe — how a clip travels from its first frame to its last: reveals, whip pans, product turns. Runs later as image-to-video with an end frame. Feed it clips whose first and last frames show the move.
A fourth kind, prompt-only text-to-video, is reachable through the API.
What the trainer wants. Video files only, 1 to 40 clips, at least ten recommended — fewer will train, but it generalises badly. Each clip stays under 200 MB and the dataset under 1.5 GB, captions are optional, and past Genfire generations can be used as clips directly. You also set a short trigger phrase, which is prepended to every training caption and then added to your prompts automatically, so you never have to remember to type it. Training is billed by step — 2,000 by default, roughly 1,000–1,500 for a clean style, higher for a stubborn subject — and Genfire quotes the exact number first. A run usually takes 20 to 60 minutes, and you can close the tab.
One hard requirement: you must confirm you hold the rights to every clip you train on. Genfire asks, and will not start without it.
Using one. Attach up to three trained models to a single Hailuo 03 generation, each with its own strength dial. The kinds have to agree on a mode — style and keyframe adapters both run image-to-video, subject adapters run reference-to-video — so mixing kinds that land on different modes is rejected rather than silently mangled. And because the adapters are trained against Hailuo 03's own weights, switching the model to H3 Max or Turbo removes them; the studio tells you it did rather than pretending otherwise.
One nice bit of glue: @-mention one of your influencers on a Hailuo 03 reference generation and Genfire binds that influencer's subject model and voice clip for you, so the character looks and sounds the same from clip to clip — see the AI influencer generator. Trained-model runs bill at their own per-second rate rather than the plain Hailuo 03 rate, and the studio quotes it before you generate.
H3 Max Director: A Live Session, Not a Render
Everything above is a queued job: you write a brief, you wait, a file arrives. H3 Max Director is a different shape of thing, and it lives on its own dashboard page rather than in the model picker.
You open a session, give it a brief, and the model generates faster than real time. The stream plays while you type new directions into a bar underneath — "the lights cut out and a dog wanders in" — and the direction lands within seconds, with characters, set, and story carrying across the cut. Format presets get you started (sitcom, cooking show, product live, anime, nature doc, late night), each with a deck of one-tap directions.
The limits, from the code that enforces them:
- Resolution: 480p, 768p, or 1080p. Aspect: 16:9, 9:16, or 1:1.
- Metered by the second while live, with a 60-second minimum held when the session opens. 1080p bills twice the standard rate.
- A single session runs up to 15 minutes, and the server trims that cap down to whatever your balance covers.
- Longer runs chain. Ask for 30 minutes or "until I stop" and the page snapshots the stage before the cap, opens the next session with that frame as its exact first frame, and cuts over when the new stream arrives. The few seconds of overlap bill on both.
- You can open with a start frame, an end frame for the first scene, and a soundtrack that plays until the stream ends, and you can attach a new end frame or swap the soundtrack with any later direction.
- Going live opens a public Genfire Live channel where viewers watch in the browser, chat, and spend credits to pick the next scene.
- Preview mode uses sample data and costs nothing — worth a look before you spend anything.
How It Compares With Hailuo 2.3
Hailuo 2.3 is still on Genfire and still a good, cheap single-take model. The generational jump is about control.
| Hailuo 03 family | Hailuo 2.3 | |
|---|---|---|
| Duration | Any whole second, 5–15s | 6 or 10 seconds |
| Resolution | 768p · 2K (Hailuo 03) · 480p–1080p (Max tiers) | 720p · 1080p |
| Aspect ratios | 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | 16:9 · 9:16 · 1:1 |
| Reference-to-video | Yes — 9 images, 3 clips, 3 audio | No |
| Last-frame keyframe | Yes | No |
| Trained models | Yes, on Hailuo 03 | No |
| Live directed sessions | Yes, via Director | No |
| Tiers | Hailuo 03 · Max · Max Turbo · Director | Standard · Pro · Fast |
A 6-second establishing shot and nothing else is fine on 2.3. Everything about consistency — a face across a campaign, a product that has to look like the product, a shot that has to land on a specific frame — is a Hailuo 03 job.
How to Run It on Genfire
In the browser studio
Open the AI video generator and choose the Minimax Hailuo family. The five picks map to what you are doing: H3 Max Turbo, H3 Max, H3 Max Reference, Hailuo 03, and Hailuo 03 Reference. Attach a first frame (and a last frame if you want one), or open the References panel and fill the image, clip, and audio pools — the studio inserts the citation chips for you. On Hailuo 03 a Trained models control appears next to the model row: pick up to three and set each one's strength. Cost is quoted before you generate, and Director has its own page under Create → Video.
Through the REST API
Three model ids: video.hailuo_03, video.hailuo_03_max, video.hailuo_03_max_turbo. Mode follows the inputs — a prompt alone is text-to-video, image_url makes it image-to-video (add end_image_url for the landing frame), and the reference_*_urls pools make it reference-to-video.
curl -X POST https://api.genfire.ai/v1/videos/generations \
-H "Authorization: Bearer $GENFIRE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "video.hailuo_03",
"prompt": "A boxer in a smoke-filled 1970s gym skips rope under a single hanging bulb, sweat catching the light — slow dolly in, then a whip to the speed bag as the bell rings.",
"resolution": "2k",
"aspect_ratio": "16:9",
"duration": 10,
"loras": [{ "id": "lora_...", "scale": 1 }]
}'loras takes up to three { id, scale } entries, and only on video.hailuo_03. The ids come from the trained-model endpoints: POST /v1/loras starts a training run, GET /v1/loras lists yours, GET /v1/loras/{id} polls one until it reads completed, POST /v1/loras/{id}/continue warm-starts a subject model, and POST /v1/loras/estimate-cost quotes a run first. The trigger phrase is prepended server-side. Runs come back queued — poll /v1/runs/{runId} or register a webhook. Keys and reference docs are at /developers.
Through the MCP server
Point Claude, ChatGPT, Cursor, or any MCP client at mcp.genfire.ai. genfire_generate_video takes the same fields including loras, genfire_list_trained_models enumerates what you have trained (only completed rows can be used), and genfire_train_video_model starts a training run — after asking you to confirm you hold the rights to the clips.
From the CLI
npm i -g @genfire/cli && genfire auth login
genfire generate video "The character in Image 1 walks the neon market in Video 1's handheld style and speaks the line in Audio 1 — same jacket, same scar, rain on the lens." \
-m video.hailuo_03_max -d 12 -r 1080p -a 21:9 \
--ref-image face.jpg --ref-video handheld.mp4 --ref-audio line.mp3 -o market.mp4--image and --end-image set the first and last frames, and --ref-video-trim cuts a window out of a reference clip so only the seconds you want are sent. Trained models are attached from the studio, the API, and MCP — the CLI has no flag for them yet.
Pricing
Genfire is pay-as-you-go. Credit packs start at $19 for 1,000 credits, purchased credits last 12 months, and every purchase includes a commercial license; optional monthly plans start at $29 a month. Video bills per second, so tier and resolution are the dials that matter: 1080p bills twice the 768p rate on both Max tiers, Turbo is half of H3 Max per second, and Hailuo 03 sits above both. Training bills per step; Director bills per second of live session with a 60-second minimum. The studio, POST /v1/models/estimate-cost, and genfire cost video quote the exact number first. Current rates are on the pricing page.
Where It Fits
Use Turbo while you are still deciding what the shot is, H3 Max for finished 1080p work and anything with references, and Hailuo 03 when the deliverable is a 2K master — or when the thing on screen is yours, and a trained model is what keeps it that way. Use Director when the point is not a file at all but a stream you and an audience steer together.
The full spec is at /hailuo-03, the older generation at /hailuo-2-3, and the studio at /ai-video-generator.