What's New on Genfire: Summer 2026 Recap
Eight video models, six image families, a rebuilt audio stack, Meshy v7, faceless channels, games, apps, book studios and a five-surface developer stack.
A Lot Happened
Summer 2026 was the busiest stretch Genfire has had. Nearly every model family on the platform gained a new generation, three whole product surfaces went live, and the developer stack grew from an API into five separate ways in.
This is the map. Each section is short, each one links to the page where the detail lives, and everything below is something that is actually shipped — no roadmap items dressed up as features.
New Video Models
Eight video families landed or moved forward this season.
[Seedance 2.5](/seedance-2-5) is the headline. ByteDance's newest model renders 4 to 30 seconds as one continuous generation — scene changes, camera moves and tempo shifts inside a single pass, with native synced audio. It takes up to 30 reference images, 10 reference clips and 10 audio takes, and it is the default in the video studio.
[Wan 3.0](/wan-3) from Alibaba matches that 30-second ceiling and starts shorter, at 2 seconds. It defaults to 1080p, carries native audio, and takes 10 image, 5 video and 5 audio references. Wan 3.0 Prime is the faster, higher-fidelity tier at 1.4x the price, with identical inputs.
[Flux 3](/flux-3) from Black Forest Labs is the routing model: one id that resolves to text-to-video, image-to-video, first-and-last-frame interpolation, keyframe pinning (up to 10 images pinned to positions in the clip), or extending an existing video. 5 to 20 seconds, 720p or 1080p, synchronized audio included at no extra cost. Flux 3 Draft renders the same shot at 720p for roughly 3.5x less, then hands back a cache reference you pass in to re-render at full 1080p.
[Hailuo 03 and the H3 family](/hailuo-03) from MiniMax cover the speed end. H3 Max renders 480p, 768p or 1080p from text, a first frame or references. H3 Max Turbo is the same output at half the H3 Max price per second, text-to-video and image-to-video only. Hailuo 03 itself is the 2K member. None of the H3 endpoints generate audio.
[Kling V3](/kling-v3) arrived in six shapes: Standard and Pro at 5 or 10 seconds with native audio, a 4K tier, and Turbo Standard (720p) and Turbo Pro (1080p) at 3 to 15 seconds — Turbo Pro is fast Kling V3 at the standard tier's price. The Turbo endpoints have no audio and no end frame. Kling V3 Motion Control is separate again: hand it a character image plus a motion reference video and it drives the performance.
[Gemini Omni Flash 1.1](/gemini-omni-flash) from Google went to a full resolution ladder — 360p, 720p, 1080p and 4K — across 3 to 10 seconds, with text, image, reference and video-edit modes, 10 reference images and 3 reference videos. It does not output audio.
[Grok Imagine v1.5](/grok-imagine) from xAI covers 1 to 15 seconds at 480p, 720p or 1080p across seven aspect ratios. And [Happy Horse v1.1](/happy-horse) is in preview, running 3 to 15 seconds with the widest aspect-ratio set on the platform — nine, including 21:9, 9:21, 5:4 and 4:5.
New Image Models
Six image families landed alongside them, all in the AI image generator.
[GPT Image 2.5](/gpt-image-2-5) arrived as two endpoints — Flare for speed (better than GPT Image 2 at roughly half the latency) and Sunburst for precision. Both take 16 reference images, edit with masked inpainting, and cost exactly what GPT Image 2 costs.
Seedream 5.0 Pro is ByteDance's flagship photoreal model with multi-reference editing up to 2K and up to 10 references; Seedream 5.0 Lite is in preview with multilingual text rendering. Both sit alongside Seedream 4.5.
[Qwen Image 3](/qwen-image-3) renders bilingual Chinese and English text up to 2K and takes 1 to 3 edit references you address by name in the prompt. [Recraft V4.1](/recraft-v4) brought sharper prompt control plus true-SVG vector output and a Utility runtime that costs the same as the standard tier but runs at pipeline throughput. Grok Imagine 2.0 offers 13 aspect ratios including 20:9 and 19.5:9. [Muse Image](/muse-image) from Meta is the instruction-follower, built for accurate text, plots and QR codes, composing from up to 10 references.
Nano Banana 2 remains the default, and remains the only image model here with a full 1K / 2K / 4K resolution control.
Audio: Speech, Music and Sound
The audio stack got rebuilt around newer defaults.
On the AI voice generator: ElevenLabs Flash v2.5 is now the default speech model — roughly 75 ms latency across 32 languages, with a 40,000-character ceiling. ElevenLabs v3 is the expressive, highest-quality option. ElevenLabs Text to Dialogue generates a multi-speaker conversation as a single file, up to 10 voices in one take with v3 audio tags. ElevenLabs Multilingual v2 covers 29 languages, and Seed Audio 1.0 from ByteDance handles multilingual speech plus full audio scenes, taking up to three reference audio clips.
On the AI music generator: ElevenLabs Music v2 is the new default — up to 10 minutes, chunk-based composition plans, inpainting into an existing song, and video-to-music. Lyria 3 Pro writes structured songs up to 3 minutes with vocals and lyrics for a flat per-generation fee. MiniMax Music 3 sings your lyrics, which are required input rather than a suggestion, and returns 44.1 kHz stereo.
Sound effects run on ElevenLabs SFX up to 30 seconds, and transcription now offers Scribe v2 — 90+ languages, up to 32 speakers, and keyterm hinting on files up to 500 MB.
3D: Meshy v7
The AI 3D model generator moved to Meshy v7 as its default. The change over v6 is ultra mode — higher-fidelity geometry with finer surface detail, available on single-image inputs.
Meshy v7 takes up to four images (it is image-driven, not text-driven), exports GLB, FBX, OBJ, USDZ, BLEND and STL, and lets you set a target polycount from 100 up to 300,000 with quad or triangle topology. Generation runs asynchronously and typically takes 5 to 10 minutes.
Faceless Studio
Faceless Studio is the biggest new product surface: you describe a channel once, and Genfire names it, writes each episode, narrates it, renders it in a locked style and posts it on a schedule.
- Two formats: Shorts (9:16) and Long-form (16:9), from 20 seconds up to 10 minutes
- Fast mode: a Quality / Fast toggle on every episode. Fast routes plain scenes to MiniMax H3 Max Turbo — faster and cheaper — while scenes with reference images stay on the Omni engine. The wizard shows the credit difference before you commit
- Motion styles: Seamless, Scene by scene, or Stills
- Scheduled channels: turn on auto-generation, set 1–6 episodes a day in your own timezone, pick generation times, let the AI choose topics or supply your own list, and auto-publish finished episodes to connected accounts
- Look and voice: 13 niches, 63 locked visual styles, 17 caption presets with five animation modes, and any ElevenLabs voice including your own clones
Two related studios live in the dashboard apps catalog and share the same pipeline ideas: the AI explainer video generator for wide narrated explainers, and the AI music video generator for putting picture to a track.
Genfire Live
Genfire Live is a new kind of stream: every frame is generated as it plays, and the audience decides what happens next. Viewers get one free vote per round and can put credits behind a direction; whichever has the most support when the round timer hits zero is what the stream renders. There is live chat with slow mode and pinned messages, 15-second clipping to your library with a share link, follows with notify-me on scheduled streams, host revenue share, recordings that stay up after the host signs off, and an embeddable player.
You can browse and watch at /live today. Hosting your own stream is still gated ahead of the full launch.
Games and Apps
Two build-and-publish surfaces went live.
[Genfire Games](/games) — describe a game in plain English and Genfire builds it; it runs directly in the browser with no download, no install and no engine. Browser-playable 2D and lightweight 3D: platformers, shooters, runners, puzzles, clickers. The public gallery needs no account. Start one from the AI game generator.
[Genfire Apps](/apps) — the same idea for software. Describe an app and Genfire builds the UI and the data behind it, then opens it at a shareable URL. CRMs, trackers, dashboards, planners, admin panels. Start one from the AI app builder.
Coloring Books and Picture Books
Two book studios shipped, both ending in a print-ready PDF.
The AI coloring book generator turns a theme into a whole book of clean black-and-white pages with blank backs, a colour cover and a spine already built to Amazon KDP's spec. The Picture Book Studio writes and illustrates a story with consistent characters across every page, sets the words in a real typeface, and exports for KDP or a tablet. Both run on the same engine, with format groups for KDP, digital and other sizes — the print cards name trim, bleed and DPI so the file lands correctly the first time.
Video Upscaling
The AI video upscaler now has two paths. Topaz Video Upscale is the default: 2x or 4x through Topaz Proteus, billed per second of output by resolution. Flux Video Upscale runs FLUX 3 super-resolution at 1.5x to 3x with two modes — precise, or creative with an optional guiding prompt — on MP4 sources up to 20 seconds. Flux is materially pricier than Topaz, so reach for it when the creative mode is the point. Still images upscale through Topaz at 2x or 4x.
Social Publishing and the Calendar
Social publishing closes the loop: create it, schedule it, post it. Three platforms are connected — Instagram (Reels, single photos, and carousels of up to 10 images), TikTok (with the privacy level you choose per post), and YouTube (Shorts or standard uploads with your own title and description).
Everything sits on one calendar in month view: drag a post to a new time, or cancel it before it goes out. Post now or schedule for later, fan out to several accounts with a different caption on each, and watch each post move through scheduled, posting, posted or failed with a link to the live result. Recurring series generate a fresh video on your topic and post it on the schedule you set.
Five Ways to Build on Genfire
The developer side grew into five distinct surfaces, all covered at /developers.
[The MCP server](/mcp) is the big one: 110+ tools at mcp.genfire.ai, connectable by OAuth or a bearer key from Claude, Claude Code, Cursor, ChatGPT, Grok, Codex, Gemini CLI or anything speaking Streamable HTTP. Results come back as live interactive widgets in Claude, and there are dedicated endpoints for ChatGPT's connector limits, for Grok, and a lite profile with trimmed schemas for token-budget agents.
[The CLI](/cli) — npm install -g @genfire/cli, or Homebrew, or a curl installer. Browser-based login, an interactive shell with a live job tracker, one-line generation with influencer @-mentions, cost quotes before you spend, and genfire mcp setup to wire the MCP server up without pasting a key.
The REST API at https://api.genfire.ai/v1/ with bearer auth, idempotency keys, signed-URL uploads and webhooks. The TypeScript SDK (@genfire/sdk) wraps it with upload and run-polling helpers. Claude Code skills load prompt-level instructions once instead of paying MCP schema cost on every call.
And five plugins put Genfire inside the tools you already use: Photoshop (select a region, describe the change, keep every other pixel), Premiere Pro (B-roll, voiceover and captions without leaving the timeline), After Effects, Blender (with a local MCP bridge for Claude Code and Claude Desktop), and DaVinci Resolve (generate straight into the media pool from the edit page). None of them ask for an API key — you approve in the browser and the key lives in your OS keychain.
Pricing Stayed Simple
Genfire is pay-as-you-go. Credit packs start at $19 for 1,000 credits, run through Starter (2,000 / $33), Creator (4,500 / $66) and Professional (8,000 / $99), and purchased credits stay valid for 12 months. Every pack unlocks the full studio — all models, commercial use, no watermarks — and includes Teams: 5 seats on one shared credit pool, so a small team spends from one balance instead of five.
Optional monthly plans start at $29 a month if you would rather have credits arrive automatically; the Business plan raises Teams to 25 seats. Nothing above requires a subscription. Full details on the pricing page.
What's Next
The pattern this summer was less about any single model and more about the gap between models closing — the same reference images, the same schedules, the same credit balance, whichever engine you pick. That is the direction: fewer places where switching tools means starting over.
The fastest way to see what changed is to open a studio and generate something. Start with the video generator or the image generator, or spin up a channel in Faceless Studio.