New MiniMax Two-Stage FastH3 2K resolution is here and it's glorious!
Text → video · Image → video · Native audio
VideoNewFast

MiniMax H3 + FastH3 Turbo

Create polished 5- to 15-second videos with dialogue, music, and sound in one generation. Choose FastH3 Turbo, up to 6× faster than Standard, for everyday creation, or Standard H3 for maximum detail. New two-stage FastH3 output renders 2K at the Standard 768p price.

Create with text, images, or your own audio. FastH3 Turbo brings fast iteration, synchronized sound, and now two-stage 2K output.

New · FastH3 two-stage 2K · Beta

MiniMax Two-Stage FastH3 2K resolution is here, and it's glorious

Twice the resolution of Standard 768p at the same price and about the same render time. FastH3 renders the clip, then Sogni's in-house latent-enhance upscale refines it to a full 2688 × 1536 canvas with the dialogue and sound intact.

2688 × 15362K output canvas · 24 fps · stereo audio
2× the resolutionof Standard 768p · four times the pixels
16 Spark ($0.08)per output second · same as Standard 768p
~15 minfor a 15-second clip · about Standard's time
2-stage 2K · FastH3 text to video · 15 s

Straight FastH3 two-stage 2K with no LoRA: 2688 × 1536 at 24 fps, with the dialogue, espresso-machine hiss, and ceramic clinks generated in the same pass. Watch the chalk lettering, the wood grain, and the crema hold their detail.

2-stage 2K + Natural Face & Speech LoRA · 7 s

The same prompt and seed as a shorter take with the Natural Face & Speech LoRA at 0.6, which keeps cheeks, brows, jaw, and lips moving together and makes the spoken lines clearer. Find it in the H3 LoRAs panel; it also works one-stage.

What the second stage does

A two-stage render starts as a normal FastH3 clip at 1344 × 768. Instead of decoding it to pixels and stretching it, the second stage enlarges the latent itself 2× with a learned 3D latent upscaler, then runs a short tiled refinement pass with the same FastH3 model at the full 2688 × 1536 canvas. Your keyframes are re-applied at the output size, colour is matched to the first pass, and the generated dialogue and sound are untouched. Fine detail holds up — skin, fabric weave, chalk lettering, crema — where a plain upscale goes soft.

It is Sogni's in-house optimized latent-enhance-upscale, built on the open-source ComfyUI MiniMax H3 Latent Upscaler by LBH-123-AI. See how it fared against FlashVSR v1.1 and three other 768p-to-1440p upscalers in The Best Open-Weight 2K Upscaler for MiniMax H3 Video? on the Sogni blog, which also features both clips above. Choose 2-stage 768p, 2-stage 1080p, or 2-stage 2K in the Video Size widget of MiniMax H3's FastH3 settings in Sogni Web, including Sound to Video, or send one of the six two-stage model IDs through the API.

Same price as Standard 768p. 2-stage 768p costs the plain FastH3 rate of 4 Spark ($0.02) per output second. 1080p adds 6 Spark ($0.03) and 2K adds 12 Spark ($0.06), so a 2K clip lands at 16 Spark ($0.08) per second — exactly what Standard 768p costs — and a 15-second clip renders in roughly the same 15 minutes, with four times the pixels. Sizes and pricing table →

API model IDs: minimax-h3-fastvideo-int8_t2v_turbo_2stage, …_i2v_turbo_2stage, …_flf2v_turbo_2stage, …_a2v_turbo_2stage, …_ia2v_turbo_2stage, and …_flfa2v_turbo_2stage. Beta: two-stage output is separate from MiniMax's hosted 2K service, and the one-stage “up to 6× faster” figure does not describe two-stage render times.

The prompt behind both clips
integrated_multimodal_description: [Shot 1] Live-action, cinematic, one continuous medium shot at eye level. A small neighborhood coffee bar at golden hour, sunlight raking through the front window across a worn wooden counter, a brass espresso machine, a chalkboard menu with hand-lettered prices, a jar of biscotti, and steam curling from a milk pitcher. Maya, a barista in her late twenties with short dark curls, a denim apron and a small silver nose ring, steams milk with practiced hands, taps the pitcher twice on the counter, and pours a rosetta into a wide cup while glancing up. Across the counter Theo, a tall man in his forties with a grey beard and a rumpled linen shirt, leans on one elbow holding a folded newspaper. The camera holds steady with a very slight push in; no cuts, no pans. Maya, a warm, quick, slightly teasing voice (S1), says: <d>[English] Extra shot, no sugar, same as every Tuesday.</d> Theo, a low, dry, amused voice (S2), says: <d>[English] You remembered the sugar part this time.</d> Maya slides the cup across, the foam art intact, and wipes the steam wand with a cloth as Theo lifts the cup. Fine detail on the wood grain, the crema, the chalk lettering and the fabric weave of the apron. There is no on-screen text, subtitle, logo or watermark, and no cut, morph, dissolve or crossfade.

overall_soundscape: Espresso machine hiss and gurgle, the clink of ceramic on wood, a faint radio somewhere in the back, street traffic muffled through the window.

non_diegetic_music: N/A
Use prompt →

Both clips were rendered on the Sogni Supernet with FastH3 two-stage 2K (seed 771314) and re-encoded for the web at their native 2688 × 1536. Render time and cost are illustrative, exclude queue time, and follow the live quote and the rendered duration.

About

MiniMax H3 creates complete short-form videos with picture and sound in one generation. Direct the action, camera, dialogue, ambience, effects, and music in the same prompt — no separate soundtrack pass required. It is a strong choice for dialogue scenes, brand films, product reveals, fashion clips, motion design, music visuals, and social content.

Start from a text prompt, animate an opening image, guide a transition with opening and closing frames, or use image, video, and audio references to shape the result. Choose the workflow that matches how much creative control you want. Sound to Video also accepts an audio track on its own, with a starting image, or with both opening and closing frames.

Start with FastH3 Turbo for everyday creation. Built on FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, it renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. It is ideal for drafts, timing tests, quick iteration, and most social content, and its text and frame-guided workflows work with the same H3 LoRAs. Sound to Video does not accept artist LoRAs. Choose Standard H3 when you want the very best fine detail and audio polish. The familiar LightX2V Turbo stays one switch away, and multi-reference video is available in Standard, Balanced, and LightX2V Turbo.

Create videos from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts, with dialogue, music, and sound generated together. Sogni offers 480p plus 544/768p-class output from the open-weights release, and FastH3 two-stage output (Beta) adds 2-stage 768p, 1080p, and 2K: a second pass refines the clip at twice its width and height with Sogni's in-house latent-enhance upscale, and 2K costs the same per second as Standard 768p. It is separate from MiniMax’s hosted 2K service. Reference to Video has its own two-stage output (Beta) on the Standard and Balanced tiers: 2-stage 768p renders at half size, then enlarges and refines the clip with the same model, so a 15-second reference video also renders on 32 GB cards such as the RTX 5090, at the same price as 1-stage 768p; 2-stage 1080p and 2K for Reference to Video are coming soon. MiniMax reports stable dialogue support in 11 languages, with additional languages supported to varying degrees.

Sogni Web makes FastH3 the default Turbo engine for text-to-video, image-to-video, and first-and-last-frame video, with a switch back to LightX2V Turbo. The Sogni API accepts an exact model ID for every Standard, Balanced, LightX2V Turbo, FastH3 Turbo, FastH3 two-stage, and two-stage Reference to Video workflow, and the Creative Agent Skill offers Standard, LightX2V Turbo, and FastH3 Turbo selectors. Multi-reference video accepts up to nine images, three videos, and three audio clips in Standard, Balanced, and LightX2V Turbo; FastH3 does not offer a multi-reference mode, and the 2-stage choices offered under the Turbo tier of Reference to Video render on Standard.

Choose your H3 workflow and speed

Choose FastH3 Turbo for fast everyday creation — up to 6× faster than Standard — LightX2V Turbo for its familiar four-step look, or Standard when you want maximum detail and audio polish. Multi-reference video is available in Standard, Balanced, and LightX2V Turbo. Sound to Video uses FastH3 Turbo with your audio and optional opening and closing frames. Two-stage 768p, 1080p, and 2K output (Beta) is a Video Size choice in FastH3 settings and has its own six model IDs. Two-stage Reference to Video (Beta) is a Video Size choice on Standard and Balanced — 2-stage 768p now, 1080p and 2K coming soon — with two more model IDs.

Workflow Details Access
Text to video Standard Workflow: Create video and audio from a written scene
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
First or last frame Standard Workflow: Animate from an opening image or converge on a closing image
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
First and last frames Standard Workflow: Direct the motion between two compositions
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Multi-reference Standard Workflow: Condition a scene on labelled image, video, and audio references
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Two-stage multi-reference Standard Workflow: Condition a scene on labelled references at half size, then enlarge and refine it with the same model
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p (1344 × 768) · 1080p and 2K coming soon · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Balanced two-stage multi-reference Balanced Workflow: Condition a scene on labelled references at half size with the Balanced recipe, then enlarge and refine it with the same model
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p (1344 × 768) · 1080p and 2K coming soon · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo text to video FastH3 Turbo Workflow: Create video and audio from text with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo first or last frame FastH3 Turbo Workflow: Animate from an opening image or converge on a closing image with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo first and last frames FastH3 Turbo Workflow: Connect two anchor images with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo text to video Turbo Workflow: Create video and audio from text with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo first or last frame Turbo Workflow: Animate from an opening image or converge on a closing image with Turbo
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo first and last frames Turbo Workflow: Connect two anchor images with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo multi-reference Turbo Workflow: Condition a scene on labelled image, video, and audio references with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 544–1344 px · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo audio to video FastH3 Turbo Workflow: Guide video with an uploaded audio track
Duration: Roughly 5–15 seconds · fixed 24 fps
Resolution: 576–1344 px · fixed 24 fps
References: One audio track
Steps: 4 default · 4–4
Spark or eligible Unlimited Create →
FastH3 Turbo image + audio to video FastH3 Turbo Workflow: Guide video with an uploaded audio track
Duration: Roughly 5–15 seconds · fixed 24 fps
Resolution: 576–1344 px · fixed 24 fps
References: One audio track and a starting image
Steps: 4 default · 4–4
Spark or eligible Unlimited Create →
FastH3 Turbo first + last frame + audio FastH3 Turbo Workflow: Guide video with an uploaded audio track
Duration: Roughly 5–15 seconds · fixed 24 fps
Resolution: 576–1344 px · fixed 24 fps
References: One audio track plus opening and closing images
Steps: 4 default · 4–4
Spark or eligible Unlimited Create →
FastH3 two-stage text to video FastH3 Turbo Workflow: Create video and audio from text, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: No image needed
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →
FastH3 two-stage first or last frame FastH3 Turbo Workflow: Animate from an opening image or converge on a closing image, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: One opening or closing frame
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →
FastH3 two-stage first and last frames FastH3 Turbo Workflow: Direct the motion between two compositions, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: Opening and closing frames
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →
FastH3 two-stage audio to video FastH3 Turbo Workflow: Guide video with an uploaded audio track, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: One audio track
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →
FastH3 two-stage image + audio to video FastH3 Turbo Workflow: Guide video with an uploaded audio track, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: One audio track and a starting image
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →
FastH3 two-stage first + last frame + audio FastH3 Turbo Workflow: Guide video with an uploaded audio track, then refine at twice the size
Duration: Roughly 5–15 seconds
Resolution: 2-stage 768p · 1080p · 2K (2688 × 1536) · fixed 24 fps
References: One audio track plus opening and closing images
Steps: 4 FastH3 + 2 refinement
Spark or eligible Unlimited Create →

Seven optional MiniMax H3 video LoRAs

Open the H3 LoRAs panel to add realism, natural speech, better motion, fight choreography, a vintage look, tighter prompt adherence, or opt-in mature-theme knowledge without changing the underlying MiniMax H3 workflow.

  • Realism People by fal — improves faces, skin texture, hands, lighting, and natural human motion.
  • Natural Face & Speech by AdaptiveVision — makes people talking on camera look and sound more natural, with cheeks, brows, jaw, and lips moving together and clearer spoken English. Built for vlogs, podcasts, interviews, and presenters.
  • Better Motion by AdaptiveVision — gives full-body movement more natural weight shifts, strides, turns, and gestures for dance, sport, and walking shots.
  • Combat Base V2 by FourBunny — strengthens continuous fight exchanges, hit reactions, grappling, takedowns, collisions, and falls. Describe the fight as a step-by-step sequence of cause and effect.
  • vh5tape Worn VHS by KennethFal — makes picture and sound look recorded off 1980s broadcast TV onto a worn VHS tape. Start the prompt with vh5tape.
  • VBVR Video Reasoning by MisticRain69 — improves literal prompt adherence and ordered actions.
  • Mystic X v4 by alcaitiff — adds opt-in mature-theme knowledge while preserving motion and fine detail.

All seven work with Standard, Balanced, LightX2V Turbo, FastH3 Turbo, and FastH3 two-stage text, image, and first-and-last-frame video, and every one except Mystic X v4 also works with multi-reference video. Sound to Video does not accept artist LoRAs. VBVR Video Reasoning and Mystic X v4 appear after the Sensitive Content Filter is turned off. For a discreet overview of supported mature-theme workflows, see Sogni's uncensored AI video generator.

Prompting tips

We recommend leaving Sogni's AI Script Writer on. It is enabled by default in many Sogni interfaces and automagically turns a simple idea into a rich script tailored to H3. You can still add shot, camera, style, dialogue, and sound details whenever you want more control.

  • Start with one sentence — With AI Script Writer on, a simple idea is enough. Say who or what is in the scene, what happens, and the vibe you want.
  • Choose the right anchor — Use text alone for T2VA, an opening frame for I2VA, a closing frame for L2VA, both endpoints for FL2VA, or loose labelled image, video, and audio references for Ref2VA.
  • Preserve exact words — Put must-keep dialogue and visible text in your request exactly as written. If you ask for speech without writing a line, Script Writer can author one concise line; otherwise it should not invent speech.
  • Direct sound as carefully as picture — Name ambience, physical effects, diegetic sound, and audience-only music separately because H3 generates native video and stereo audio together.
  • Give every reference one job — For Ref2VA, say which asset controls identity, style, composition, motion, pacing, voice, ambience, or music, and state which source wins if they conflict.
  • Let Script Writer place endpoint alignment — I2VA, L2VA, and FL2VA require MiniMax's exact opening line. It names the referenced picture, actual final shot, and snapped endpoint; T2VA and Ref2VA have no alignment line.
  • Focus on one finished beat — H3 works best when each 5- to 15-second clip has a clear action, development, and payoff instead of trying to compress an entire story.

MiniMax H3 prompting guide

For everyday creation, leave Sogni's AI Script Writer on and write a natural brief. It expands the idea into MiniMax's model-native structure. Advanced callers can write that structure directly using the guide below; small formatting imperfections are worth correcting, but they should not replace the creative substance of a useful prompt.

1. Pick the input shape

  • T2VA — text only; establish the whole scene.
  • I2VA — one opening frame; continue forward from it.
  • L2VA — one closing frame; infer a plausible earlier state and converge on it. Sogni routes this through the I2V model with a closing-frame role, not a separate L2V model ID.
  • FL2VA — opening and closing frames; describe the continuous physical path between both endpoints.
  • Ref2VA — loose labelled image, video, and audio references; assign every asset a specific job.

2. Use the three Base fields

integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...

[Shot 1] has no timestamp. Every later cut uses [Shot N] At MM:SS.mmm, ..., with contiguous shot numbers and strictly increasing times inside the rendered duration. Describe visible action, camera, dialogue, singing, diegetic music, and synchronized events in the main description. Put ambience, effects, and non-verbal human sound in overall_soundscape. Put only audience-only score in non_diegetic_music, or use N/A when there is no score.

3. Preserve speech and text precisely

Give each vocal source a stable ID such as (S1). Keep identity, action, and delivery outside the dialogue tag; put only the language and words inside <d>[English] Exact words.</d>. Preserve user-supplied dialogue character-for-character. If the request explicitly asks for speaking, dialogue, lyrics, or a vocal performance without supplying words, author one concise line; with no vocal intent, invent no speech. Use <scenetrans> at both connection points when one line crosses a cut and plain <cutoff> only when the video ending truncates speech. Never write tokenizer-internal <|...|> controls in a prompt, and never fix one by merely removing its pipe characters: caption markers become exact visible text in double quotes, lyrics markers become an ordinary <d>[Language] exact words</d> singing block, and <|cutoff|> becomes plain <cutoff>. Plain caption or lyrics boundary tags are not valid substitutes.

4. Add the endpoint alignment line

I2VA begins with the exact opening-frame alignment sentence. L2VA and FL2VA begin with their duration-aware alignment sentence naming the actual final shot and snapped endpoint. T2VA and Ref2VA use no alignment sentence. The Script Writer handles these lines automatically when it knows which frames are attached.

5. Use the six Ref2VA fields

subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:

Keep <Subject N>, <Picture N>, <Video N>, and <Audio N> meanings stable. The summary starts with task types chosen from keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference; join multiple types with exactly + and never repeat one. Attached clips are loose references unless the runtime explicitly provides an edit or continuation relationship, so do not promise source-video transformation from file presence alone.

Use fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visual retention; use fully_copy, partially_copy, reference, or weak_reference for audio. Bind a reference voice to its subject's (Sx) in subject_definitions, not in retention_analysis. A timbre, rhythm, emotion, or delivery reference does not authorize copying its words. Preserve explicitly reused speech and mark unintelligible spans [unclear] instead of guessing.

Ref2VA requires at least one image or video. It accepts up to nine images, three videos, three standalone audio clips, and 12 files total. Each reference video and audio clip must be 2–15 seconds; reference videos may total at most 15 seconds, and reference audio may separately total at most 15 seconds.

6. Stay within the open-weights model's real limits

H3 renders at fixed 24 fps on a 124 + n×17 frame grid, about 5.17–15.08 seconds, with a 7,000-character prompt limit. Preserve required fields, exact user text, reference jobs, retention markers, and shot timing before trimming redundant adjectives or audio prose. Sogni offers 480p plus 544/768p-class output from the open-weights release, FastH3 two-stage 768p, 1080p, and 2K output (Beta), and 2-stage 768p output for Ref2VA (Beta); two-stage 1080p and 2K for Ref2VA are coming soon. Changing the prompt cannot enable an output resolution that is not yet open.

Source: MiniMax's official H3 prompt-writing skill, pinned to the reviewed open-weights revision.

Measured on one RTX 5090: the same 15-second 768p image-to-video prompt rendered in 2 min 6 s on FastH3 Turbo and 11 min 55 s on Standard, with Balanced and LightX2V Turbo in between. Every clip is embedded in FastH3 Speed Test: Four MiniMax H3 Tiers, One RTX 5090, One Cat in a Bath.

See Standard and Turbo side by side in the Sogni Engineering field guide, Your Prompt Is Now a Director.

Go fast with FastH3 Turbo

FastH3 Turbo is our go-to for fast everyday creation. It runs FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, timing tests, quick iterations, and most social content. Choose Standard H3 when maximum fine detail and audio polish matter most.

FastH3 is the default Turbo engine in Sogni Web for text, image, and first-and-last-frame video, with a switch back to the familiar LightX2V Turbo look. Its text and frame-guided workflows work with the H3 LoRAs; Sound to Video does not. It costs 4 Spark ($0.02) per output second, and has no multi-reference mode. Shorter clips see a smaller speed-up than 15-second clips.

MiniMax H3 speed comparison

For 15-second 768p-class video, use Standard as the 1× baseline. Balanced runs the MiniMax H3 LightX2V 8-step recipe, Turbo runs the LightX2V 4-step recipe, and FastH3 Turbo runs FastVideo FastH3 4-step.

Sogni speed Engine Relative speed Illustrative 15-second render 544/768p price
Standard MiniMax H3 · 20 steps 1× baseline 15 min 16 Spark ($0.08) per output second
Balanced LightX2V · 8 steps 7.5 min 10 Spark ($0.05) per output second
Turbo LightX2V · 4 steps 3.75 min 6 Spark ($0.03) per output second
FastH3 Turbo FastVideo FastH3 · 4 steps 2.5 min 4 Spark ($0.02) per output second

The times show relative render speed: if Standard takes 15 minutes, the faster tiers scale from that baseline. They exclude queue time and are not guarantees; workflow, input, worker hardware, and cache state can change results. For an illustrative 15.0-second cost comparison, 544/768p-class output is 240 Spark ($1.20) on Standard, 150 Spark ($0.75) on Balanced, 90 Spark ($0.45) on Turbo, or 60 Spark ($0.30) on FastH3 Turbo; actual billing follows the rendered frame duration, up to 15.08 seconds. At 480p, Standard costs 10 Spark ($0.05), Balanced 6 Spark ($0.03), and both Turbo engines 4 Spark ($0.02) per output second.

Measured, not illustrative. On a single RTX 5090 worker, the same 15.08-second 768 × 1024 image-to-video prompt rendered hot in 11 min 55 s on Standard, 5 min 30 s on Balanced, 3 min 0 s on LightX2V Turbo, and 2 min 6 s on FastH3 Turbo, with 3 min 39 s, 1 min 46 s, 1 min 7 s, and 55 s at 480p. Every clip is embedded so you can judge the quality trade yourself: read the FastH3 speed test.

Make a finished beat, not a silent motion test

MiniMax H3 creates video and sound together. Direct dialogue, ambience, effects, and music in the same brief, then get a complete clip without a separate soundtrack pass.

Choose the control your shot needs

Start from text, an opening frame, a closing frame, both endpoints, or a loose labelled reference set. Standard and Turbo Ref2VA can combine up to nine images, three videos, and three audio clips.

A powerful alternative for open creative work

MiniMax H3 is its own MiniMax model, not a Seedance derivative. Sogni offers 480p plus 544/768p-class output from the published weights, and FastH3 two-stage output (Beta) adds 2-stage 768p, 1080p, and 2K at the Standard 768p price. Sogni's in-house latent-enhance upscale is separate from MiniMax’s hosted 2K service. Creators 18 or older can opt into mature-content creation on Sogni, subject to Sogni's Terms of Use and the applicable MiniMax, LightX2V, and FastVideo licenses.

Where each workflow is available

Sogni Web offers T2VA, I2VA, L2VA, and FL2VA in Standard, Balanced, LightX2V Turbo, and FastH3 Turbo, plus Ref2VA in Standard, Balanced, and LightX2V Turbo. Sound to Video supports audio alone, image + audio, and first-and-last-frame + audio through FastH3 Turbo. FastH3 two-stage 768p, 1080p, and 2K output (Beta) is a Video Size choice in FastH3 settings, including Sound to Video. Two-stage Reference to Video (Beta) is a Video Size choice on Standard and Balanced: 2-stage 768p now, with 1080p and 2K coming soon. The Sogni API accepts every workflow by exact model ID, and the Creative Agent Skill exposes Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors.

Simple base pricing

FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output; an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, even with Turbo output. FastH3 two-stage output adds $0.03 per second for 1080p and $0.06 for 2K, so 2K costs $0.08 per second, the same as Standard 768p; 2-stage 768p has no surcharge. Two-stage Reference to Video costs the tier's own rate at 768p, with the same 1080p and 2K surcharges once those classes open. Sogni shows the current quote before rendering.

Sources and licenses

Review the official MiniMax H3 model card, MiniMax H3 Community License, official MiniMax H3 API pricing, LightX2V H3 Turbo model card, and LightX2V Turbo implementation, FastVideo FastH3 4-step Preview v1 VSA DataFree model card, FastVideo implementation, Hao AI Lab FastH3 preview post, and Kijai INT8 ConvRot conversion used by Sogni. Sogni has received written authorization from MiniMax to offer H3 through the platform.

FastH3 two-stage output: sizes and pricing

Two-stage output (Beta) renders a FastH3 clip, then refines it at twice the width and height with Sogni's in-house latent-enhance upscale, preserving the frame count, keyframes, and audio. Choose 2-stage 768p, 1080p, or 2K in the Video Size widget of FastH3 settings, including Sound to Video; the price class follows the delivered pixels.

Output class Example render canvas Delivered pixels Price per output second
2-stage 768p 672 × 384 1344 × 768 4 Spark ($0.02)
2-stage 1080p 960 × 544 1920 × 1088 10 Spark ($0.05)
2-stage 2K 1344 × 768 2688 × 1536 16 Spark ($0.08)

Prices include FastH3 generation and the second-stage surcharge at the exact canvases shown: no surcharge for 768p, 6 Spark ($0.03) per second for 1080p, and 12 Spark ($0.06) for 2K. Other aspect ratios can fall into a different price class; Sogni shows the live quote before rendering. The one-stage “up to 6× faster” comparison above does not describe two-stage render times: a 2K clip takes roughly as long as Standard 768p.

The second stage is built on the open-source ComfyUI MiniMax H3 Latent Upscaler and its published weights. See the 2K samples, or check pricing for every generation option.

Two-stage Reference to Video (Beta)

Reference to Video has its own two-stage output on the Standard and Balanced tiers, chosen in the same Video Size widget; 1-stage 768p stays available. 2-stage 768p renders at half size, then enlarges and refines the clip with the same model, so a 15-second reference video renders on 32 GB cards such as the RTX 5090 instead of waiting for the largest workers. 2-stage 1080p and 2-stage 2K render the first pass on the tier's model and hand the 2× enlargement to FastH3 for refinement; both are coming soon. Under the Turbo tier, the 2-stage choices render on Standard at the Standard rate.

Output class Example render canvas Delivered pixels Price per output second Status
2-stage 768p 672 × 384 1344 × 768 Tier rate: Standard 16 Spark ($0.08) · Balanced 10 Spark ($0.05) Available
2-stage 1080p 960 × 544 1920 × 1088 Tier rate + 6 Spark ($0.03) Coming soon
2-stage 2K 1344 × 768 2688 × 1536 Tier rate + 12 Spark ($0.06) Coming soon

2-stage 768p costs exactly what 1-stage 768p does on the same tier. API model IDs: minimax-h3-ref2va-fp8_r2v_2stage (Standard) and minimax-h3-ref2va-fp8_r2v_balanced_2stage (Balanced); send a canvas with a 384 px short edge, such as 672 × 384, for 768p output.

Pricing

FastH3 Turbo, built on FastVideo FastH3 4-step Preview v1, is Sogni's fastest H3 path — up to 6× faster than Standard — and costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output. An exact 8-second, 192-frame FastH3 clip costs 32 Spark ($0.16). Balanced uses the qualified 8-step LightX2V recipe for text and frame-guided video and is about 2× faster than 20-step Standard; quality is still being evaluated. At 480p, Standard output costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at $0.05 per second for 480p or $0.08 for 544/768p, regardless of output tier. Standard and Balanced multi-reference jobs can also include the reference-image charge shown in the live quote. FastH3 two-stage output (Beta) adds 6 Spark ($0.03) per output second for 1080p and 12 Spark ($0.06) for 2K on top of the FastH3 rate, so 2K costs 16 Spark ($0.08) per second, the same as Standard 768p; 2-stage 768p has no surcharge. Two-stage Reference to Video (Beta) costs the tier's own rate: 2-stage 768p is the same price as 1-stage 768p, and once 1080p and 2K open they add the same 6 Spark ($0.03) and 12 Spark ($0.06) per output second. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.

API

Start with a natural creative brief. Creative Agent expands it into the production prompt MiniMax H3 needs and runs the matching workflow.

const response = await fetch('https://api.sogni.ai/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.SOGNI_API_KEY}`,
  },
  body: JSON.stringify({
    messages: [{ role: 'user', content: "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
    sogni_tools: 'creative-agent',
    sogni_tool_execution: true,
  }),
});

if (!response.ok) throw new Error(await response.text());
const { choices } = await response.json();
console.log(choices[0].message.content);
curl https://api.sogni.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SOGNI_API_KEY" \
  -d '{
    "messages": [{ "role": "user", "content": "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
    "sogni_tools": "creative-agent",
    "sogni_tool_execution": true
  }'

Creative Agent accepts the short brief above and prepares MiniMax H3's structured production prompt. Calling a worker directly? Follow the MiniMax H3 integration guide and use one of these exact model IDs: minimax-h3-fl2va-fp8_t2v, minimax-h3-fl2va-fp8_i2v, minimax-h3-fl2va-fp8_flf2v, minimax-h3-ref2va-fp8_r2v, minimax-h3-ref2va-fp8_r2v_2stage, minimax-h3-ref2va-fp8_r2v_balanced_2stage, minimax-h3-fastvideo-int8_t2v_turbo, minimax-h3-fastvideo-int8_i2v_turbo, minimax-h3-fastvideo-int8_flf2v_turbo, minimax-h3-fl2va-fp8_t2v_turbo, minimax-h3-fl2va-fp8_i2v_turbo, minimax-h3-fl2va-fp8_flf2v_turbo, minimax-h3-ref2va-fp8_r2v_turbo, minimax-h3-fastvideo-int8_a2v_turbo, minimax-h3-fastvideo-int8_ia2v_turbo, minimax-h3-fastvideo-int8_flfa2v_turbo, minimax-h3-fastvideo-int8_t2v_turbo_2stage, minimax-h3-fastvideo-int8_i2v_turbo_2stage, minimax-h3-fastvideo-int8_flf2v_turbo_2stage, minimax-h3-fastvideo-int8_a2v_turbo_2stage, minimax-h3-fastvideo-int8_ia2v_turbo_2stage, minimax-h3-fastvideo-int8_flfa2v_turbo_2stage. Full reference at docs.sogni.ai.

Why run it on Sogni

Subscriptions or Spark

Use a flat monthly plan for credit-free fair-use generation, or buy Spark packs when pay-as-you-go fits better. Both run on the same creator-owned GPU network.

Unlimited plans

One flat price in the app. Generate under fair use without a per-image meter.

🧩

200+ models

Image, video, music, and language models in one workspace and one API key.

Pay-as-you-go Spark

Prefer pay-as-you-go? Call MiniMax H3 by id and pay with Spark packs.

🌐

Powered by people

Runs on a decentralized GPU network where workers share subscription revenue.

FAQ

MiniMax H3 on Sogni

Is MiniMax H3 an uncensored alternative to Seedance 2.0?

Yes — if you are looking for uncensored or NSFW AI video with short-form references, dialogue, built-in sound, and an explicit mature-theme opt-in, MiniMax H3 is a strong alternative. H3 is its own MiniMax model, not a Seedance derivative or modified checkpoint. All creation remains subject to Sogni's Terms of Use and MiniMax's rules. Compare Sogni's opt-in mature-theme video options.

What is MiniMax H3 best for?

H3 is a strong fit for short brand films, dialogue scenes, product reveals, fashion clips, animated posters and title cards, motion design, interface animation, music visuals, and stylized social video where sound matters as much as motion.

Which MiniMax H3 video LoRAs are available on Sogni?

Seven. Realism People improves faces, skin, hands, lighting, and natural human motion. Natural Face & Speech makes people talking on camera move their whole face naturally and speak more clearly. Better Motion gives full-body movement more natural weight shifts, strides, and gestures. Combat Base V2 strengthens fight exchanges, hit reactions, grappling, and falls. vh5tape Worn VHS gives footage and sound a worn 1980s VHS look. VBVR Video Reasoning tightens literal prompt adherence, and Mystic X v4 adds opt-in mature-theme knowledge. All seven work with Standard, Balanced, LightX2V Turbo, FastH3 Turbo, and FastH3 two-stage text, image, and first-and-last-frame video; every one except Mystic X v4 also works with multi-reference video. Sound to Video does not accept artist LoRAs.

How do I enable opt-in mature or NSFW MiniMax H3 generation?

At app.sogni.ai, open your username menu, switch off Sensitive Content Filter under Content preferences, and confirm that you are 18 or older. This reveals 18+ models and the VBVR Video Reasoning and Mystic X v4 LoRAs. Mature-theme generation is strictly opt-in and remains subject to Sogni's Terms of Use and each model's license.

What is FastH3 Turbo, and how much faster is it?

FastH3 Turbo is Sogni's fastest MiniMax H3 path and the default Turbo engine in Sogni Web. It runs the FastVideo FastH3 4-step Preview v1 VSA DataFree checkpoint, a four-step distillation of MiniMax H3 from the FastVideo team at Hao AI Lab, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, quick iterations, and most social content, and choose Standard H3 when you want maximum fine detail and audio polish. The familiar LightX2V Turbo stays one switch away.

What is FastVideo FastH3 4-step Preview v1 VSA DataFree?

It is the FastH3 checkpoint published by the FastVideo team at Hao AI Lab: a preview distillation of MiniMax H3 that generates synchronized video and audio in four transformer passes instead of twenty. VSA is video sparse attention, which skips most attention work at 90% sparsity, and DataFree means the distillation ran without an external training dataset. Sogni runs Kijai's INT8 ConvRot conversion of this checkpoint as FastH3 Turbo for text-to-video, image-to-video, and first-and-last-frame video. Shorter clips see a smaller speed-up than the up to 6× measured on 15-second clips.

Does MiniMax H3 support 2K resolution on Sogni?

Yes. FastH3 two-stage output (Beta) renders 2-stage 768p, 1080p, and 2K: FastH3 generates the clip at the render canvas, then Sogni's in-house latent-enhance upscale, built on the open-source ComfyUI MiniMax H3 Latent Upscaler, refines it at twice the width and height while preserving frame count, keyframes, and audio. A 1344 × 768 canvas delivers 2688 × 1536 pixels, twice the resolution of Standard 768p, for the same 16 Spark ($0.08) per output second and roughly the same render time. Choose it in the Video Size widget of FastH3 settings, including Sound to Video. It is separate from MiniMax’s hosted 2K service. Reference to Video's own 2-stage 1080p and 2K output — first pass on Standard or Balanced, enlargement refined by FastH3 — is coming soon.

How much does two-stage 2K MiniMax H3 cost, and how long does it take?

Two-stage 768p costs the plain FastH3 rate of 4 Spark ($0.02) per output second. 1080p adds 6 Spark ($0.03) per second and 2K adds 12 Spark ($0.06), so a 2K clip costs 16 Spark ($0.08) per second, exactly the Standard 768p rate: an illustrative 15-second 2K clip is 240 Spark ($1.20). Render time is roughly the same as Standard 768p, about 15 minutes for a 15-second clip once rendering starts, because the second stage refines four times the pixels in a short tiled pass. Times exclude queue time and vary with worker hardware and cache state; Sogni shows the live quote before rendering.

Can I create MiniMax H3 video from my own audio?

Yes. FastH3 Sound to Video accepts an audio track alone, with a starting image, or with opening and closing frames. These are separate workflows from multi-reference Ref2VA, and they do not support artist LoRAs. One-stage output and two-stage 768p, 1080p, and 2K output are all available for Sound to Video.

Does Reference to Video offer two-stage output?

Yes, in Beta. On the Standard and Balanced tiers, the Video Size widget offers 2-stage 768p: Reference to Video renders at half size, then enlarges and refines the clip with the same model. It delivers the same 1344 × 768-class output as 1-stage 768p at the same price, and it lets a 15-second reference video render on 32 GB cards such as the RTX 5090 instead of waiting for the largest workers. 1-stage 768p stays available. 2-stage 1080p and 2K for Reference to Video — first pass on the tier's model, enlargement refined by FastH3 — are coming soon; once open, they add 6 Spark ($0.03) and 12 Spark ($0.06) per output second to the tier's rate. Under the Turbo tier, the 2-stage choices render on Standard at the Standard rate. API model IDs: minimax-h3-ref2va-fp8_r2v_2stage and minimax-h3-ref2va-fp8_r2v_balanced_2stage.

How long does a 15-second MiniMax H3 video take to render?

For a 15-second 768p-class clip, use Standard as the 1× baseline, Balanced at 2×, LightX2V Turbo at 4×, and FastH3 Turbo at 6×. If Standard takes 15 minutes, that works out to about 7.5 minutes on Balanced, 3.75 minutes on LightX2V Turbo, or 2.5 minutes on FastH3 Turbo once rendering starts. These are relative render-time examples, not guarantees; actual time varies by workflow, input, worker hardware, cache state, and queue demand. Pay-as-you-go jobs get top queue priority, followed by Unlimited Pro and Unlimited.

Does MiniMax H3 generate audio?

Yes. H3 creates the video, dialogue, music, ambience, and sound effects together, so your clip can come out ready to watch and share.

Which MiniMax H3 generation modes does Sogni support?

H3 supports text-only T2VA, opening-frame I2VA, closing-frame L2VA, two-endpoint FL2VA, and loose-reference Ref2VA. L2VA reuses the I2V model with a closing-frame role rather than inventing a separate model ID. FastH3 Turbo covers T2VA, I2VA, L2VA, and FL2VA; Ref2VA stays on Standard, Balanced, and LightX2V Turbo, with two-stage output (Beta) on Standard and Balanced. The API exposes every endpoint shape with exact model IDs for Standard, Balanced, LightX2V Turbo, FastH3 Turbo, and two-stage output, and the Creative Agent Skill offers Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors (minimax-h3-fasth3-turbo).

How long are MiniMax H3 videos, and which layouts can I use?

Create clips from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts. Developers can also choose custom sizes through the API.

How much does MiniMax H3 cost on Sogni?

FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output, so an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, regardless of output speed. Sogni shows the current live quote before rendering. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.

How many full-quality 15-second H3 videos can I generate per day on Unlimited?

There is no fixed video count. Video draws on the same daily fair-use capacity as other covered models, which resets every 24 hours on your own plan schedule; Relaxed rendering has no daily capacity, and Premium Spark is unaffected. Unlimited lets you generate one Standard H3 video at a time, while Unlimited Pro lets you generate two and carries 4x the daily capacity. You can keep adding videos to your queue throughout the day, subject to fair use, and see your current standing as a percentage at app.sogni.ai/usage.

What happens if I use MiniMax H3 heavily on Unlimited?

If your daily use is much heavier than normal, the number of videos you can generate at once may be reduced until the next UTC day. Turbo is less affected than Standard, and more than 90% of subscribers do not currently hit a fair-use slowdown. Speeds and plan terms are subject to change as Sogni balances supply and demand to keep renders fast for artists and rewarding for the people who share their GPUs through our people-powered render network. Sogni Unlimited is the best deal in town for frequent creators, and we plan to keep it that way.

Can I continuously queue MiniMax H3 generations throughout the day?

Yes. Individual creators can keep a substantial queue moving throughout the day, including through the API and Creative Agent. Unlimited is not intended to power unattended 24/7 or multi-user production systems; use pay-as-you-go Spark or an Enterprise plan for those workloads. Read the Creative Agent Skill guide.

Does Sogni's H3 release support image, video, or audio references?

Yes. Standard and Turbo Ref2VA let you guide a clip with up to nine images, three videos, and three audio clips — 12 files total, with at least one image or video. Video and audio references must each be 2–15 seconds; video references may total at most 15 seconds, and audio references have a separate 15-second total. Choose 2-stage 768p in Video Size on Standard or Balanced to render a 15-second reference video on 32 GB cards at the same price as 1-stage 768p.

Can I use MiniMax H3 in Sogni Web, the API, and Creative Agent?

Yes. Sogni Web, the API, and Creative Agent support text-to-video, image-to-video, first-and-last-frame video, and multi-reference video across Standard and Turbo. Pay as you go with Spark or create on an Unlimited plan.

Is MiniMax H3 licensed for use on Sogni?

Yes. MiniMax has given Sogni written authorization to offer H3 in the United States, European Union, United Kingdom, and South Korea. The MiniMax H3 Community License applies to people who self-host the published model, not to videos you generate on Sogni.

How should I structure a MiniMax H3 prompt?

Leave Sogni's AI Script Writer on for a natural-language brief, or follow the full prompting guide on this page. Base modes use three ordered fields; Ref2VA uses six. Shot 1 has no timestamp, later cuts use [Shot N] At MM:SS.mmm, vocal sources keep stable (S1) IDs, and dialogue uses <d>[Language] exact words</d>. Preserve supplied words exactly; only author dialogue when the user explicitly asks for speech without providing a line.

What languages can MiniMax H3 dialogue use?

MiniMax reports stable dialogue support for Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish, with additional languages supported to varying degrees.

How private is my use of this model on Sogni?

Sogni is built with privacy and creative freedom in mind. Your work remains your own, and inference runs through the Sogni Supernet, a decentralized network of creator GPUs, instead of requiring local hardware or a separate model host. Use is still governed by Sogni's Privacy Policy and Terms of Use.

Start with MiniMax H3 today

Create in the app, or build with the API. Your call.