Text → video · Image → video · Native audio
VideoNewFast

MiniMax H3 + FastH3 Turbo

Create polished 5- to 15-second videos with dialogue, music, and sound in one generation. Choose FastH3 Turbo, up to 6× faster than Standard, for everyday creation, or Standard H3 for maximum detail.

Create a full 15-second video up to 6× faster with FastH3 Turbo, or choose Standard for maximum polish.

About

MiniMax H3 creates complete short-form videos with picture and sound in one generation. Direct the action, camera, dialogue, ambience, effects, and music in the same prompt — no separate soundtrack pass required. It is a strong choice for dialogue scenes, brand films, product reveals, fashion clips, motion design, music visuals, and social content.

Start from a text prompt, animate an opening image, guide a transition with opening and closing frames, or use image, video, and audio references to shape the result. Choose the workflow that matches how much creative control you want.

Start with FastH3 Turbo for everyday creation. Built on FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, it renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. It is ideal for drafts, timing tests, quick iteration, and most social content, and it works with the same H3 LoRAs. Choose Standard H3 when you want the very best fine detail and audio polish. The familiar LightX2V Turbo stays one switch away, and multi-reference video is available in Standard and LightX2V Turbo.

Create videos from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts, with dialogue, music, and sound generated together. Sogni offers 480p plus 544/768p-class output from the open-weights release; MiniMax's hosted 2K stage is not part of the published weights. MiniMax reports stable dialogue support in 11 languages, with additional languages supported to varying degrees.

Sogni Web makes FastH3 the default Turbo engine for text-to-video, image-to-video, and first-and-last-frame video, with a switch back to LightX2V Turbo. The Sogni API accepts an exact model ID for every Standard, Balanced, LightX2V Turbo, and FastH3 Turbo workflow, and the Creative Agent Skill offers Standard, LightX2V Turbo, and FastH3 Turbo selectors. Multi-reference video accepts up to nine images, three videos, and three audio clips in Standard and LightX2V Turbo; FastH3 does not offer a multi-reference mode.

Choose your H3 workflow and speed

Choose FastH3 Turbo for fast everyday creation — up to 6× faster than Standard — LightX2V Turbo for its familiar four-step look, or Standard when you want maximum detail and audio polish. Multi-reference video is available in Standard and LightX2V Turbo.

Workflow Details Access
Text to video Standard Workflow: Create video and audio from a written scene
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
First or last frame Standard Workflow: Animate from an opening image or converge on a closing image
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
First and last frames Standard Workflow: Direct the motion between two compositions
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited API / agent workflow
Multi-reference Standard Workflow: Condition a scene on labelled image, video, and audio references
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo text to video FastH3 Turbo Workflow: Create video and audio from text with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo first or last frame FastH3 Turbo Workflow: Animate from an opening image or converge on a closing image with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
FastH3 Turbo first and last frames FastH3 Turbo Workflow: Connect two anchor images with the FastH3 engine
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo text to video Turbo Workflow: Create video and audio from text with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: No image needed
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo first or last frame Turbo Workflow: Animate from an opening image or converge on a closing image with Turbo
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening or closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo first and last frames Turbo Workflow: Connect two anchor images with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 576–1344 px · fixed 24 fps
References: One opening frame and one closing frame
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →
Turbo multi-reference Turbo Workflow: Condition a scene on labelled image, video, and audio references with the Turbo path
Duration: Roughly 5–15 seconds
Resolution: 544–1344 px · fixed 24 fps
References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total
Audio: 32 kHz stereo · video at 24 fps
Spark or eligible Unlimited Create →

Three optional MiniMax H3 video LoRAs

Open the H3 LoRAs panel to add realism, tighter prompt adherence, or opt-in mature-theme knowledge without changing the underlying MiniMax H3 workflow.

  • Realism People — improves faces, skin texture, hands, lighting, and natural human motion. It supports Standard and Turbo text-to-video, image-to-video, first-and-last-frame, and reference-to-video.
  • VBVR Video Reasoning — improves literal prompt adherence and ordered actions for Standard and Turbo text-to-video and image-to-video.
  • Mystic X v4 — adds opt-in mature-theme knowledge while preserving motion and fine detail for Standard and Turbo text, image, and first-and-last-frame video.

VBVR Video Reasoning and Mystic X v4 appear after the Sensitive Content Filter is turned off. All three also work with FastH3 Turbo. Balanced beta LoRA support is still being evaluated. For a discreet overview of supported mature-theme workflows, see Sogni's uncensored AI video generator.

Prompting tips

We recommend leaving Sogni's AI Script Writer on. It is enabled by default in many Sogni interfaces and automagically turns a simple idea into a rich script tailored to H3. You can still add shot, camera, style, dialogue, and sound details whenever you want more control.

  • Start with one sentence — With AI Script Writer on, a simple idea is enough. Say who or what is in the scene, what happens, and the vibe you want.
  • Choose the right anchor — Use text alone for T2VA, an opening frame for I2VA, a closing frame for L2VA, both endpoints for FL2VA, or loose labelled image, video, and audio references for Ref2VA.
  • Preserve exact words — Put must-keep dialogue and visible text in your request exactly as written. If you ask for speech without writing a line, Script Writer can author one concise line; otherwise it should not invent speech.
  • Direct sound as carefully as picture — Name ambience, physical effects, diegetic sound, and audience-only music separately because H3 generates native video and stereo audio together.
  • Give every reference one job — For Ref2VA, say which asset controls identity, style, composition, motion, pacing, voice, ambience, or music, and state which source wins if they conflict.
  • Let Script Writer place endpoint alignment — I2VA, L2VA, and FL2VA require MiniMax's exact opening line. It names the referenced picture, actual final shot, and snapped endpoint; T2VA and Ref2VA have no alignment line.
  • Focus on one finished beat — H3 works best when each 5- to 15-second clip has a clear action, development, and payoff instead of trying to compress an entire story.

MiniMax H3 prompting guide

For everyday creation, leave Sogni's AI Script Writer on and write a natural brief. It expands the idea into MiniMax's model-native structure. Advanced callers can write that structure directly using the guide below; small formatting imperfections are worth correcting, but they should not replace the creative substance of a useful prompt.

1. Pick the input shape

  • T2VA — text only; establish the whole scene.
  • I2VA — one opening frame; continue forward from it.
  • L2VA — one closing frame; infer a plausible earlier state and converge on it. Sogni routes this through the I2V model with a closing-frame role, not a separate L2V model ID.
  • FL2VA — opening and closing frames; describe the continuous physical path between both endpoints.
  • Ref2VA — loose labelled image, video, and audio references; assign every asset a specific job.

2. Use the three Base fields

integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...

[Shot 1] has no timestamp. Every later cut uses [Shot N] At MM:SS.mmm, ..., with contiguous shot numbers and strictly increasing times inside the rendered duration. Describe visible action, camera, dialogue, singing, diegetic music, and synchronized events in the main description. Put ambience, effects, and non-verbal human sound in overall_soundscape. Put only audience-only score in non_diegetic_music, or use N/A when there is no score.

3. Preserve speech and text precisely

Give each vocal source a stable ID such as (S1). Keep identity, action, and delivery outside the dialogue tag; put only the language and words inside <d>[English] Exact words.</d>. Preserve user-supplied dialogue character-for-character. If the request explicitly asks for speaking, dialogue, lyrics, or a vocal performance without supplying words, author one concise line; with no vocal intent, invent no speech. Use <scenetrans> at both connection points when one line crosses a cut and plain <cutoff> only when the video ending truncates speech. Never write tokenizer-internal <|...|> controls in a prompt, and never fix one by merely removing its pipe characters: caption markers become exact visible text in double quotes, lyrics markers become an ordinary <d>[Language] exact words</d> singing block, and <|cutoff|> becomes plain <cutoff>. Plain caption or lyrics boundary tags are not valid substitutes.

4. Add the endpoint alignment line

I2VA begins with the exact opening-frame alignment sentence. L2VA and FL2VA begin with their duration-aware alignment sentence naming the actual final shot and snapped endpoint. T2VA and Ref2VA use no alignment sentence. The Script Writer handles these lines automatically when it knows which frames are attached.

5. Use the six Ref2VA fields

subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:

Keep <Subject N>, <Picture N>, <Video N>, and <Audio N> meanings stable. The summary starts with task types chosen from keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference; join multiple types with exactly + and never repeat one. Attached clips are loose references unless the runtime explicitly provides an edit or continuation relationship, so do not promise source-video transformation from file presence alone.

Use fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visual retention; use fully_copy, partially_copy, reference, or weak_reference for audio. Bind a reference voice to its subject's (Sx) in subject_definitions, not in retention_analysis. A timbre, rhythm, emotion, or delivery reference does not authorize copying its words. Preserve explicitly reused speech and mark unintelligible spans [unclear] instead of guessing.

Ref2VA requires at least one image or video. It accepts up to nine images, three videos, three standalone audio clips, and 12 files total. Each reference video and audio clip must be 2–15 seconds; reference videos may total at most 15 seconds, and reference audio may separately total at most 15 seconds.

6. Stay within the open-weights model's real limits

H3 renders at fixed 24 fps on a 124 + n×17 frame grid, about 5.17–15.08 seconds, with a 7,000-character prompt limit. Preserve required fields, exact user text, reference jobs, retention markers, and shot timing before trimming redundant adjectives or audio prose. Sogni offers 480p plus 544/768p-class output from the open-weights release; do not prompt for or promise the hosted-only 2K stage.

Source: MiniMax's official H3 prompt-writing skill, pinned to the reviewed open-weights revision.

Measured on one RTX 5090: the same 15-second 768p image-to-video prompt rendered in 2 min 6 s on FastH3 Turbo and 11 min 55 s on Standard, with Balanced and LightX2V Turbo in between. Every clip is embedded in FastH3 Speed Test: Four MiniMax H3 Tiers, One RTX 5090, One Cat in a Bath.

See Standard and Turbo side by side in the Sogni Engineering field guide, Your Prompt Is Now a Director.

Go fast with FastH3 Turbo

FastH3 Turbo is our go-to for fast everyday creation. It runs FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, timing tests, quick iterations, and most social content. Choose Standard H3 when maximum fine detail and audio polish matter most.

FastH3 is the default Turbo engine in Sogni Web for text, image, and first-and-last-frame video, with a switch back to the familiar LightX2V Turbo look. It works with the H3 LoRAs, costs 4 Spark ($0.02) per output second, and has no multi-reference mode. Shorter clips see a smaller speed-up than 15-second clips.

MiniMax H3 speed comparison

For 15-second 768p-class video, use Standard as the 1× baseline. Balanced runs the MiniMax H3 LightX2V 8-step recipe, Turbo runs the LightX2V 4-step recipe, and FastH3 Turbo runs FastVideo FastH3 4-step.

Sogni speed Engine Relative speed Illustrative 15-second render 544/768p price
Standard MiniMax H3 · 20 steps 1× baseline 15 min 16 Spark ($0.08) per output second
Balanced LightX2V · 8 steps 7.5 min 10 Spark ($0.05) per output second
Turbo LightX2V · 4 steps 3.75 min 6 Spark ($0.03) per output second
FastH3 Turbo FastVideo FastH3 · 4 steps 2.5 min 4 Spark ($0.02) per output second

The times show relative render speed: if Standard takes 15 minutes, the faster tiers scale from that baseline. They exclude queue time and are not guarantees; workflow, input, worker hardware, and cache state can change results. For an illustrative 15.0-second cost comparison, 544/768p-class output is 240 Spark ($1.20) on Standard, 150 Spark ($0.75) on Balanced, 90 Spark ($0.45) on Turbo, or 60 Spark ($0.30) on FastH3 Turbo; actual billing follows the rendered frame duration, up to 15.08 seconds. At 480p, Standard costs 10 Spark ($0.05), Balanced 6 Spark ($0.03), and both Turbo engines 4 Spark ($0.02) per output second.

Measured, not illustrative. On a single RTX 5090 worker, the same 15.08-second 768 × 1024 image-to-video prompt rendered hot in 11 min 55 s on Standard, 5 min 30 s on Balanced, 3 min 0 s on LightX2V Turbo, and 2 min 6 s on FastH3 Turbo, with 3 min 39 s, 1 min 46 s, 1 min 7 s, and 55 s at 480p. Every clip is embedded so you can judge the quality trade yourself: read the FastH3 speed test.

Make a finished beat, not a silent motion test

MiniMax H3 creates video and sound together. Direct dialogue, ambience, effects, and music in the same brief, then get a complete clip without a separate soundtrack pass.

Choose the control your shot needs

Start from text, an opening frame, a closing frame, both endpoints, or a loose labelled reference set. Standard and Turbo Ref2VA can combine up to nine images, three videos, and three audio clips.

A powerful alternative for open creative work

MiniMax H3 is its own MiniMax model, not a Seedance derivative. Sogni offers 480p plus 544/768p-class output from the published weights; MiniMax's hosted 2K stage is not part of the open release. Creators 18 or older can opt into mature-content creation on Sogni, subject to Sogni's Terms of Use and the applicable MiniMax, LightX2V, and FastVideo licenses.

Where each workflow is available

Sogni Web offers T2VA, I2VA, L2VA, and FL2VA in Standard, Balanced, LightX2V Turbo, and FastH3 Turbo, plus Ref2VA in Standard, Balanced, and LightX2V Turbo. The Sogni API accepts every workflow by exact model ID, and the Creative Agent Skill exposes Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors.

Simple base pricing

FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output; an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, even with Turbo output. Sogni shows the current quote before rendering.

Sources and licenses

Review the official MiniMax H3 model card, MiniMax H3 Community License, official MiniMax H3 API pricing, LightX2V H3 Turbo model card, and LightX2V Turbo implementation, FastVideo FastH3 4-step Preview v1 VSA DataFree model card, FastVideo implementation, Hao AI Lab FastH3 preview post, and Kijai INT8 ConvRot conversion used by Sogni. Sogni has received written authorization from MiniMax to offer H3 through the platform.

Pricing

FastH3 Turbo, built on FastVideo FastH3 4-step Preview v1, is Sogni's fastest H3 path — up to 6× faster than Standard — and costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output. An exact 8-second, 192-frame FastH3 clip costs 32 Spark ($0.16). Balanced uses the qualified 8-step LightX2V recipe for text and frame-guided video and is about 2× faster than 20-step Standard; quality is still being evaluated. At 480p, Standard output costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at $0.05 per second for 480p or $0.08 for 544/768p, regardless of output tier. Standard and Balanced multi-reference jobs can also include the reference-image charge shown in the live quote. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.

API

Start with a natural creative brief. Creative Agent expands it into the production prompt MiniMax H3 needs and runs the matching workflow.

const response = await fetch('https://api.sogni.ai/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.SOGNI_API_KEY}`,
  },
  body: JSON.stringify({
    messages: [{ role: 'user', content: "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
    sogni_tools: 'creative-agent',
    sogni_tool_execution: true,
  }),
});

if (!response.ok) throw new Error(await response.text());
const { choices } = await response.json();
console.log(choices[0].message.content);
curl https://api.sogni.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SOGNI_API_KEY" \
  -d '{
    "messages": [{ "role": "user", "content": "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
    "sogni_tools": "creative-agent",
    "sogni_tool_execution": true
  }'

Creative Agent accepts the short brief above and prepares MiniMax H3's structured production prompt. Calling a worker directly? Follow the MiniMax H3 integration guide and use one of these exact model IDs: minimax-h3-fl2va-fp8_t2v, minimax-h3-fl2va-fp8_i2v, minimax-h3-fl2va-fp8_flf2v, minimax-h3-ref2va-fp8_r2v, minimax-h3-fastvideo-int8_t2v_turbo, minimax-h3-fastvideo-int8_i2v_turbo, minimax-h3-fastvideo-int8_flf2v_turbo, minimax-h3-fl2va-fp8_t2v_turbo, minimax-h3-fl2va-fp8_i2v_turbo, minimax-h3-fl2va-fp8_flf2v_turbo, minimax-h3-ref2va-fp8_r2v_turbo. Full reference at docs.sogni.ai.

Why run it on Sogni

Subscriptions or Spark

Use a flat monthly plan for credit-free fair-use generation, or buy Spark packs when pay-as-you-go fits better. Both run on the same creator-owned GPU network.

Unlimited plans

One flat price in the app. Generate under fair use without a per-image meter.

🧩

200+ models

Image, video, music, and language models in one workspace and one API key.

Pay-as-you-go Spark

Prefer pay-as-you-go? Call MiniMax H3 by id and pay with Spark packs.

🌐

Powered by people

Runs on a decentralized GPU network where workers share subscription revenue.

FAQ

MiniMax H3 on Sogni

Is MiniMax H3 an uncensored alternative to Seedance 2.0?

Yes — if you are looking for uncensored or NSFW AI video with short-form references, dialogue, built-in sound, and an explicit mature-theme opt-in, MiniMax H3 is a strong alternative. H3 is its own MiniMax model, not a Seedance derivative or modified checkpoint. All creation remains subject to Sogni's Terms of Use and MiniMax's rules. Compare Sogni's opt-in mature-theme video options.

What is MiniMax H3 best for?

H3 is a strong fit for short brand films, dialogue scenes, product reveals, fashion clips, animated posters and title cards, motion design, interface animation, music visuals, and stylized social video where sound matters as much as motion.

Which three MiniMax H3 video LoRAs are available on Sogni?

Realism People improves faces, skin texture, hands, lighting, and natural human motion across Standard and Turbo text, image, first-and-last-frame, and reference video. VBVR Video Reasoning improves literal prompt adherence for Standard and Turbo text-to-video and image-to-video. Mystic X v4 adds opt-in mature-theme knowledge for Standard and Turbo text, image, and first-and-last-frame video. All three also work with FastH3 Turbo. Balanced beta LoRA support is still being evaluated.

How do I enable opt-in mature or NSFW MiniMax H3 generation?

At app.sogni.ai, open your username menu, switch off Sensitive Content Filter under Content preferences, and confirm that you are 18 or older. This reveals 18+ models and the VBVR Video Reasoning and Mystic X v4 LoRAs. Mature-theme generation is strictly opt-in and remains subject to Sogni's Terms of Use and each model's license.

What is FastH3 Turbo, and how much faster is it?

FastH3 Turbo is Sogni's fastest MiniMax H3 path and the default Turbo engine in Sogni Web. It runs the FastVideo FastH3 4-step Preview v1 VSA DataFree checkpoint, a four-step distillation of MiniMax H3 from the FastVideo team at Hao AI Lab, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, quick iterations, and most social content, and choose Standard H3 when you want maximum fine detail and audio polish. The familiar LightX2V Turbo stays one switch away.

What is FastVideo FastH3 4-step Preview v1 VSA DataFree?

It is the FastH3 checkpoint published by the FastVideo team at Hao AI Lab: a preview distillation of MiniMax H3 that generates synchronized video and audio in four transformer passes instead of twenty. VSA is video sparse attention, which skips most attention work at 90% sparsity, and DataFree means the distillation ran without an external training dataset. Sogni runs Kijai's INT8 ConvRot conversion of this checkpoint as FastH3 Turbo for text-to-video, image-to-video, and first-and-last-frame video. Shorter clips see a smaller speed-up than the up to 6× measured on 15-second clips.

Does MiniMax H3 support 2K resolution on Sogni?

No. Sogni offers 480p plus 544/768p-class output from MiniMax H3's published open weights. MiniMax's 2K stage is hosted-only and is not part of the open release, so Sogni does not advertise or promise 2K output from this model.

How long does a 15-second MiniMax H3 video take to render?

For a 15-second 768p-class clip, use Standard as the 1× baseline, Balanced at 2×, LightX2V Turbo at 4×, and FastH3 Turbo at 6×. If Standard takes 15 minutes, that works out to about 7.5 minutes on Balanced, 3.75 minutes on LightX2V Turbo, or 2.5 minutes on FastH3 Turbo once rendering starts. These are relative render-time examples, not guarantees; actual time varies by workflow, input, worker hardware, cache state, and queue demand. Pay-as-you-go jobs get top queue priority, followed by Unlimited Pro and Unlimited.

Does MiniMax H3 generate audio?

Yes. H3 creates the video, dialogue, music, ambience, and sound effects together, so your clip can come out ready to watch and share.

Which MiniMax H3 generation modes does Sogni support?

H3 supports text-only T2VA, opening-frame I2VA, closing-frame L2VA, two-endpoint FL2VA, and loose-reference Ref2VA. L2VA reuses the I2V model with a closing-frame role rather than inventing a separate model ID. FastH3 Turbo covers T2VA, I2VA, L2VA, and FL2VA; Ref2VA stays on Standard, Balanced, and LightX2V Turbo. The API exposes every endpoint shape with exact model IDs for Standard, Balanced, LightX2V Turbo, and FastH3 Turbo, and the Creative Agent Skill offers Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors (minimax-h3-fasth3-turbo).

How long are MiniMax H3 videos, and which layouts can I use?

Create clips from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts. Developers can also choose custom sizes through the API.

How much does MiniMax H3 cost on Sogni?

FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output, so an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, regardless of output speed. Sogni shows the current live quote before rendering. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.

How many full-quality 15-second H3 videos can I generate per day on Unlimited?

There is no fixed video count. Video draws on the same daily fair-use capacity as other covered models, which resets every 24 hours on your own plan schedule; Relaxed rendering has no daily capacity, and Premium Spark is unaffected. Unlimited lets you generate one Standard H3 video at a time, while Unlimited Pro lets you generate two and carries 4x the daily capacity. You can keep adding videos to your queue throughout the day, subject to fair use, and see your current standing as a percentage at app.sogni.ai/usage.

What happens if I use MiniMax H3 heavily on Unlimited?

If your daily use is much heavier than normal, the number of videos you can generate at once may be reduced until the next UTC day. Turbo is less affected than Standard, and more than 90% of subscribers do not currently hit a fair-use slowdown. Speeds and plan terms are subject to change as Sogni balances supply and demand to keep renders fast for artists and rewarding for the people who share their GPUs through our people-powered render network. Sogni Unlimited is the best deal in town for frequent creators, and we plan to keep it that way.

Can I continuously queue MiniMax H3 generations throughout the day?

Yes. Individual creators can keep a substantial queue moving throughout the day, including through the API and Creative Agent. Unlimited is not intended to power unattended 24/7 or multi-user production systems; use pay-as-you-go Spark or an Enterprise plan for those workloads. Read the Creative Agent Skill guide.

Does Sogni's H3 release support image, video, or audio references?

Yes. Standard and Turbo Ref2VA let you guide a clip with up to nine images, three videos, and three audio clips — 12 files total, with at least one image or video. Video and audio references must each be 2–15 seconds; video references may total at most 15 seconds, and audio references have a separate 15-second total.

Can I use MiniMax H3 in Sogni Web, the API, and Creative Agent?

Yes. Sogni Web, the API, and Creative Agent support text-to-video, image-to-video, first-and-last-frame video, and multi-reference video across Standard and Turbo. Pay as you go with Spark or create on an Unlimited plan.

Is MiniMax H3 licensed for use on Sogni?

Yes. MiniMax has given Sogni written authorization to offer H3 in the United States, European Union, United Kingdom, and South Korea. The MiniMax H3 Community License applies to people who self-host the published model, not to videos you generate on Sogni.

How should I structure a MiniMax H3 prompt?

Leave Sogni's AI Script Writer on for a natural-language brief, or follow the full prompting guide on this page. Base modes use three ordered fields; Ref2VA uses six. Shot 1 has no timestamp, later cuts use [Shot N] At MM:SS.mmm, vocal sources keep stable (S1) IDs, and dialogue uses <d>[Language] exact words</d>. Preserve supplied words exactly; only author dialogue when the user explicitly asks for speech without providing a line.

What languages can MiniMax H3 dialogue use?

MiniMax reports stable dialogue support for Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish, with additional languages supported to varying degrees.

How private is my use of this model on Sogni?

Sogni is built with privacy and creative freedom in mind. Your work remains your own, and inference runs through the Sogni Supernet, a decentralized network of creator GPUs, instead of requiring local hardware or a separate model host. Use is still governed by Sogni's Privacy Policy and Terms of Use.

Start with MiniMax H3 today

Create in the app, or build with the API. Your call.