Create polished 5- to 15-second videos with dialogue, music, and sound in one generation. Choose FastH3 Turbo, up to 6× faster than Standard, for everyday creation, or Standard H3 for maximum detail.
Create a full 15-second video up to 6× faster with FastH3 Turbo, or choose Standard for maximum polish.
MiniMax H3 creates complete short-form videos with picture and sound in one generation. Direct the action, camera, dialogue, ambience, effects, and music in the same prompt — no separate soundtrack pass required. It is a strong choice for dialogue scenes, brand films, product reveals, fashion clips, motion design, music visuals, and social content.
Start from a text prompt, animate an opening image, guide a transition with opening and closing frames, or use image, video, and audio references to shape the result. Choose the workflow that matches how much creative control you want.
Start with FastH3 Turbo for everyday creation. Built on FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, it renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. It is ideal for drafts, timing tests, quick iteration, and most social content, and it works with the same H3 LoRAs. Choose Standard H3 when you want the very best fine detail and audio polish. The familiar LightX2V Turbo stays one switch away, and multi-reference video is available in Standard and LightX2V Turbo.
Create videos from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts, with dialogue, music, and sound generated together. Sogni offers 480p plus 544/768p-class output from the open-weights release; MiniMax's hosted 2K stage is not part of the published weights. MiniMax reports stable dialogue support in 11 languages, with additional languages supported to varying degrees.
Sogni Web makes FastH3 the default Turbo engine for text-to-video, image-to-video, and first-and-last-frame video, with a switch back to LightX2V Turbo. The Sogni API accepts an exact model ID for every Standard, Balanced, LightX2V Turbo, and FastH3 Turbo workflow, and the Creative Agent Skill offers Standard, LightX2V Turbo, and FastH3 Turbo selectors. Multi-reference video accepts up to nine images, three videos, and three audio clips in Standard and LightX2V Turbo; FastH3 does not offer a multi-reference mode.
Choose FastH3 Turbo for fast everyday creation — up to 6× faster than Standard — LightX2V Turbo for its familiar four-step look, or Standard when you want maximum detail and audio polish. Multi-reference video is available in Standard and LightX2V Turbo.
| Workflow | Details | Access | |
|---|---|---|---|
| Text to video Standard | Workflow: Create video and audio from a written scene Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: No image needed Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| First or last frame Standard | Workflow: Animate from an opening image or converge on a closing image Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening or closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| First and last frames Standard | Workflow: Direct the motion between two compositions Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame and one closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | API / agent workflow |
| Multi-reference Standard | Workflow: Condition a scene on labelled image, video, and audio references Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| FastH3 Turbo text to video FastH3 Turbo | Workflow: Create video and audio from text with the FastH3 engine Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: No image needed Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| FastH3 Turbo first or last frame FastH3 Turbo | Workflow: Animate from an opening image or converge on a closing image with the FastH3 engine Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening or closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| FastH3 Turbo first and last frames FastH3 Turbo | Workflow: Connect two anchor images with the FastH3 engine Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame and one closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| Turbo text to video Turbo | Workflow: Create video and audio from text with the Turbo path Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: No image needed Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| Turbo first or last frame Turbo | Workflow: Animate from an opening image or converge on a closing image with Turbo Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening or closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| Turbo first and last frames Turbo | Workflow: Connect two anchor images with the Turbo path Duration: Roughly 5–15 seconds Resolution: 576–1344 px · fixed 24 fps References: One opening frame and one closing frame Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
| Turbo multi-reference Turbo | Workflow: Condition a scene on labelled image, video, and audio references with the Turbo path Duration: Roughly 5–15 seconds Resolution: 544–1344 px · fixed 24 fps References: 0–9 images · up to 3 videos · up to 3 audio clips · 12 files total Audio: 32 kHz stereo · video at 24 fps | Spark or eligible Unlimited | Create → |
Open the H3 LoRAs panel to add realism, tighter prompt adherence, or opt-in mature-theme knowledge without changing the underlying MiniMax H3 workflow.
VBVR Video Reasoning and Mystic X v4 appear after the Sensitive Content Filter is turned off. All three also work with FastH3 Turbo. Balanced beta LoRA support is still being evaluated. For a discreet overview of supported mature-theme workflows, see Sogni's uncensored AI video generator.
We recommend leaving Sogni's AI Script Writer on. It is enabled by default in many Sogni interfaces and automagically turns a simple idea into a rich script tailored to H3. You can still add shot, camera, style, dialogue, and sound details whenever you want more control.
For everyday creation, leave Sogni's AI Script Writer on and write a natural brief. It expands the idea into MiniMax's model-native structure. Advanced callers can write that structure directly using the guide below; small formatting imperfections are worth correcting, but they should not replace the creative substance of a useful prompt.
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ... [Shot 1] has no timestamp. Every later cut uses [Shot N] At MM:SS.mmm, ..., with contiguous shot numbers and strictly increasing times inside the rendered duration. Describe visible action, camera, dialogue, singing, diegetic music, and synchronized events in the main description. Put ambience, effects, and non-verbal human sound in overall_soundscape. Put only audience-only score in non_diegetic_music, or use N/A when there is no score.
Give each vocal source a stable ID such as (S1). Keep identity, action, and delivery outside the dialogue tag; put only the language and words inside <d>[English] Exact words.</d>. Preserve user-supplied dialogue character-for-character. If the request explicitly asks for speaking, dialogue, lyrics, or a vocal performance without supplying words, author one concise line; with no vocal intent, invent no speech. Use <scenetrans> at both connection points when one line crosses a cut and plain <cutoff> only when the video ending truncates speech. Never write tokenizer-internal <|...|> controls in a prompt, and never fix one by merely removing its pipe characters: caption markers become exact visible text in double quotes, lyrics markers become an ordinary <d>[Language] exact words</d> singing block, and <|cutoff|> becomes plain <cutoff>. Plain caption or lyrics boundary tags are not valid substitutes.
I2VA begins with the exact opening-frame alignment sentence. L2VA and FL2VA begin with their duration-aware alignment sentence naming the actual final shot and snapped endpoint. T2VA and Ref2VA use no alignment sentence. The Script Writer handles these lines automatically when it knows which frames are attached.
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music: Keep <Subject N>, <Picture N>, <Video N>, and <Audio N> meanings stable. The summary starts with task types chosen from keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference; join multiple types with exactly + and never repeat one. Attached clips are loose references unless the runtime explicitly provides an edit or continuation relationship, so do not promise source-video transformation from file presence alone.
Use fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visual retention; use fully_copy, partially_copy, reference, or weak_reference for audio. Bind a reference voice to its subject's (Sx) in subject_definitions, not in retention_analysis. A timbre, rhythm, emotion, or delivery reference does not authorize copying its words. Preserve explicitly reused speech and mark unintelligible spans [unclear] instead of guessing.
Ref2VA requires at least one image or video. It accepts up to nine images, three videos, three standalone audio clips, and 12 files total. Each reference video and audio clip must be 2–15 seconds; reference videos may total at most 15 seconds, and reference audio may separately total at most 15 seconds.
H3 renders at fixed 24 fps on a 124 + n×17 frame grid, about 5.17–15.08 seconds, with a 7,000-character prompt limit. Preserve required fields, exact user text, reference jobs, retention markers, and shot timing before trimming redundant adjectives or audio prose. Sogni offers 480p plus 544/768p-class output from the open-weights release; do not prompt for or promise the hosted-only 2K stage.
Source: MiniMax's official H3 prompt-writing skill, pinned to the reviewed open-weights revision.
Measured on one RTX 5090: the same 15-second 768p image-to-video prompt rendered in 2 min 6 s on FastH3 Turbo and 11 min 55 s on Standard, with Balanced and LightX2V Turbo in between. Every clip is embedded in FastH3 Speed Test: Four MiniMax H3 Tiers, One RTX 5090, One Cat in a Bath.
See Standard and Turbo side by side in the Sogni Engineering field guide, Your Prompt Is Now a Director.
FastH3 Turbo is our go-to for fast everyday creation. It runs FastVideo FastH3 4-step Preview v1 VSA DataFree, the FastVideo team's four-step distillation of MiniMax H3, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, timing tests, quick iterations, and most social content. Choose Standard H3 when maximum fine detail and audio polish matter most.
FastH3 is the default Turbo engine in Sogni Web for text, image, and first-and-last-frame video, with a switch back to the familiar LightX2V Turbo look. It works with the H3 LoRAs, costs 4 Spark ($0.02) per output second, and has no multi-reference mode. Shorter clips see a smaller speed-up than 15-second clips.
For 15-second 768p-class video, use Standard as the 1× baseline. Balanced runs the MiniMax H3 LightX2V 8-step recipe, Turbo runs the LightX2V 4-step recipe, and FastH3 Turbo runs FastVideo FastH3 4-step.
| Sogni speed | Engine | Relative speed | Illustrative 15-second render | 544/768p price |
|---|---|---|---|---|
| Standard | MiniMax H3 · 20 steps | 1× baseline | 15 min | 16 Spark ($0.08) per output second |
| Balanced | LightX2V · 8 steps | 2× | 7.5 min | 10 Spark ($0.05) per output second |
| Turbo | LightX2V · 4 steps | 4× | 3.75 min | 6 Spark ($0.03) per output second |
| FastH3 Turbo | FastVideo FastH3 · 4 steps | 6× | 2.5 min | 4 Spark ($0.02) per output second |
The times show relative render speed: if Standard takes 15 minutes, the faster tiers scale from that baseline. They exclude queue time and are not guarantees; workflow, input, worker hardware, and cache state can change results. For an illustrative 15.0-second cost comparison, 544/768p-class output is 240 Spark ($1.20) on Standard, 150 Spark ($0.75) on Balanced, 90 Spark ($0.45) on Turbo, or 60 Spark ($0.30) on FastH3 Turbo; actual billing follows the rendered frame duration, up to 15.08 seconds. At 480p, Standard costs 10 Spark ($0.05), Balanced 6 Spark ($0.03), and both Turbo engines 4 Spark ($0.02) per output second.
Measured, not illustrative. On a single RTX 5090 worker, the same 15.08-second 768 × 1024 image-to-video prompt rendered hot in 11 min 55 s on Standard, 5 min 30 s on Balanced, 3 min 0 s on LightX2V Turbo, and 2 min 6 s on FastH3 Turbo, with 3 min 39 s, 1 min 46 s, 1 min 7 s, and 55 s at 480p. Every clip is embedded so you can judge the quality trade yourself: read the FastH3 speed test.
MiniMax H3 creates video and sound together. Direct dialogue, ambience, effects, and music in the same brief, then get a complete clip without a separate soundtrack pass.
Start from text, an opening frame, a closing frame, both endpoints, or a loose labelled reference set. Standard and Turbo Ref2VA can combine up to nine images, three videos, and three audio clips.
MiniMax H3 is its own MiniMax model, not a Seedance derivative. Sogni offers 480p plus 544/768p-class output from the published weights; MiniMax's hosted 2K stage is not part of the open release. Creators 18 or older can opt into mature-content creation on Sogni, subject to Sogni's Terms of Use and the applicable MiniMax, LightX2V, and FastVideo licenses.
Sogni Web offers T2VA, I2VA, L2VA, and FL2VA in Standard, Balanced, LightX2V Turbo, and FastH3 Turbo, plus Ref2VA in Standard, Balanced, and LightX2V Turbo. The Sogni API accepts every workflow by exact model ID, and the Creative Agent Skill exposes Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors.
FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output; an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, even with Turbo output. Sogni shows the current quote before rendering.
Review the official MiniMax H3 model card, MiniMax H3 Community License, official MiniMax H3 API pricing, LightX2V H3 Turbo model card, and LightX2V Turbo implementation, FastVideo FastH3 4-step Preview v1 VSA DataFree model card, FastVideo implementation, Hao AI Lab FastH3 preview post, and Kijai INT8 ConvRot conversion used by Sogni. Sogni has received written authorization from MiniMax to offer H3 through the platform.
FastH3 Turbo, built on FastVideo FastH3 4-step Preview v1, is Sogni's fastest H3 path — up to 6× faster than Standard — and costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output. An exact 8-second, 192-frame FastH3 clip costs 32 Spark ($0.16). Balanced uses the qualified 8-step LightX2V recipe for text and frame-guided video and is about 2× faster than 20-step Standard; quality is still being evaluated. At 480p, Standard output costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at $0.05 per second for 480p or $0.08 for 544/768p, regardless of output tier. Standard and Balanced multi-reference jobs can also include the reference-image charge shown in the live quote. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.
Start with a natural creative brief. Creative Agent expands it into the production prompt MiniMax H3 needs and runs the matching workflow.
const response = await fetch('https://api.sogni.ai/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.SOGNI_API_KEY}`,
},
body: JSON.stringify({
messages: [{ role: 'user', content: "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
sogni_tools: 'creative-agent',
sogni_tool_execution: true,
}),
});
if (!response.ok) throw new Error(await response.text());
const { choices } = await response.json();
console.log(choices[0].message.content); curl https://api.sogni.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $SOGNI_API_KEY" \
-d '{
"messages": [{ "role": "user", "content": "Create an 8-second MiniMax H3 video from this brief: A ceramic artist opens a glowing kiln as the camera slowly pushes in; fire crackles, tools clink softly, and a restrained string score begins" }],
"sogni_tools": "creative-agent",
"sogni_tool_execution": true
}' Creative Agent accepts the short brief above and prepares MiniMax H3's structured production prompt. Calling a worker directly? Follow the MiniMax H3 integration guide and use one of these exact model IDs: minimax-h3-fl2va-fp8_t2v, minimax-h3-fl2va-fp8_i2v, minimax-h3-fl2va-fp8_flf2v, minimax-h3-ref2va-fp8_r2v, minimax-h3-fastvideo-int8_t2v_turbo, minimax-h3-fastvideo-int8_i2v_turbo, minimax-h3-fastvideo-int8_flf2v_turbo, minimax-h3-fl2va-fp8_t2v_turbo, minimax-h3-fl2va-fp8_i2v_turbo, minimax-h3-fl2va-fp8_flf2v_turbo, minimax-h3-ref2va-fp8_r2v_turbo. Full reference at docs.sogni.ai.
Real generations from Sogni. Hit Use prompt to open the app with MiniMax H3 selected and the full prompt preloaded.
Use a flat monthly plan for credit-free fair-use generation, or buy Spark packs when pay-as-you-go fits better. Both run on the same creator-owned GPU network.
One flat price in the app. Generate under fair use without a per-image meter.
Image, video, music, and language models in one workspace and one API key.
Prefer pay-as-you-go? Call MiniMax H3 by id and pay with Spark packs.
Runs on a decentralized GPU network where workers share subscription revenue.
Yes — if you are looking for uncensored or NSFW AI video with short-form references, dialogue, built-in sound, and an explicit mature-theme opt-in, MiniMax H3 is a strong alternative. H3 is its own MiniMax model, not a Seedance derivative or modified checkpoint. All creation remains subject to Sogni's Terms of Use and MiniMax's rules. Compare Sogni's opt-in mature-theme video options.
H3 is a strong fit for short brand films, dialogue scenes, product reveals, fashion clips, animated posters and title cards, motion design, interface animation, music visuals, and stylized social video where sound matters as much as motion.
Realism People improves faces, skin texture, hands, lighting, and natural human motion across Standard and Turbo text, image, first-and-last-frame, and reference video. VBVR Video Reasoning improves literal prompt adherence for Standard and Turbo text-to-video and image-to-video. Mystic X v4 adds opt-in mature-theme knowledge for Standard and Turbo text, image, and first-and-last-frame video. All three also work with FastH3 Turbo. Balanced beta LoRA support is still being evaluated.
At app.sogni.ai, open your username menu, switch off Sensitive Content Filter under Content preferences, and confirm that you are 18 or older. This reveals 18+ models and the VBVR Video Reasoning and Mystic X v4 LoRAs. Mature-theme generation is strictly opt-in and remains subject to Sogni's Terms of Use and each model's license.
FastH3 Turbo is Sogni's fastest MiniMax H3 path and the default Turbo engine in Sogni Web. It runs the FastVideo FastH3 4-step Preview v1 VSA DataFree checkpoint, a four-step distillation of MiniMax H3 from the FastVideo team at Hao AI Lab, and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard 768p H3 on a full 15-second clip. Use it for drafts, quick iterations, and most social content, and choose Standard H3 when you want maximum fine detail and audio polish. The familiar LightX2V Turbo stays one switch away.
It is the FastH3 checkpoint published by the FastVideo team at Hao AI Lab: a preview distillation of MiniMax H3 that generates synchronized video and audio in four transformer passes instead of twenty. VSA is video sparse attention, which skips most attention work at 90% sparsity, and DataFree means the distillation ran without an external training dataset. Sogni runs Kijai's INT8 ConvRot conversion of this checkpoint as FastH3 Turbo for text-to-video, image-to-video, and first-and-last-frame video. Shorter clips see a smaller speed-up than the up to 6× measured on 15-second clips.
No. Sogni offers 480p plus 544/768p-class output from MiniMax H3's published open weights. MiniMax's 2K stage is hosted-only and is not part of the open release, so Sogni does not advertise or promise 2K output from this model.
For a 15-second 768p-class clip, use Standard as the 1× baseline, Balanced at 2×, LightX2V Turbo at 4×, and FastH3 Turbo at 6×. If Standard takes 15 minutes, that works out to about 7.5 minutes on Balanced, 3.75 minutes on LightX2V Turbo, or 2.5 minutes on FastH3 Turbo once rendering starts. These are relative render-time examples, not guarantees; actual time varies by workflow, input, worker hardware, cache state, and queue demand. Pay-as-you-go jobs get top queue priority, followed by Unlimited Pro and Unlimited.
Yes. H3 creates the video, dialogue, music, ambience, and sound effects together, so your clip can come out ready to watch and share.
H3 supports text-only T2VA, opening-frame I2VA, closing-frame L2VA, two-endpoint FL2VA, and loose-reference Ref2VA. L2VA reuses the I2V model with a closing-frame role rather than inventing a separate model ID. FastH3 Turbo covers T2VA, I2VA, L2VA, and FL2VA; Ref2VA stays on Standard, Balanced, and LightX2V Turbo. The API exposes every endpoint shape with exact model IDs for Standard, Balanced, LightX2V Turbo, and FastH3 Turbo, and the Creative Agent Skill offers Standard, Balanced, LightX2V Turbo, and FastH3 Turbo selectors (minimax-h3-fasth3-turbo).
Create clips from about 5 to 15 seconds in square, landscape, cinema, or portrait layouts. Developers can also choose custom sizes through the API.
FastH3 costs 4 Spark ($0.02) per output second at both 480p and 544/768p-class output, so an exact 8-second, 192-frame clip costs 32 Spark ($0.16). At 480p, Standard H3 costs $0.05 per second, Balanced costs $0.03, and LightX2V Turbo costs $0.02. At 544/768p, Standard costs $0.08 per second, Balanced costs $0.05, and LightX2V Turbo costs $0.03. Ref2VA reference-video input is billed by exact duration at the full resolution rate: $0.05 per second at 480p or $0.08 at 544/768p, regardless of output speed. Sogni shows the current live quote before rendering. Pay as you go with Spark, or create with any tier on an Unlimited plan, subject to fair use.
There is no fixed video count. Video draws on the same daily fair-use capacity as other covered models, which resets every 24 hours on your own plan schedule; Relaxed rendering has no daily capacity, and Premium Spark is unaffected. Unlimited lets you generate one Standard H3 video at a time, while Unlimited Pro lets you generate two and carries 4x the daily capacity. You can keep adding videos to your queue throughout the day, subject to fair use, and see your current standing as a percentage at app.sogni.ai/usage.
If your daily use is much heavier than normal, the number of videos you can generate at once may be reduced until the next UTC day. Turbo is less affected than Standard, and more than 90% of subscribers do not currently hit a fair-use slowdown. Speeds and plan terms are subject to change as Sogni balances supply and demand to keep renders fast for artists and rewarding for the people who share their GPUs through our people-powered render network. Sogni Unlimited is the best deal in town for frequent creators, and we plan to keep it that way.
Yes. Individual creators can keep a substantial queue moving throughout the day, including through the API and Creative Agent. Unlimited is not intended to power unattended 24/7 or multi-user production systems; use pay-as-you-go Spark or an Enterprise plan for those workloads. Read the Creative Agent Skill guide.
Yes. Standard and Turbo Ref2VA let you guide a clip with up to nine images, three videos, and three audio clips — 12 files total, with at least one image or video. Video and audio references must each be 2–15 seconds; video references may total at most 15 seconds, and audio references have a separate 15-second total.
Yes. Sogni Web, the API, and Creative Agent support text-to-video, image-to-video, first-and-last-frame video, and multi-reference video across Standard and Turbo. Pay as you go with Spark or create on an Unlimited plan.
Yes. MiniMax has given Sogni written authorization to offer H3 in the United States, European Union, United Kingdom, and South Korea. The MiniMax H3 Community License applies to people who self-host the published model, not to videos you generate on Sogni.
Leave Sogni's AI Script Writer on for a natural-language brief, or follow the full prompting guide on this page. Base modes use three ordered fields; Ref2VA uses six. Shot 1 has no timestamp, later cuts use [Shot N] At MM:SS.mmm, vocal sources keep stable (S1) IDs, and dialogue uses <d>[Language] exact words</d>. Preserve supplied words exactly; only author dialogue when the user explicitly asks for speech without providing a line.
MiniMax reports stable dialogue support for Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish, with additional languages supported to varying degrees.
Sogni is built with privacy and creative freedom in mind. Your work remains your own, and inference runs through the Sogni Supernet, a decentralized network of creator GPUs, instead of requiring local hardware or a separate model host. Use is still governed by Sogni's Privacy Policy and Terms of Use.
Create in the app, or build with the API. Your call.