01
Text to Video
Describe the subject, action, camera, lighting, pacing, and sound. MiniMax H3 builds the scene without a source asset.
Multimodal video generation with MiniMax H3
Create 4–15 second cinematic videos from text, images, video, and audio—with native stereo sound and output up to 2K.
Generated Result
One model, one creative brief
MiniMax H3 is a multimodal video model from MiniMax. This AI video generator reads text, images, video, and audio in one creative brief, then applies each reference to the subject, motion, camera, or sound.
The model supports text-to-video, reference-based generation, and editing. Create 4–15 second clips with native stereo sound, then use H3Art to manage prompts, references, and output settings without an API workflow.
Choose the right starting point
Start from a prompt, a frame, mixed references, or an existing clip. Choose the workflow that matches your source material and control needs.
01
Describe the subject, action, camera, lighting, pacing, and sound. MiniMax H3 builds the scene without a source asset.
02
Animate a first frame or guide both ends of the shot with first and last frames. Preserve composition while adding motion and audio.
03
Assign separate images, video, and audio to the subject, camera motion, voice, or rhythm within one request.
04
Upload a source clip and use natural-language instructions to change its motion, style, content, or sound.
Multimodal direction
Add only the references that control the result. A first frame holds composition, images establish identity, video supplies movement, and audio supplies voice or rhythm.
Name each file in the prompt and assign its role. MiniMax H3 can combine a character image, a camera reference, and a voice sample in one output.
Reference Video

Voice, sound, or rhythm
Generated Video
Motion and camera direction
Character, product, or visual identity
Voice, sound, or rhythm
A unified result from MiniMax H3
Picture and sound, generated together
MiniMax H3 generates stereo speech, effects, ambience, and music with the picture. A spoken line can guide mouth movement and pacing.
Output duration runs from 4 to 15 seconds, giving short drafts and finished shots the same multimodal control.
Native audio
Listen for speech, effects, ambience, and music timed to the MiniMax H3 shot.
Source fidelity
Inspect product, interface, and environmental detail preserved from the source context.
Representative outputs
Each example shows a different input strategy and result to inspect.
01 / Titles
Use title frames and a camera brief to hold the typography, timing, and mood. Add audio for synchronized visual beats.
02 / Product
Turn a product image or interface concept into a moving page while a reference clip supplies the scroll rhythm.
03 / Campaign
Keep the poster composition while adding depth, selective movement, and sound for a short campaign loop.
04 / Commerce
Create short product shots from existing assets while your prompt controls labels, camera direction, and pacing.
Four steps from brief to clip
Write the subject, action, camera direction, and sound. MiniMax H3 also accepts text-only requests.
Upload first and last frames or a mixed media set for identity, motion, style, voice, or rhythm.
Mention each file in the prompt. Assign the image to identity, the clip to motion, and the audio to voice or rhythm.
Set aspect ratio, duration, resolution, and audio. Generate, review the result, then revise the prompt or references.
Model architecture
Contextual Omni Representation describes how each input should affect the output, helping MiniMax H3 follow instructions across media types.
H3-VAE compresses source information while preserving detail. MiniMax reports four times the effective sequence length, supporting native 2K generation.
H3-Omni Transformer separates understanding and generation workloads. MiniMax reports nearly 30 percent higher training throughput, while in-context regeneration restores fine detail from the original references.
Plans and credits
Use a monthly plan for a steady production rhythm or choose a credit pack for individual projects.
220 credits/month
1090 credits/month
2650 credits/month
4800 credits/month
Questions, answered in full
Yes. It can generate stereo speech, effects, ambience, and music. Turn audio on or off per task.
Upload one or more images to guide the subject, composition, or visual style. You can also use first and last frames when you want to control how the shot begins and ends. MiniMax H3 adds motion and can generate stereo audio from the prompt and references.
No. H3Art connects to the model, so you do not supply your own MiniMax API key.
Generation uses H3Art credits. Duration, resolution, and reference-video length affect the estimate shown before submission.
Yes. MiniMax positions H3 for commercial content creation, and H3Art does not add its own watermark. You are responsible for the rights to uploaded references and for following applicable terms and laws.
No. MiniMax develops the model; H3Art provides an independent creation interface.
Start from one creative brief
Add your prompt and references, direct the motion and sound, then generate a 4–15 second clip with MiniMax H3.
Create with MiniMax H3