Multimodal AI video generator

MiniMax H3 AI Video Generator

MiniMax H3 brings prompts, images, motion references, and audio direction into one creative brief. Plan a video up to 15 seconds, direct native stereo sound, and work toward a detailed 2K result without splitting every creative task into a different tool.

Create MiniMax H3 Video
Reference Images(max 9)
Add Image
Reference Videos(max 3)
0s / 15s
Add Video
Reference Audios(max 3)
0s / 15s
Add Audio
Reference assets are addressed as Image 1…N, Video 1…N, Audio 1…N.0 / 7000
MiniMax H3 AI Video Result
Generation takes about 5 min. Please don't close this tab.

Your MiniMax H3 result will appear here

Choose a mode, complete the required inputs, and generate a native 2K video up to 15 seconds long.

Fill in the form on the left and click "Generate" to create your own video.

Download links expire in 24h — please save your videos promptly.

View My Video History →

Multi Reference Guide

MiniMax H3 accepts an image, video, or audio track on its own, or up to 12 references across all three types.

How to Reference Files in Your Prompt

Address references in fixed per-type order: Image 1…N, Video 1…N, then Audio 1…N.

Example: Keep the character from Image 1 and follow the camera movement from Video 1.

Upload Limits

  • Images — up to 9 files, JPG / JPEG / PNG / WebP / HEIC / HEIF, max 30 MB each, 256–5760px, aspect ratio 0.4–2.5
  • The first 5 reference images are included; Images 6–9 add 60 credits each
  • Videos — up to 3 files, MP4 / MOV, 2–15s each, combined max 15s, max 50 MB each, 256–5760px, ratio 0.4–2.5, 23.976–60 fps
  • Audio — up to 3 files, WAV / MP3, 2–15s each, combined max 15s, max 15 MB each
  • Total materials across all types: 12 max

Accepted Input Combinations

  • Text + Image only
  • Text + Video only
  • Text + Audio only
  • Text + any mix of images, videos, and audio

Billing: 200 credits × (output seconds + reference video seconds). The first 5 images are included; Images 6–9 add 60 credits each.

The live generator routes all three modes to MiniMax H3 through EvoLink. Mode-specific fields, upload limits, and billing come from the EvoLink API specification; the model capabilities below are based on MiniMax's official July 31, 2026 announcement.

One creative context

What MiniMax H3 changes in video production

Traditional generators separate text-to-video, image animation, subject reference, motion transfer, voice guidance, and editing. MiniMax H3 is designed as a general-purpose model that learns those relationships together. The practical benefit is a clearer workflow: collect the evidence for the shot, explain how the pieces interact, generate, inspect, and revise.

One context for every reference

MiniMax H3 can interpret text, images, video, and audio together. Instead of forcing every asset into a separate tool, you can explain what each reference should control and describe the relationship between the sources and the final shot.

Native 2K video output

MiniMax H3 offers 2K output by default, according to MiniMax. The model regenerates its own lower-resolution result with the original context available again, which is designed to recover small details rather than simply enlarge pixels.

Stereo sound generated with the video

MiniMax H3 jointly models picture and sound. Dialogue, environmental effects, music, and the visual action can be planned as one result, while every generated audio output is described by MiniMax as native stereo.

Reference and editing in natural language

It treats reference and editing as general instructions instead of a long list of isolated modes. Tell the model which identity, movement, voice, layout, or detail should remain, then describe what needs to change.

Input role map

Give every reference one job

The best MiniMax H3 prompt does more than list attractive adjectives. It identifies the target, assigns authority to each source, and resolves conflicts before generation. A reference should control identity, motion, sound, composition, or style—not vaguely influence everything at once.

Text

Define the target

Use the prompt to name the subject, action, setting, camera behavior, sound, sequence of events, and the job of every uploaded reference. MiniMax H3 uses language as the bridge between those inputs.

Images

Anchor identity and design

Upload a character, product, key visual, interface, poster, or style frame. Tell MiniMax H3 whether the image controls appearance, composition, typography, material, color, or another specific property.

Video

Transfer motion and camera language

A reference clip can communicate a gesture, performance, transition, camera move, edit rhythm, or physical interaction. MiniMax H3 is officially positioned for video-to-video motion transfer and multimodal editing.

Audio

Guide voice, rhythm, and atmosphere

Provide a vocal, sound, or music reference and state how it relates to the image and motion. It can use audio as context while generating an audiovisual result with synchronized stereo sound.

From brief to deliverable

MiniMax H3 use cases for creative teams

MiniMax names advertising, branding, e-commerce, product design, UI/UX, gaming, and other commercial workflows in its launch post. These are practical starting points for turning the model's multimodal context into a defined production task.

Advertising and e-commerce

Turn product images, brand rules, a motion reference, and a short offer into one campaign clip. MiniMax H3 is suited to product reveals, paid social hooks, storefront loops, and localized creative variations where labels and brand details matter.

Product and interface concepts

Move from a static mockup to a presentation-ready concept. MiniMax H3 can help teams animate product interactions, website sections, interface transitions, packaging details, and speculative design behavior before full production begins.

Titles, posters, and story beats

Build a film opening, animate a poster, or test a short multi-shot beat. MiniMax H3 supports native multi-shot modeling, making it useful when a concept needs an ordered visual progression instead of one disconnected motion loop.

Brand systems and campaign variants

Use approved art direction as context, then change the product, setting, language, or format. MiniMax H3 emphasizes accurate text and brand rendering, but every logo and claim should still pass a human review before publishing.

Motion and performance transfer

Reference an existing camera move or performance while replacing the subject and setting. MiniMax H3 can connect the motion from a video, the identity from an image, and the timing from an audio reference in one written instruction.

Targeted creative revisions

Describe what should change while keeping the accepted context available. The model is designed around generalized reference and editing, so revision requests can be expressed as creative intent rather than a rigid task preset.

A repeatable method

How to prepare a MiniMax H3 generation

More inputs do not automatically create more control. Use this four-step workflow to reduce conflicting direction and make each revision easier to diagnose.

01

Choose one production goal

Start with a single deliverable: a product reveal, animated key visual, interface teaser, motion transfer, or 15-second story beat. MiniMax H3 can handle broad context, but a clear target gives every reference a reason to exist.

02

Assign one role to each asset

Label what the image, video, and audio should control. If two sources disagree about identity, camera, wardrobe, pace, or lighting, decide which one wins before asking MiniMax H3 to combine them.

03

Write relationships, not a keyword pile

Explain the sequence in normal language: use the first image for the character, follow the camera path from the clip, match the vocal timing from the audio, and preserve the package text. MiniMax H3 is built to interpret these contextual relationships.

04

Review the result like a final draft

Watch at full speed, then inspect frames for identity, small text, hands, product geometry, audio timing, and continuity. A strong MiniMax H3 result can reduce production work, but it does not remove brand, rights, or accuracy checks.

Prompt starters

Write a MiniMax H3 brief that connects the inputs

These examples focus on relationships, protected details, timing, and reviewable outcomes. Replace the numbered assets and product facts with your own approved material.

Product launch spot

01

Use Image 1 as the exact product and packaging reference. Follow the slow orbital camera movement from Video 2. Place the product on a dark reflective surface with controlled studio highlights. Reveal the brand mark only after the camera settles. Match the final impact to Audio 3, preserve all package proportions, and generate a concise stereo sound bed without narration.

Performance transfer

02

Use the person in Image 1 as the lead character and keep their face, hair, and clothing consistent. Transfer only the body movement and handheld camera rhythm from Video 2. Set the performance in the location from Image 3. Align the character's singing to Audio 4, keep the voice reference recognizable, and preserve natural room ambience in stereo.

Animated interface concept

03

Treat Image 1 as the approved desktop interface and Image 2 as the mobile layout. Animate a cursor opening the product panel, selecting one item, and completing checkout. Keep every label legible, preserve spacing and brand colors, use restrained camera movement, add subtle interaction sounds, and finish on the unchanged logo lockup.

Confirmed output details

MiniMax H3 was announced on July 31, 2026 as a general-purpose multimodal generation model. The official post confirms video up to 15 seconds, default 2K output, native stereo sound, multimodal understanding, native multi-shot modeling, reference and editing, and video-to-video motion transfer.

MiniMax also reports that MiniMax H3 costs less than one-third as much per second at 2K as unnamed mainstream models, and that 768p costs less than half as much as those models at 720p. The post does not publish a dollar price or a complete comparison method, so treat those figures as provider claims rather than a checkout quote for this site.

Read the official announcement

Review before publishing

MiniMax H3 can reduce the distance between a creative brief and a finished-looking clip. It cannot decide whether your source material is licensed, whether a likeness is authorized, or whether a generated product claim is accurate.

Compare faces, wardrobe, products, and key props across every shot.
Pause on logos, labels, interface text, captions, and legal copy.
Listen for lip-sync drift, clipped speech, music conflicts, and abrupt ambience.
Check hands, contact points, reflections, shadows, and object geometry.
Confirm rights for every reference image, clip, voice, track, and trademark.

Clear answers

MiniMax H3 FAQ

Official capabilities, practical workflow notes, and the boundary between the model announcement and this site's live generator.

1

What is MiniMax H3?

MiniMax H3 is a general-purpose multimodal generation model released by MiniMax on July 31, 2026. It understands context across text, images, video, and audio, and it can generate video with native stereo sound for a maximum duration of 15 seconds at up to 2K resolution.

2

Can MiniMax H3 generate audio with video?

Yes. MiniMax says MiniMax H3 jointly generates audio and video, with all generated audio delivered as native stereo. You can describe dialogue, voice, sound effects, music, and their timing in relation to the visual action.

3

How long can a generated video be?

The official MiniMax H3 launch post states that generated video can be up to 15 seconds. For a longer campaign or narrative, plan several self-contained shots and carry approved identity, product, style, and audio references from one generation to the next.

4

Does the model support 2K resolution?

Yes. MiniMax presents 2K as the default output resolution for MiniMax H3. Its in-context regeneration process revisits the original multimodal context when producing the higher-resolution result, with the goal of recovering details such as small text more faithfully.

5

What can I use as a reference?

The official examples combine text, image, video, and audio context. With MiniMax H3, one source might define a camera move, another a character, and a third a vocal performance. The prompt explains how those sources relate to the target video.

6

Is MiniMax H3 open source?

MiniMax announced plans to release the MiniMax H3 model weights in the days following launch, subject to applicable laws and regulations. That is a stated plan, not a guarantee that every weight, license, inference package, or hardware configuration is already available today.

7

Is the official API available on this site?

Yes. The generator above connects MiniMax H3 through EvoLink for text-to-video, image-to-video, and reference-to-video. Each mode sends only the fields accepted by its H3 endpoint and uses the upload and billing rules shown in the generator.

8

Can I use the output for commercial work?

MiniMax positions MiniMax H3 for commercial content creation, including advertising, branding, e-commerce, product design, UI/UX, and gaming. Your right to publish still depends on the service terms, model license, source-asset rights, likeness permissions, music rights, and applicable law.

Turn your MiniMax H3 brief into a first video draft

Organize the subject, motion, visual references, and audio intent, then use the live generator above to create a native 2K H3 video through the connected EvoLink workflow.

Open the generator