MiniMax H3 is a multimodal video generation model that produces 4 to 15 second clips at 768p or 2K, 24 fps, each with a…
MiniMax H3 is a multimodal video generation model that produces 4 to 15 second clips at 768p or 2K, 24 fps, each with a native stereo audio track. It supports Text-to-Video, Image-to-Video, First/Last Frame, and Reference-to-Video, auto-detected from your attachments or set manually. Input Limits & Requirements - Prompt: up to 7,000 characters. Can describe action, camera work, dialogue, sound effects, and on-screen text - Images: JPEG, PNG, WEBP, or HEIC. 256 to 5760px per side, aspect ratio within 1:2.5-2.5:1, max 30MB each - Videos: MP4, MOV, MKV, or WEBM. 2 to 15 seconds each, max 50MB each - Audio: WAV or MP3. 2 to 15 seconds each, max 15MB each, and requires an accompanying image or video reference - Image-to-Video: 1 image (the opening frame). First/Last Frame: exactly 2 images - Reference-to-Video: max 9 images, 3 videos, and 3 audio clips, 12 files total This bot supports optional parameters for additional customization.
EmpirioLabs AI
web
free
No benchmark results have been added yet.