MOSS-Video-and-Audio
MOSS Video and Audio (MOVA) is an open-source foundation model developed by OpenMOSS that generates synchronized, high-…
Description
MOSS Video and Audio (MOVA) is an open-source foundation model developed by OpenMOSS that generates synchronized, high-fidelity video and audio in a single end-to-end inference step. Built on a 32-billion parameter Mixture-of-Experts (MoE) architecture, it employs an asymmetric dual-tower design to achieve precise lip-synchronization and eliminate the error accumulation found in traditional cascaded systems. This model is hosted by EmpirioLabs.ai and is exclusive to the Poe platform. Learn more: https://mosi.cn/models/mova Notes: - Generations may take upwards of 20 minutes to complete. - Image-to-Video generations will likely yield superior results. - Supported image upload file types: .jpg, .jpeg, .png, .webp, .heic, .heif, .bmp, .tiff, .tif, .gif - Only 1 image attachment is supported (first frame), and video attachments are not supported. Parameter controls available: 1. Generation - Resolution `360p` or q720p` (default: 360p) - Aspect ratio `landscape` or `portrait` or `square` (default: landscape; auto-detected for uploaded images) - Duration [2-8] video length in seconds (default: 4) - t2v quality 1fast` or `quality` (default: quality; only for text-to-video) 2. Advanced - Inference steps [10-50] more steps = higher quality but slower (default: 25, step: 5) - CFG scale [1.0-10.0] prompt adherence strength, higher = more literal (default: 5.0, step: 0.5) - Sigma_shift [1.0-10.0] noise schedule shift, 360p only (default: 5.0, step: 0.5) - Seed [number] for reproducibility, omit for random - Negative prompt "your negative prompt" leave blank for optimized default
Author
EmpirioLabs AI
Platform
web
Pricing model
free
Categories
Tags
Capabilities
- Text input
- Video generation
- By EmpirioLabs AI