A Quick MiniMaxH3 Setup Guide for SwarmUI
MiniMax H3 in SwarmUI Installation & First Video Guide ⚠️ IMPORTANT HARDWARE WARNING MiniMax H3 is a very large video model. This is not a lightweight model that can be casually installed on an 8 GB GPU. The primary H3 FL2AV INT8 model is approximately 21 GB on disk, and the H3 text encoder is also very large. During generation, H3 can require substantial system RAM in addition to VRAM because parts of the model may need to be offloaded from the GPU. Recommended: 16 GB+ VRAM and 64 GB system RAM Possible: 12–16 GB VRAM with 32 GB system RAM, depending on settings and offloading Not recommended: 8 GB VRAM / 16 GB system RAM Expect large downloads, significant disk usage, and potentially long generation times on lower-VRAM GPUs.
MiniMax H3 Installation Guide for SwarmUI
1. What Is MiniMax H3?
MiniMax H3 is a large video-generation model capable of generating both video and audio.
It supports:
- Text-to-Video
- Image-to-Video
- First-frame / last-frame generation
- Reference-to-Video
- Dialogue
- Sound effects
- Music
- Environmental audio
H3 is a very large model, however, so hardware requirements are substantially higher than most image-generation models.
2. Hardware Requirements
There is no single hard VRAM requirement because memory usage depends on the model variant, quantization, resolution, duration, and offloading.
For a practical local setup:
| Hardware | Recommendation |
|---|---|
| GPU VRAM | 16 GB+ recommended starting point |
| System RAM | 32 GB workable, 64 GB recommended |
| Free storage | 50 GB+ recommended |
| Batch size | 1 to start |
For a 16 GB GPU, start with the pruned INT8 ConvRot H3 model and the NVFP4/AWQ Qwen3-VL encoder.
Expect H3 to use substantial system RAM as well as VRAM.
3. Your SwarmUI Folder
Find your SwarmUI installation.
It should look approximately like:
SwarmUI/
├── Models/
├── Backend/
├── Data/
├── src/
└── ...
The folder we care about is:
SwarmUI/Models/
This is your SwarmUI model root.
Do not put the H3 files in arbitrary folders outside this structure.
4. H3 Model Folder Structure in SwarmUI
For H3, create/use these folders:
SwarmUI/
└── Models/
├── diffusion_models/
├── text_encoders/
├── VAE/
│ └── MiniMaxH3/
└── loras/
The important distinction is that SwarmUI's model root is Models/.
The paths below are therefore relative to:
SwarmUI/Models/
5. Download the Main H3 Model
For the standard H3 Text-to-Video / Image-to-Video workflow, download:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Put it here:
SwarmUI/Models/diffusion_models/
Your final path should be:
SwarmUI/
└── Models/
└── diffusion_models/
└── minimax_h3_fl2va_pruned_int8_convrot.safetensors
This is the main H3 video-generation model.
The file is approximately 21 GB, so make sure you have enough storage before downloading it.
6. Download the Qwen3-VL Text Encoder
Download:
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Put it here:
SwarmUI/Models/text_encoders/
Your final path should be:
SwarmUI/
└── Models/
└── text_encoders/
└── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
This is the text/multimodal encoder H3 uses to interpret your prompt and reference information.
The NVFP4/AWQ version is particularly useful for newer NVIDIA GPUs.
7. Download the H3 Video VAE
SwarmUI currently registers a MiniMax H3 Video VAE under its own H3-specific folder.
Download:
minimax_h3_video_vae_fp16.safetensors
Put it here:
SwarmUI/Models/VAE/MiniMaxH3/
Your final path should be:
SwarmUI/
└── Models/
└── VAE/
└── MiniMaxH3/
└── minimax_h3_video_vae_fp16.safetensors
This is a SwarmUI path.
Do not put it in:
SwarmUI/Models/vae/
for this guide.
SwarmUI's current model registry specifically defines the H3 video VAE as:
MiniMaxH3/minimax_h3_video_vae_fp16.safetensors
under its VAE model set.
8. Download the H3 Audio VAE
Download:
minimax_h3_audio_vae_fp32.safetensors
Put it in the same SwarmUI H3 VAE folder:
SwarmUI/Models/VAE/MiniMaxH3/
Your folder should now contain:
SwarmUI/
└── Models/
└── VAE/
└── MiniMaxH3/
├── minimax_h3_video_vae_fp16.safetensors
└── minimax_h3_audio_vae_fp32.safetensors
SwarmUI currently registers this audio VAE in the same MiniMaxH3 VAE directory.
9. Optional: H3 Turbo LoRA
Once you have the normal H3 workflow working, you can install the Turbo LoRA.
Download:
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Put it here:
SwarmUI/Models/loras/
So:
SwarmUI/
└── Models/
└── loras/
└── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Do not install Turbo yet if this is your first H3 setup.
Get the normal workflow working first.
10. Your SwarmUI H3 Folder Should Look Like This
After installing the basic H3 files:
SwarmUI/
└── Models/
│
├── diffusion_models/
│ └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
│
├── text_encoders/
│ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│
├── VAE/
│ └── MiniMaxH3/
│ ├── minimax_h3_video_vae_fp16.safetensors
│ └── minimax_h3_audio_vae_fp32.safetensors
│
└── loras/
└── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
The Turbo LoRA is optional.
That is the SwarmUI layout.
11. Configure SwarmUI's ModelRoot
Normally, SwarmUI already uses its own:
SwarmUI/Models/
as its model root.
To verify it:
Server
→ Server Configuration
Look under:
Paths
Find:
ModelRoot
For a normal SwarmUI installation, this should point to your SwarmUI model directory.
For example:
C:/AI/SwarmUI/Models
or:
D:/AI/SwarmUI/Models
Do not point ModelRoot at SwarmUI/Models/diffusion_models.
It should point to:
SwarmUI/Models
The folders such as diffusion_models, text_encoders, and VAE live underneath that root.
SwarmUI's documentation defines ModelRoot as the root model directory and then uses separate subfolder settings beneath it.
12. Do NOT Change SDModelFolder for H3
This is another place where things can get confusing.
SDModelFolder is not your H3 diffusion-model folder.
It is the folder SwarmUI uses for traditional Stable Diffusion-style checkpoint models.
H3's main model belongs in:
SwarmUI/Models/diffusion_models/
So don't move the H3 model into:
SwarmUI/Models/Stable-Diffusion/
just because you see SDModelFolder in the settings.
SwarmUI itself distinguishes the traditional SD model folder from the modern diffusion_models folder.
13. Restart SwarmUI
After installing the files:
- Save your settings.
- Shut down SwarmUI.
- Start SwarmUI again.
- Refresh the browser.
This gives SwarmUI a clean opportunity to rescan the model directories.
14. Finding the H3 Model in SwarmUI
Go to:
Generate
Then open the model selector.
Your H3 diffusion model should be discoverable from the:
diffusion_models
folder.
Select:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
If it appears, SwarmUI has successfully found the main H3 model.
15. First Generation
Do not start with an elaborate 15-second cinematic masterpiece.
We're testing the installation first.
Start around:
Resolution:
1344 × 768
Steps:
20
Batch:
1
H3's documented native canvas uses a 768-pixel short edge, with 1344×768 being the standard 16:9 example.
16. First Test Prompt
Use something simple:
A woman standing on a quiet city street at sunset.
She gently brushes her hair behind her ear while the camera slowly moves toward her.
Warm evening light reflects from the surrounding buildings.
Natural movement, subtle city ambience, gentle footsteps.
The first test is simply checking whether:
- H3 loads
- The text encoder loads
- The video VAE loads
- The audio VAE loads
- Video generation works
- Audio generation works
If you get a successful video with sound:
Congratulations. H3 is installed.
17. H3 Audio
One of H3's major features is simultaneous video and audio generation.
You can describe audio directly in your prompt.
For example:
A woman walks through a rainy city street at night.
Rain falls steadily around her.
Her footsteps splash through shallow puddles.
Distant cars pass through the intersection.
She quietly says, "I should have stayed home."
A soft piano melody plays in the background.
H3 can interpret the visual and audio descriptions together.
18. Image-to-Video
Once your basic Text-to-Video generation works, try Image-to-Video.
Provide an image as the starting frame.
H3 can use:
first frame
and optionally:
last frame
The model generates the movement between the provided keyframes.
The official H3 workflow supports optional first- and last-frame inputs.
19. Turbo Mode
After you've confirmed normal H3 generation works, try Turbo.
The standard H3 workflow uses around:
20 steps
Turbo reduces this substantially, with the official workflow using an 8-step Turbo configuration.
Turbo is useful when you want faster experimentation.
Start with normal H3 first.
20. Reference-to-Video
H3 also has a separate Reference-to-Video model.
If you want R2V, you will need:
minimax_h3_ref2va_pruned_int8_convrot.safetensors
Put that additional diffusion model in:
SwarmUI/Models/diffusion_models/
Your folder then becomes:
SwarmUI/
└── Models/
└── diffusion_models/
├── minimax_h3_fl2va_pruned_int8_convrot.safetensors
└── minimax_h3_ref2va_pruned_int8_convrot.safetensors
The two models are separate.
FL2VA
Used for:
- Text-to-Video
- Image-to-Video
- First-frame generation
- Last-frame generation
REF2VA
Used for:
- Reference-to-Video
Do not replace FL2VA with REF2VA.
Keep both if you want both workflows.
21. Troubleshooting: H3 Model Doesn't Appear
Check the physical file location.
The main model should be:
SwarmUI/Models/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors
The text encoder should be:
SwarmUI/Models/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
The video VAE should be:
SwarmUI/Models/VAE/MiniMaxH3/minimax_h3_video_vae_fp16.safetensors
The audio VAE should be:
SwarmUI/Models/VAE/MiniMaxH3/minimax_h3_audio_vae_fp32.safetensors
The Turbo LoRA should be:
SwarmUI/Models/loras/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
If those paths are correct, restart SwarmUI and check again.
22. Troubleshooting: Out of VRAM
If your GPU runs out of VRAM:
- Make sure you're using the pruned INT8 ConvRot diffusion model.
- Use batch size 1.
- Lower the resolution.
- Reduce video duration.
- Close other GPU-heavy programs.
- Give SwarmUI/backend more opportunity to offload to system RAM.
H3 is simply a gigantic model.
A 16 GB GPU does not mean the entire H3 workload will fit into 16 GB of total memory.
23. Troubleshooting: High System RAM Usage
High system RAM usage is expected with a model this large.
The model can offload portions of the workload from VRAM into system memory.
For that reason:
32 GB RAM
can be workable, but:
64 GB RAM
is a much more comfortable target.
24. Final SwarmUI Checklist
Before troubleshooting anything, verify these five locations:
Main H3 model
SwarmUI/Models/diffusion_models/
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Text encoder
SwarmUI/Models/text_encoders/
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Video VAE
SwarmUI/Models/VAE/MiniMaxH3/
minimax_h3_video_vae_fp16.safetensors
Audio VAE
SwarmUI/Models/VAE/MiniMaxH3/
minimax_h3_audio_vae_fp32.safetensors
Optional Turbo LoRA
SwarmUI/Models/loras/
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
25. The Whole Setup at a Glance
Your SwarmUI installation should ultimately look like:
SwarmUI/
│
├── Models/
│ │
│ ├── diffusion_models/
│ │ ├── minimax_h3_fl2va_pruned_int8_convrot.safetensors
│ │ └── minimax_h3_ref2va_pruned_int8_convrot.safetensors ← optional
│ │
│ ├── text_encoders/
│ │ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│ │
│ ├── VAE/
│ │ └── MiniMaxH3/
│ │ ├── minimax_h3_video_vae_fp16.safetensors
│ │ └── minimax_h3_audio_vae_fp32.safetensors
│ │
│ └── loras/
│ └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
│
├── Backend/
├── Data/
└── ...
This guide is using SwarmUI's model structure throughout.
You should not need to translate these paths into another application's folder structure.
26. Recommended Starting Configuration
For a 16 GB NVIDIA GPU, start with:
Diffusion:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Text Encoder:
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Video VAE:
minimax_h3_video_vae_fp16.safetensors
Audio VAE:
minimax_h3_audio_vae_fp32.safetensors
Resolution:
1344 × 768
Steps:
20
Batch:
1
Turbo:
OFF
Reference images:
OFF
First/last frame:
OFF for the first test
Once that works:
Normal H3
↓
Image-to-Video
↓
First/Last Frame
↓
Turbo
↓
Reference-to-Video
That gives you a sane progression without throwing five new variables at the GPU at once.
27. Official Resources
Use the official SwarmUI documentation for model-root configuration and the official MiniMax H3 repository for the actual model files.
