A Quick MiniMaxH3 Setup Guide for SwarmUI

Last edit 28 Sept 2026, 16:30 by CroissantAfterDarkHistory

MiniMax H3 in SwarmUI Installation & First Video Guide ⚠️ IMPORTANT HARDWARE WARNING MiniMax H3 is a very large video model. This is not a lightweight model that can be casually installed on an 8 GB GPU. The primary H3 FL2AV INT8 model is approximately 21 GB on disk, and the H3 text encoder is also very large. During generation, H3 can require substantial system RAM in addition to VRAM because parts of the model may need to be offloaded from the GPU. Recommended: 16 GB+ VRAM and 64 GB system RAM Possible: 12–16 GB VRAM with 32 GB system RAM, depending on settings and offloading Not recommended: 8 GB VRAM / 16 GB system RAM Expect large downloads, significant disk usage, and potentially long generation times on lower-VRAM GPUs.

MiniMax H3 Installation Guide for SwarmUI


1. What Is MiniMax H3?

MiniMax H3 is a large video-generation model capable of generating both video and audio.

It supports:

  • Text-to-Video
  • Image-to-Video
  • First-frame / last-frame generation
  • Reference-to-Video
  • Dialogue
  • Sound effects
  • Music
  • Environmental audio

H3 is a very large model, however, so hardware requirements are substantially higher than most image-generation models.


2. Hardware Requirements

There is no single hard VRAM requirement because memory usage depends on the model variant, quantization, resolution, duration, and offloading.

For a practical local setup:

HardwareRecommendation
GPU VRAM16 GB+ recommended starting point
System RAM32 GB workable, 64 GB recommended
Free storage50 GB+ recommended
Batch size1 to start

For a 16 GB GPU, start with the pruned INT8 ConvRot H3 model and the NVFP4/AWQ Qwen3-VL encoder.

Expect H3 to use substantial system RAM as well as VRAM.


3. Your SwarmUI Folder

Find your SwarmUI installation.

It should look approximately like:

text
SwarmUI/
├── Models/
├── Backend/
├── Data/
├── src/
└── ...

The folder we care about is:

text
SwarmUI/Models/

This is your SwarmUI model root.

Do not put the H3 files in arbitrary folders outside this structure.


4. H3 Model Folder Structure in SwarmUI

For H3, create/use these folders:

text
SwarmUI/
└── Models/
    ├── diffusion_models/
    ├── text_encoders/
    ├── VAE/
    │   └── MiniMaxH3/
    └── loras/

The important distinction is that SwarmUI's model root is Models/.

The paths below are therefore relative to:

text
SwarmUI/Models/

5. Download the Main H3 Model

For the standard H3 Text-to-Video / Image-to-Video workflow, download:

text
minimax_h3_fl2va_pruned_int8_convrot.safetensors

Put it here:

text
SwarmUI/Models/diffusion_models/

Your final path should be:

text
SwarmUI/
└── Models/
    └── diffusion_models/
        └── minimax_h3_fl2va_pruned_int8_convrot.safetensors

This is the main H3 video-generation model.

The file is approximately 21 GB, so make sure you have enough storage before downloading it.


6. Download the Qwen3-VL Text Encoder

Download:

text
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

Put it here:

text
SwarmUI/Models/text_encoders/

Your final path should be:

text
SwarmUI/
└── Models/
    └── text_encoders/
        └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

This is the text/multimodal encoder H3 uses to interpret your prompt and reference information.

The NVFP4/AWQ version is particularly useful for newer NVIDIA GPUs.


7. Download the H3 Video VAE

SwarmUI currently registers a MiniMax H3 Video VAE under its own H3-specific folder.

Download:

text
minimax_h3_video_vae_fp16.safetensors

Put it here:

text
SwarmUI/Models/VAE/MiniMaxH3/

Your final path should be:

text
SwarmUI/
└── Models/
    └── VAE/
        └── MiniMaxH3/
            └── minimax_h3_video_vae_fp16.safetensors

This is a SwarmUI path.

Do not put it in:

text
SwarmUI/Models/vae/

for this guide.

SwarmUI's current model registry specifically defines the H3 video VAE as:

text
MiniMaxH3/minimax_h3_video_vae_fp16.safetensors

under its VAE model set.


8. Download the H3 Audio VAE

Download:

text
minimax_h3_audio_vae_fp32.safetensors

Put it in the same SwarmUI H3 VAE folder:

text
SwarmUI/Models/VAE/MiniMaxH3/

Your folder should now contain:

text
SwarmUI/
└── Models/
    └── VAE/
        └── MiniMaxH3/
            ├── minimax_h3_video_vae_fp16.safetensors
            └── minimax_h3_audio_vae_fp32.safetensors

SwarmUI currently registers this audio VAE in the same MiniMaxH3 VAE directory.


9. Optional: H3 Turbo LoRA

Once you have the normal H3 workflow working, you can install the Turbo LoRA.

Download:

text
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

Put it here:

text
SwarmUI/Models/loras/

So:

text
SwarmUI/
└── Models/
    └── loras/
        └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

Do not install Turbo yet if this is your first H3 setup.

Get the normal workflow working first.


10. Your SwarmUI H3 Folder Should Look Like This

After installing the basic H3 files:

text
SwarmUI/
└── Models/
    │
    ├── diffusion_models/
    │   └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
    │
    ├── text_encoders/
    │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
    │
    ├── VAE/
    │   └── MiniMaxH3/
    │       ├── minimax_h3_video_vae_fp16.safetensors
    │       └── minimax_h3_audio_vae_fp32.safetensors
    │
    └── loras/
        └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

The Turbo LoRA is optional.

That is the SwarmUI layout.


11. Configure SwarmUI's ModelRoot

Normally, SwarmUI already uses its own:

text
SwarmUI/Models/

as its model root.

To verify it:

text
Server
→ Server Configuration

Look under:

text
Paths

Find:

text
ModelRoot

For a normal SwarmUI installation, this should point to your SwarmUI model directory.

For example:

text
C:/AI/SwarmUI/Models

or:

text
D:/AI/SwarmUI/Models

Do not point ModelRoot at SwarmUI/Models/diffusion_models.

It should point to:

text
SwarmUI/Models

The folders such as diffusion_models, text_encoders, and VAE live underneath that root.

SwarmUI's documentation defines ModelRoot as the root model directory and then uses separate subfolder settings beneath it.


12. Do NOT Change SDModelFolder for H3

This is another place where things can get confusing.

SDModelFolder is not your H3 diffusion-model folder.

It is the folder SwarmUI uses for traditional Stable Diffusion-style checkpoint models.

H3's main model belongs in:

text
SwarmUI/Models/diffusion_models/

So don't move the H3 model into:

text
SwarmUI/Models/Stable-Diffusion/

just because you see SDModelFolder in the settings.

SwarmUI itself distinguishes the traditional SD model folder from the modern diffusion_models folder.


13. Restart SwarmUI

After installing the files:

  1. Save your settings.
  2. Shut down SwarmUI.
  3. Start SwarmUI again.
  4. Refresh the browser.

This gives SwarmUI a clean opportunity to rescan the model directories.


14. Finding the H3 Model in SwarmUI

Go to:

text
Generate

Then open the model selector.

Your H3 diffusion model should be discoverable from the:

text
diffusion_models

folder.

Select:

text
minimax_h3_fl2va_pruned_int8_convrot.safetensors

If it appears, SwarmUI has successfully found the main H3 model.


15. First Generation

Do not start with an elaborate 15-second cinematic masterpiece.

We're testing the installation first.

Start around:

text
Resolution:
1344 × 768

Steps:
20

Batch:
1

H3's documented native canvas uses a 768-pixel short edge, with 1344×768 being the standard 16:9 example.


16. First Test Prompt

Use something simple:

text
A woman standing on a quiet city street at sunset. 
She gently brushes her hair behind her ear while the camera slowly moves toward her. 
Warm evening light reflects from the surrounding buildings. 
Natural movement, subtle city ambience, gentle footsteps.

The first test is simply checking whether:

  • H3 loads
  • The text encoder loads
  • The video VAE loads
  • The audio VAE loads
  • Video generation works
  • Audio generation works

If you get a successful video with sound:

Congratulations. H3 is installed.


17. H3 Audio

One of H3's major features is simultaneous video and audio generation.

You can describe audio directly in your prompt.

For example:

text
A woman walks through a rainy city street at night.
Rain falls steadily around her.
Her footsteps splash through shallow puddles.
Distant cars pass through the intersection.
She quietly says, "I should have stayed home."
A soft piano melody plays in the background.

H3 can interpret the visual and audio descriptions together.


18. Image-to-Video

Once your basic Text-to-Video generation works, try Image-to-Video.

Provide an image as the starting frame.

H3 can use:

text
first frame

and optionally:

text
last frame

The model generates the movement between the provided keyframes.

The official H3 workflow supports optional first- and last-frame inputs.


19. Turbo Mode

After you've confirmed normal H3 generation works, try Turbo.

The standard H3 workflow uses around:

text
20 steps

Turbo reduces this substantially, with the official workflow using an 8-step Turbo configuration.

Turbo is useful when you want faster experimentation.

Start with normal H3 first.


20. Reference-to-Video

H3 also has a separate Reference-to-Video model.

If you want R2V, you will need:

text
minimax_h3_ref2va_pruned_int8_convrot.safetensors

Put that additional diffusion model in:

text
SwarmUI/Models/diffusion_models/

Your folder then becomes:

text
SwarmUI/
└── Models/
    └── diffusion_models/
        ├── minimax_h3_fl2va_pruned_int8_convrot.safetensors
        └── minimax_h3_ref2va_pruned_int8_convrot.safetensors

The two models are separate.

FL2VA

Used for:

  • Text-to-Video
  • Image-to-Video
  • First-frame generation
  • Last-frame generation

REF2VA

Used for:

  • Reference-to-Video

Do not replace FL2VA with REF2VA.

Keep both if you want both workflows.


21. Troubleshooting: H3 Model Doesn't Appear

Check the physical file location.

The main model should be:

text
SwarmUI/Models/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors

The text encoder should be:

text
SwarmUI/Models/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

The video VAE should be:

text
SwarmUI/Models/VAE/MiniMaxH3/minimax_h3_video_vae_fp16.safetensors

The audio VAE should be:

text
SwarmUI/Models/VAE/MiniMaxH3/minimax_h3_audio_vae_fp32.safetensors

The Turbo LoRA should be:

text
SwarmUI/Models/loras/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

If those paths are correct, restart SwarmUI and check again.


22. Troubleshooting: Out of VRAM

If your GPU runs out of VRAM:

  1. Make sure you're using the pruned INT8 ConvRot diffusion model.
  2. Use batch size 1.
  3. Lower the resolution.
  4. Reduce video duration.
  5. Close other GPU-heavy programs.
  6. Give SwarmUI/backend more opportunity to offload to system RAM.

H3 is simply a gigantic model.

A 16 GB GPU does not mean the entire H3 workload will fit into 16 GB of total memory.


23. Troubleshooting: High System RAM Usage

High system RAM usage is expected with a model this large.

The model can offload portions of the workload from VRAM into system memory.

For that reason:

text
32 GB RAM

can be workable, but:

text
64 GB RAM

is a much more comfortable target.


24. Final SwarmUI Checklist

Before troubleshooting anything, verify these five locations:

Main H3 model

text
SwarmUI/Models/diffusion_models/
minimax_h3_fl2va_pruned_int8_convrot.safetensors

Text encoder

text
SwarmUI/Models/text_encoders/
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

Video VAE

text
SwarmUI/Models/VAE/MiniMaxH3/
minimax_h3_video_vae_fp16.safetensors

Audio VAE

text
SwarmUI/Models/VAE/MiniMaxH3/
minimax_h3_audio_vae_fp32.safetensors

Optional Turbo LoRA

text
SwarmUI/Models/loras/
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

25. The Whole Setup at a Glance

Your SwarmUI installation should ultimately look like:

text
SwarmUI/
│
├── Models/
│   │
│   ├── diffusion_models/
│   │   ├── minimax_h3_fl2va_pruned_int8_convrot.safetensors
│   │   └── minimax_h3_ref2va_pruned_int8_convrot.safetensors    ← optional
│   │
│   ├── text_encoders/
│   │   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│   │
│   ├── VAE/
│   │   └── MiniMaxH3/
│   │       ├── minimax_h3_video_vae_fp16.safetensors
│   │       └── minimax_h3_audio_vae_fp32.safetensors
│   │
│   └── loras/
│       └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
│
├── Backend/
├── Data/
└── ...

This guide is using SwarmUI's model structure throughout.

You should not need to translate these paths into another application's folder structure.


26. Recommended Starting Configuration

For a 16 GB NVIDIA GPU, start with:

text
Diffusion:
minimax_h3_fl2va_pruned_int8_convrot.safetensors

Text Encoder:
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

Video VAE:
minimax_h3_video_vae_fp16.safetensors

Audio VAE:
minimax_h3_audio_vae_fp32.safetensors

Resolution:
1344 × 768

Steps:
20

Batch:
1

Turbo:
OFF

Reference images:
OFF

First/last frame:
OFF for the first test

Once that works:

text
Normal H3
    ↓
Image-to-Video
    ↓
First/Last Frame
    ↓
Turbo
    ↓
Reference-to-Video

That gives you a sane progression without throwing five new variables at the GPU at once.

27. Official Resources

Use the official SwarmUI documentation for model-root configuration and the official MiniMax H3 repository for the actual model files.

Official Sexter creator documentation. Older Unity / SmGF folder-mod docs remain at docs.unzipped.games. Creator Corner is for community models and production workflows.