DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

4-step MiniMax-H3 students for joint audio-video generation

Project Page Paper Code Demo Video

Zhengming Yu1,2, Junkun Yuan2, Haotian Yang2, Gordon Guocheng Qian2, Yizhi Wang2, Angtian Wang2, Yiding Yang2, Bo Liu2, Xin Li1, Wenping Wang1, Chongyang Ma2
1Texas A&M University, 2ByteDance

Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)

This repository holds the DMAD students of the paper: 4-step students of MiniMax-H3 (33B, text-to-audio-video) and of Wan2.1-T2V (1.3B and 14B), 4- and 1-step students of SDXL, and 1-step students of the EDM ImageNet-64 teacher. The inference and training code is in the code repository (train/h3, train/wan, train/image).

MiniMax-H3 (text-to-audio-video)

Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio.

File Checkpoint Size
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors the checkpoint of the paper: EMA of the student at iteration 800 of the main run 1.4 GB
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench 1.4 GB
minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors lora_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) 2.0 GB
minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors full_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) 2.0 GB

LoRA layout of the first two: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B [out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2 token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.

The inference code lives in the code repository: inference.py with the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps.

git clone https://github.com/Yzmblog/DMAD.git && cd DMAD   # code + environment setup (see its README)
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
    --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater

ComfyUI

minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors are the same two LoRAs converted exactly to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into attn.qkv_proj adapters of rank 384, alpha = rank, scale 1.0 β€” no rank reduction; 2.0 GB each). Load with LoraLoaderModelOnly at strength 1.0, cfg 1.0, ModelSamplingMiniMaxH3 with shift 12 / audio shift 2, and sample with ComfyUI's lcm sampler and simple scheduler (the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their operating point), or with the equivalent DMAD Sampler + DMAD Sigmas nodes from comfyui/ComfyUI-DMAD.

hf download ZhengmingYu/DMAD --include "minimax_h3/*_comfyui.safetensors" --local-dir /path/to/ComfyUI/models/loras

Wan2.1 (text-to-video)

Full generators (EMA, spectral norm folded in) in the .pth format of train/wan: 4 steps, 480p, 81 frames, no classifier-free guidance.

File Model Size
wan2.1/dmad_wan2pt1_1pt3B.pth Wan2.1-T2V-1.3B student, iteration 17k 2.8 GB
wan2.1/dmad_wan2pt1_14B.pth Wan2.1-T2V-14B student, iteration 19.5k 29 GB
hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt
bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples   # in train/wan

SDXL (text-to-image)

UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see train/image for the sampling code.

File Model Size
sdxl/dmad_sdxl_4step_unet_fp16.bin 4-step student (backward simulation), iteration 17k 5.1 GB
sdxl/dmad_sdxl_1step_unet_fp16.bin 1-step student (ODE init), iteration 49k 5.1 GB
sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin 1-step student (ODE init, frozen critic backbone), iteration 17.5k 5.1 GB
import torch
from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel
from huggingface_hub import hf_hub_download

base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16)
unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0]
# 1-step models: num_inference_steps=1, timesteps=[399]

ImageNet-64 (class-conditional)

1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the checkpoint_model_<iteration>/pytorch_model_ema.bin layout that train/image's evaluation reads directly.

Folder Setting Size
imagenet64/dmad_imagenet_gaproute/ teacher-UNet critic + gap routing, iteration 568k 1.2 GB
imagenet64/dmad_imagenet_pgcritic/ pretrained-feature critic, iteration 108k 1.2 GB
imagenet64/dmad_imagenet_frozencritic/ frozen teacher-UNet critic + gap routing, iteration 221k 1.2 GB
hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt
python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ...   # in train/image

Citation

@misc{yu2026dmad,
  title         = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation},
  author        = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and
                   Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma},
  year          = {2026},
  eprint        = {2610.02188},
  archivePrefix = {arXiv}
}

License

Each family of weights is a derivative of its base model and is distributed under that model's license:

Images and videos produced with these weights are AI-generated.

Downloads last month
-
Inference Providers NEW

This task can take several minutes

Model tree for ZhengmingYu/DMAD

Adapter
(117)
this model

Spaces using ZhengmingYu/DMAD 3

Paper for ZhengmingYu/DMAD