DMax-Math-MMD

License: Apache-2.0 Arxiv GitHub: Code

Representation-Space MMD for Diffusion Language Models

DMax-Math-MMD is a 16B diffusion language model for mathematical reasoning, obtained by MMD post-training of DMax-Math-16B. It builds on LLaDA2.0-mini and uses DMax's hybrid masked–uniform block diffusion.

The post-training objective minimizes Maximum Mean Discrepancy (MMD) between model samples and reference responses, measured in the representation space of a frozen diffusion language model. The project reports increased tokens per forward pass while maintaining or improving accuracy on the benchmarks below.

Reference results

Results reported in the project README, with decoding threshold 0.85. Each entry is accuracy (%) / tokens per forward pass (TPF). Baseline results are attributed to the original DMax paper in the project README.

Method GSM8K MATH500 Minerva-Algebra ASDIV
DMax-Math 92.1 / 5.48 75.4 / 5.94 91.5 / 7.03 92.5 / 5.62
DMax-Math-MMD 92.1 / 6.15 76.0 / 6.84 92.1 / 8.19 92.9 / 6.20

Results can vary with hardware, tensor parallelism, and library versions. TPF measures decoding parallelism; wall-clock speed also depends on the runtime.

Evaluation and inference

Use DMax's dInfer evaluation pipeline. After following the evaluation environment setup in the project README, run from the project repository root:

conda activate dinfer
DOMAIN=math MODEL_PATH=yresearch/DMax-Math-MMD bash scripts/eval.sh

The project's math evaluation uses threshold 0.85 and evaluates GSM8K, MATH500, Minerva-Algebra, and ASDIV.

The checkpoint includes nine Safetensors shards with BF16 parameters and FP32 router bias buffers, plus its tokenizer, chat template, and custom model code. Direct Transformers loading requires trust_remote_code=True. Pass dtype=torch.bfloat16 for BF16 loading; the current configuration declares FP32.

Generation requires the DMax/dInfer diffusion decoder. The included custom backbone does not implement the standard Transformers .generate() interface.

License and acknowledgements

The model follows the Apache-2.0 license of the DMax-Math base checkpoint and LLaDA2.0-mini. The included model implementation retains its Apache-2.0 notices. The separate MMD training repository is MIT-licensed, with Apache-2.0 third-party components.

We thank the authors of DMax, LLaDA2.0-mini, and dInfer for releasing their models, data, and code.

Citation

If you find this work useful in your research, please consider citing:

@article{drobyshevskiy2026mmd,
  title  = {Representation-Space MMD for Diffusion Language Models},
  author = {Drobyshevskiy, Ilya and Sudakov, Ilia and Semenov, Maksim and Kuznedelev, Denis and
            Ignatov, Maksim and Temirchev, Pavel and Balagansky, Nikita and
            Meshchaninov, Viacheslav and Gushchin, Nikita and Baranchuk, Dmitry},
  journal = {arXiv preprint arXiv:2610.06648},
  year    = {2026}
}
Downloads last month
324
Safetensors
Model size
16B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yresearch/DMax-Math-MMD

Finetuned
(1)
this model

Dataset used to train yresearch/DMax-Math-MMD

Collection including yresearch/DMax-Math-MMD

Paper for yresearch/DMax-Math-MMD