Instructions to use yresearch/DMax-Math-MMD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yresearch/DMax-Math-MMD with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yresearch/DMax-Math-MMD", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("yresearch/DMax-Math-MMD", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yresearch/DMax-Math-MMD with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yresearch/DMax-Math-MMD" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yresearch/DMax-Math-MMD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/yresearch/DMax-Math-MMD
- SGLang
How to use yresearch/DMax-Math-MMD with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yresearch/DMax-Math-MMD" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yresearch/DMax-Math-MMD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yresearch/DMax-Math-MMD" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yresearch/DMax-Math-MMD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use yresearch/DMax-Math-MMD with Docker Model Runner:
docker model run hf.co/yresearch/DMax-Math-MMD
DMax-Math-MMD
Representation-Space MMD for Diffusion Language Models
DMax-Math-MMD is a 16B diffusion language model for mathematical reasoning, obtained by MMD post-training of DMax-Math-16B. It builds on LLaDA2.0-mini and uses DMax's hybrid masked–uniform block diffusion.
The post-training objective minimizes Maximum Mean Discrepancy (MMD) between model samples and reference responses, measured in the representation space of a frozen diffusion language model. The project reports increased tokens per forward pass while maintaining or improving accuracy on the benchmarks below.
Reference results
Results reported in the project README, with decoding threshold 0.85. Each entry is accuracy (%) / tokens per forward pass (TPF). Baseline results are attributed to the original DMax paper in the project README.
| Method | GSM8K | MATH500 | Minerva-Algebra | ASDIV |
|---|---|---|---|---|
| DMax-Math | 92.1 / 5.48 | 75.4 / 5.94 | 91.5 / 7.03 | 92.5 / 5.62 |
| DMax-Math-MMD | 92.1 / 6.15 | 76.0 / 6.84 | 92.1 / 8.19 | 92.9 / 6.20 |
Results can vary with hardware, tensor parallelism, and library versions. TPF measures decoding parallelism; wall-clock speed also depends on the runtime.
Evaluation and inference
Use DMax's dInfer evaluation pipeline. After following the evaluation environment setup in the project README, run from the project repository root:
conda activate dinfer
DOMAIN=math MODEL_PATH=yresearch/DMax-Math-MMD bash scripts/eval.sh
The project's math evaluation uses threshold 0.85 and evaluates GSM8K, MATH500, Minerva-Algebra, and ASDIV.
The checkpoint includes nine Safetensors shards with BF16 parameters and FP32
router bias buffers, plus its tokenizer, chat template, and custom model code.
Direct Transformers loading requires trust_remote_code=True. Pass
dtype=torch.bfloat16 for BF16 loading; the current configuration declares FP32.
Generation requires the DMax/dInfer diffusion decoder. The included custom
backbone does not implement the standard Transformers .generate() interface.
License and acknowledgements
The model follows the Apache-2.0 license of the DMax-Math base checkpoint and LLaDA2.0-mini. The included model implementation retains its Apache-2.0 notices. The separate MMD training repository is MIT-licensed, with Apache-2.0 third-party components.
We thank the authors of DMax, LLaDA2.0-mini, and dInfer for releasing their models, data, and code.
Citation
If you find this work useful in your research, please consider citing:
@article{drobyshevskiy2026mmd,
title = {Representation-Space MMD for Diffusion Language Models},
author = {Drobyshevskiy, Ilya and Sudakov, Ilia and Semenov, Maksim and Kuznedelev, Denis and
Ignatov, Maksim and Temirchev, Pavel and Balagansky, Nikita and
Meshchaninov, Viacheslav and Gushchin, Nikita and Baranchuk, Dmitry},
journal = {arXiv preprint arXiv:2610.06648},
year = {2026}
}
- Downloads last month
- 324
docker model run hf.co/yresearch/DMax-Math-MMD