view post Post 4100 We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡Qwen3.6-27B NVFP4 runs on 24GB VRAM.35B-A3B can hit 17,561 tok/s (B200).We also improved accuracy, tool calling, agent use, and looping.Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 See translation 1 reply · 🚀 15 15 🔥 11 11 🤗 1 1 + Reply
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 Text Generation • 45B • Updated 15 days ago • 54.3k • 122