Hikari07jp/Ternary-Bonsai-27B-Abliterated-LowDeg-GGUF Text Generation • 27B • Updated 1 day ago • 780 • 10
JBrightmanAI/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED 9B • Updated 6 days ago • 28 • 1
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper • 2607.07769 • Published 16 days ago • 10
view post Post 7679 Frontier models use distillation as a step of their post-training pipelines. In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL! See translation 3 replies · 👍 14 14 🔥 7 7 ❤️ 3 3 + Reply
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2 Text Generation • 12B • Updated 23 days ago • 160k • 72
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF Text Generation • 12B • Updated Jun 19 • 482k • 1.27k