Instella: Fully Open Language Models with Stellar Performance Paper • 2511.10628 • Published Nov 13, 2025 • 7
view article Article SmolLM - blazingly fast and remarkably powerful +1 loubnabnl, anton-l, eliebak • Jul 16, 2024 • 467
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • 28 days ago • 40
view article Article DAVIDAU DELETED MY POST AND BANNED ME FOR ASKING ABOUT BENCHMARKS — HERE IS THE FULL ANALYSIS (9B + 27B + 390 MODELS) DedeProGames • 27 days ago • 16
FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset Paper • 2510.08022 • Published Oct 9, 2025 • 1
SWE-RL Collection Software Engineering Tasks for RL - Images served by Prime's registry • 9 items • Updated Jul 22 • 2
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery Paper • 2602.09892 • Published Feb 10 • 6
PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost Paper • 2603.21383 • Published Mar 22 • 20
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 214
🤏 Smol-Data Collection Tried and tested mixes for strong pretraining. Inspired by https://huggingface.co/blog/codelion/optimal-dataset-mixing • 14 items • Updated Mar 2 • 19
SmolLM3 pretraining datasets Collection datasets used in SmolLM3 pretraining • 15 items • Updated Aug 12, 2025 • 56
Embarrassingly Simple Self-Distillation Improves Code Generation Paper • 2604.01193 • Published Apr 1 • 57