Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 6 days ago • 140
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Paper • 2607.29209 • Published Jul 31 • 34
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 108
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published Jul 2 • 309
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion Paper • 2606.14885 • Published Jun 12 • 13
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 77
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion Paper • 2606.14885 • Published Jun 12 • 13
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion Paper • 2606.14885 • Published Jun 12 • 13