MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations Paper • 2610.02480 • Published 7 days ago • 6
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 22 days ago • 67
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions Paper • 2510.08999 • Published about 1 month ago • 19
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Paper • 2608.17744 • Published Aug 18 • 16
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published Aug 19 • 21
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published Aug 13 • 143
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Paper • 2608.13430 • Published Aug 13 • 13