The Problem Is the Problem: Towards Scalable Mathematical Discovery Paper • 2608.16977 • Published 4 days ago • 3
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published 8 days ago • 83
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 7 days ago • 28
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 6 days ago • 428
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents Paper • 2608.15008 • Published 6 days ago • 13
Personalized Auto-Research: Towards a True AI Co-Scientist Paper • 2608.14881 • Published 7 days ago • 4
Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection Paper • 2608.16393 • Published 3 days ago • 5
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 3 days ago • 96
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published 3 days ago • 55
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published 4 days ago • 67
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published 4 days ago • 15
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 21 days ago • 14
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Paper • 2608.13430 • Published 8 days ago • 12
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published 8 days ago • 51
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Paper • 2608.12743 • Published 8 days ago • 42
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published 11 days ago • 48
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 8 days ago • 66
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published 21 days ago • 106