The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents Paper • 2608.24358 • Published 7 days ago • 13
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 7 days ago • 143
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published 27 days ago • 68
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents Paper • 2605.05191 • Published May 6
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards Paper • 2506.07736 • Published Jun 9, 2025
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Paper • 2410.07095 • Published Oct 9, 2024 • 9
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published 27 days ago • 68
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 210
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data Paper • 2603.15594 • Published Mar 16 • 150