Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 3 days ago • 1
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 3 days ago • 1
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing Paper • 2610.00313 • Published 3 days ago • 1
Rules to Tools: Executable Checks for LLM Agents in Scientific Computing Paper • 2610.00313 • Published 3 days ago • 1
ControlScope: Workflow Revision and Reliability in LLM Agents Paper • 2609.34313 • Published 4 days ago • 2
ControlScope: Workflow Revision and Reliability in LLM Agents Paper • 2609.34313 • Published 4 days ago • 2
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 9 days ago • 14
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 9 days ago • 14
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 25 days ago • 16
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 25 days ago • 16
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 25 days ago • 16
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published Aug 24 • 4
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes Paper • 2605.05724 • Published May 7 • 16
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks Paper • 2604.20087 • Published Apr 22 • 15
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks Paper • 2604.20087 • Published Apr 22 • 15
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines Paper • 2604.01029 • Published Apr 1 • 6