SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 7 days ago • 129
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 10 days ago • 212
Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 24 days ago • 167
WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published 24 days ago • 10
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 24 days ago • 42
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 23 days ago • 220
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 29 days ago • 174
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published about 1 month ago • 34