Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments Paper • 2608.24099 • Published 4 days ago • 12
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 16 days ago • 46
Can Large Language Models Execute Parent Orders? Paper • 2607.28410 • Published about 1 month ago • 18
Can Large Language Models Execute Parent Orders? Paper • 2607.28410 • Published about 1 month ago • 18
Can Large Language Models Execute Parent Orders? Paper • 2607.28410 • Published about 1 month ago • 18
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism Paper • 2606.07512 • Published Jun 5 • 40
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism Paper • 2606.07512 • Published Jun 5 • 40
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism Paper • 2606.07512 • Published Jun 5 • 40
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Paper • 2604.03307 • Published Mar 31 • 14
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Paper • 2602.06422 • Published Feb 6 • 47
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors Paper • 2512.16915 • Published Dec 18, 2025 • 38
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention Paper • 2510.13940 • Published Oct 15, 2025 • 7
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention Paper • 2510.13940 • Published Oct 15, 2025 • 7