CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies Paper • 2609.24118 • Published 2 days ago • 11
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 6 days ago • 41
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Paper • 2608.25864 • Published 28 days ago • 9
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Paper • 2608.15265 • Published Aug 15 • 60
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective Paper • 2509.18905 • Published Sep 23, 2025 • 31