CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 4 days ago • 22
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 4 days ago • 33
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 7 days ago • 36
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 8 days ago • 166
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 65
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos Paper • 2601.00393 • Published Jan 1 • 133
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper • 2512.04324 • Published Dec 3, 2025 • 160
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs Paper • 2511.07250 • Published Nov 10, 2025 • 18
Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence Paper • 2510.20579 • Published Oct 23, 2025 • 56
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs Paper • 2510.18876 • Published Oct 21, 2025 • 37
Grasp-Any-Region Collection Models and datasets for Grasp-Any-Region • 4 items • Updated Oct 22, 2025 • 3
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR Paper • 2508.14029 • Published Aug 19, 2025 • 119
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology Paper • 2507.07999 • Published Jul 10, 2025 • 51