CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 3 days ago • 21
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 3 days ago • 32
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 6 days ago • 35
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 7 days ago • 166
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 65
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos Paper • 2601.00393 • Published Jan 1 • 133
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper • 2512.04324 • Published Dec 3, 2025 • 160
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness Paper • 2504.01901 • Published Apr 2, 2025
PairUni: Pairwise Training for Unified Multimodal Language Models Paper • 2510.25682 • Published Oct 29, 2025 • 15
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs Paper • 2511.07250 • Published Nov 10, 2025 • 18
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs Paper • 2511.07250 • Published Nov 10, 2025 • 18