Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Paper • 2608.19567 • Published 11 days ago • 32
Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Paper • 2608.19567 • Published 11 days ago • 32
PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle Paper • 2608.01354 • Published 28 days ago
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation Paper • 2605.26525 • Published May 26
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning Paper • 2605.18109 • Published May 18
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction Paper • 2605.26115 • Published May 25 • 54
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction Paper • 2605.26115 • Published May 25 • 54
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation Paper • 2605.09430 • Published May 10 • 2
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization Paper • 2605.15980 • Published May 15 • 36
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Paper • 2604.24764 • Published Apr 27 • 120
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Paper • 2604.24764 • Published Apr 27 • 120
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective Paper • 2604.14025 • Published Apr 15 • 15
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective Paper • 2604.14025 • Published Apr 15 • 15
Less Detail, Better Answers: Degradation-Driven Prompting for VQA Paper • 2604.04838 • Published Apr 6 • 13
Less Detail, Better Answers: Degradation-Driven Prompting for VQA Paper • 2604.04838 • Published Apr 6 • 13