UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published 23 days ago • 86
On-Policy Self-Distillation in Diffusion Models Paper • 2608.24646 • Published 25 days ago • 69
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published Aug 13 • 55
V-RAE: Rethinking Video Latent Spaces for Generation Paper • 2608.13556 • Published Aug 13 • 30
HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Paper • 2607.15255 • Published Jul 16 • 17
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist Paper • 2511.08521 • Published Nov 11, 2025 • 39
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding Paper • 2509.11866 • Published Sep 15, 2025 • 2
V-RAE: Rethinking Video Latent Spaces for Generation Paper • 2608.13556 • Published Aug 13 • 30
V-RAE: Rethinking Video Latent Spaces for Generation Paper • 2608.13556 • Published Aug 13 • 30
SemanticGen: Video Generation in Semantic Space Paper • 2512.20619 • Published Dec 23, 2025 • 95
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder Paper • 2512.11749 • Published Dec 12, 2025 • 39
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist Paper • 2511.08521 • Published Nov 11, 2025 • 39
Latent Diffusion Model without Variational Autoencoder Paper • 2510.15301 • Published Oct 17, 2025 • 50