MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Paper • 2607.25948 • Published 16 days ago • 18
Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 16 days ago • 17
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Paper • 2607.24743 • Published 17 days ago • 11
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published 17 days ago • 77
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 21 days ago • 16
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion Paper • 2607.20417 • Published 22 days ago • 9
Self-Improvements in Modern Agentic Systems: A Survey Paper • 2607.13104 • Published 30 days ago • 33
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Paper • 2607.12752 • Published 29 days ago • 20
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors Paper • 2607.12000 • Published about 1 month ago • 40
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 87