TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 10 days ago • 139
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 9 days ago • 302
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 12 days ago • 35
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 21 days ago • 194
view article Article The OlmoEarth Platform: Geospatial inference at planetary scale allenai • 10 days ago • 39
view article Article NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics nvidia • 12 days ago • 57
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 18 days ago • 77