Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild Paper • 2605.22064 • Published May 21 • 5
MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning Paper • 2602.03320 • Published Feb 3 • 2
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering Paper • 2602.23952 • Published Feb 27 • 6
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining Paper • 2609.35457 • Published 9 days ago • 66
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding Paper • 2603.10863 • Published Jun 26
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining Paper • 2609.35457 • Published 9 days ago • 66
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering Paper • 2510.14605 • Published Oct 16, 2025 • 8
Taming Modality Entanglement in Continual Audio-Visual Segmentation Paper • 2510.17234 • Published Oct 20, 2025 • 5
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Paper • 2508.21113 • Published Aug 28, 2025 • 111
Continuous Speculative Decoding for Autoregressive Image Generation Paper • 2411.11925 • Published Nov 18, 2024 • 16
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought Paper • 2505.15431 • Published May 21, 2025 • 2
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger Paper • 2506.07785 • Published Jun 9, 2025 • 3
Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis Paper • 2409.06135 • Published Sep 10, 2024 • 16
AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual Segmentation Paper • 2408.01708 • Published Aug 3, 2024 • 4
Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation Paper • 2312.06462 • Published Dec 11, 2023