Running Featured 508 Jev Decision Index 🔬 508 Benchmarks and news on various repros of TypeSafe's Jev
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks Paper • 2610.08048 • Published 2 days ago • 9
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks Paper • 2610.08048 • Published 2 days ago • 9
ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios Paper • 2601.08620 • Published Jan 13 • 12
Experiential Reflective Learning for Self-Improving LLM Agents Paper • 2603.24639 • Published Mar 25 • 3
Experiential Reflective Learning for Self-Improving LLM Agents Paper • 2603.24639 • Published Mar 25 • 3
view article Article A framework and leaderboard for Retrieval Pipelines evaluation on ViDoRe v3 antoineedy • Feb 27 • 13
view article Article Small Yet Mighty: Improve Accuracy In Multimodal Search and Visual Document Retrieval with Llama Nemotron RAG Models nvidia • Jan 6 • 31
view article Article Nemotron ColEmbed V2: Raising the Bar for Multimodal Retrieval with ViDoRe V3’s Top Model nvidia • Feb 4 • 29
ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios Paper • 2601.08620 • Published Jan 13 • 12
ViDoRe Benchmark V3 Collection ViDoRe V3 is our latest benchmark, engineered to set a new industry gold standard for multi-modal, enterprise document retrieval evaluation. • 8 items • Updated Jan 14 • 27