view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • 24 days ago • 155
Running Repro: XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation 🎯 Collaborate on a research logbook with an AI coding agent
Running Repro: XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation 🎯 Collaborate on a research logbook with an AI coding agent
SearchLM Collection NL2BM25: teaching Qwen2.5-3B to generate Tantivy boolean queries via SFT + GRPO. Covers reward hacking (GRPO v1) and the shaped-reward fix (GRPO v2). • 4 items • Updated Jun 27