Rubric-based RL that normalizes judged quality over only the responses satisfying the hard constraints.
🛵 Should they come looking for me, I inten
Aman Behera
beingamanforever
·
AI & ML interests
Long Horizon Agentic RL, OPD, Generative Engine Optimization, Performance Optimisation
Recent Activity
liked a model 5 days ago
infly/Infinity-Parser2-Pro liked a model 6 days ago
FireRedTeam/FireRed-OCR liked a dataset 6 days ago
allenai/olmOCR-benchOrganizations
None yet