-
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
Paper β’ 2510.01132 β’ Published β’ 6 -
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Paper β’ 2604.13602 β’ Published β’ 31 -
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
Paper β’ 2410.05254 β’ Published β’ 85
Yashaswi Sharma
yashu2000
AI & ML interests
Large Language Models, Vision Models, Applied Models, MLOPS, Differentiable Economics, Differentiable Physics
Recent Activity
liked a model about 14 hours ago
convaiinnovations/laya updated a collection 26 days ago
Thesis Papers updated a collection 26 days ago
Thesis PapersOrganizations
None yet