Models from the paper "LaSeR: Reinforcement Learning with Last-Token Self-Rewarding"
Wenkai Yang
Keven16
AI & ML interests
None yet
Recent Activity
upvoted a paper about 4 hours ago
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression upvoted a paper 3 months ago
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation authored a paper 3 months ago
Rethinking Continual Experience Internalization for Self-Evolving LLM AgentsOrganizations
None yet