arxiv:2606.14885
Tom Lu
eigentom
AI & ML interests
MLLM, Reinforcement Learning, Agentic RL
Recent Activity
upvoted a paper 2 days ago
MemTrain: Self-Supervised Context Memory Training upvoted a paper 3 days ago
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement published a model 21 days ago
eigentom/qwen35-4b-dci-rl-rewardv2-step10