CLI-agent trajectories on ALE-Bench (arXiv 2506.09050): ahc026+ahc039 lite, raw streams + judged submissions + final results.
AI & ML interests
None defined yet.
Recent Activity
View all activity
Phase-2 evals: agents playing from token-truncated ARAs, one repo per game.
Live Hard Task Bench trajectories, one repo per model.
-
AgentNativeResearchLab/lhtb-gemini3.1-pro-high-trajectories
Preview • Updated • 109 -
AgentNativeResearchLab/lhtb-glm5.2-trajectories
Preview • Updated • 220 -
AgentNativeResearchLab/lhtb-grok4.5-trajectories
Preview • Updated • 161 -
AgentNativeResearchLab/lhtb-kimi-k3-trajectories
Preview • Updated • 191
CLI-agent trajectories on ALE-Bench (arXiv 2506.09050): ahc026+ahc039 lite, raw streams + judged submissions + final results.
Full record per harness×model×game: trajectories, recordings, replays + the ARA the agent built live. Same game across models = comparison unit.
Phase-2 evals: agents playing from token-truncated ARAs, one repo per game.
ARAs from agents solving DiscoverPhysics, one repo per model.
Live Hard Task Bench trajectories, one repo per model.
-
AgentNativeResearchLab/lhtb-gemini3.1-pro-high-trajectories
Preview • Updated • 109 -
AgentNativeResearchLab/lhtb-glm5.2-trajectories
Preview • Updated • 220 -
AgentNativeResearchLab/lhtb-grok4.5-trajectories
Preview • Updated • 161 -
AgentNativeResearchLab/lhtb-kimi-k3-trajectories
Preview • Updated • 191
Foundation-model open-problems trajectories, one repo per model.