One Wordle game packaged for OpenEnv, ORS and NeMo Gym, from FineEnvs' 00-environments-101: the same env, three RL frameworks side by side.
AI & ML interests
Open RL Environments at Scale
Recent Activity
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 38 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 2.02k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 1.09k -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 45 โข 1
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐249Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐41Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
-
FineEnvs/MiMo-V2.6-RL-harbor-code
RL Environment โข Updated โข 3.88k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-cyber
RL Environment โข Updated โข 1.21k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-general
RL Environment โข Updated โข 1.33k -
FineEnvs/MiMo-V2.6-RL-harbor-terminal
RL Environment โข Updated โข 1.38k
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 249 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 72 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 82
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
One Wordle game packaged for OpenEnv, ORS and NeMo Gym, from FineEnvs' 00-environments-101: the same env, three RL frameworks side by side.
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
-
FineEnvs/MiMo-V2.6-RL-harbor-code
RL Environment โข Updated โข 3.88k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-cyber
RL Environment โข Updated โข 1.21k โข 2 -
FineEnvs/MiMo-V2.6-RL-harbor-general
RL Environment โข Updated โข 1.33k -
FineEnvs/MiMo-V2.6-RL-harbor-terminal
RL Environment โข Updated โข 1.38k
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
-
Geoguesser Environment
๐1Interact with a GeoGuessrโstyle environment through custom actions
-
FineEnvs/geoguesser-tasks
RL Environment โข Updated โข 3.65k โข 249 -
FineEnvs/geoguesser-qwen3.5-4b-grpo
Image-Text-to-Text โข Updated โข 72 -
FineEnvs/geoguesser-qwen3.5-4b-grpo-v3
Image-Text-to-Text โข Updated โข 82
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
FineEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 38 โข 1 -
FineEnvs/watercolour-reference-pool
RL Environment โข Updated โข 178 โข 2.02k โข 2 -
FineEnvs/watercolour-rollouts-hps-only
RL Environment โข Updated โข 470 โข 1.09k -
FineEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated โข 45 โข 1
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
-
The ultimate guide to RL environments: building and scaling them in the LLM era
๐249Building and scaling RL environments for LLM training
-
RL Environments 101 โ Slides
๐41Explore RL Environments 101 slide presentation
-
How to turn a game into an RL environment
๐68From an idea to a trained 4B, with the dead ends left in