OpenEnv documentation
Learn
Get Started
Concepts
Deploying an EnvAuto-DiscoveryCatalog DiscoveryRewardsCore ConceptsMCP EnvironmentsRubrics
Advanced Concepts
Learn
Basics
Training
OverviewTrain a Reasoning ModelRL Training with TRLRL Training with UnslothSFT Training with Environments
Harnesses
OverviewWhite-box: Train a Web Agent (BrowserGym)Black-box: Train Real Agents (Harbor)Agent in Your Env: Evaluate Claude CodeOpenCode (deprecated)Pi (deprecated)
Evals
Environments
API Reference
Project
You are viewing main version, which requires installation from source. If you'd like
regular pip install, checkout the latest stable version (v0.8.0).
Learn
Choose a learning path from the sidebar:
- Basics: start with Hello World and build Your First Environment.
- Training: the ways to train and the supported frameworks, then train a reasoning model, use TRL or Unsloth, or collect rollouts for SFT.
- Harnesses: pick a path: train with the trainer’s own loop on BrowserGym (white-box), train real agents through Harbor (black-box), or evaluate Claude Code inside an environment. The OpenCode and Pi tutorials are deprecated.
- Evals: follow Evaluating with Environments.
For MCP environments and rubrics, see Concepts in the sidebar.
New to OpenEnv? Start Here
The Getting Started Series walks you from zero to deploying your own environment in five short parts. No GPU required.
| Part | What it covers | Notebook |
|---|---|---|
| 1 — Introduction & Quick Start | What OpenEnv is, why it exists, and your first environment in under 10 minutes | |
| 2 — Using Environments | Connect to environments, create policies, run evaluations | |
| 3 — Building Environments | Create a custom environment from scratch | |
| 4 — Packaging & Deploying | Package with Docker and deploy to Hugging Face | — |
| 5 — Contributing to Hugging Face | Publish, fork, and share environments on the Hub | — |
Topic Tutorials
Already familiar with the basics? These tutorials cover specific workflows in depth.
| Tutorial | What it covers | GPU | Notebook |
|---|---|---|---|
| OpenEnv Tutorial | Full introduction to OpenEnv: install, connect to a hosted environment, step through an episode, define a reward function, and run a basic training loop. | No | |
| End-to-end walkthrough | The full pipeline: connect to reasoning_gym, wire it into TRL via environment_factory, fine-tune with GRPO, and push the checkpoint to the Hub. | Yes | |
| Building and using MCP environments | Consume and build MCP-backed environments: list and call tools through step(), register Python functions as tools with FastMCP. | No | |
| Rubrics | Compose reward functions from reusable pieces using Gate, WeightedSum, LLMJudge, and TrajectoryRubric. | No | |
| Wordle GRPO | Train an agent to play Wordle using GRPO via TRL’s environment_factory. | Yes | |
| RL Training with 2048 | Train a language model to play 2048 using GRPO. Covers game-state representation and reward shaping. | Yes | — |
| Evaluating with Environments | Wrap an OpenEnv environment in an Inspect AI Task, run it via InspectAIHarness, and get a structured EvalResult. | No | |
| White-box: Train a Web Agent on BrowserGym | Train a vision-language model on BrowserGym web tasks with TRL’s GRPOTrainer and environment_factory, so TRL runs the multi-turn tool loop. | Yes | — |
| Black-box: Train Real Agents with Harbor | Run real agents (OpenCode, Claude Code, Codex, …) on Harbor tasks with their token ids and logprobs captured, and train the model behind them with TRL’s AsyncGRPOTrainer. | Yes | — |
| Evaluate Claude Code in an Environment | Run Claude Code inside an OpenEnv environment with HarnessEnvironment (RFC 005), inject the environment’s tools over MCP, and evaluate it on τ²-bench against a simulated customer. Also serves it in production mode. | No | — |
| Training a Real Coding Agent (deprecated) | Deprecated, removed in OpenEnv 0.9.0: use Harbor with harness="opencode". Train the actual OpenCode agent (black-box, loop-owning) with TRL’s AsyncGRPOTrainer: a transparent proxy captures each turn’s token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| Collecting rollouts for supervised training | Run a teacher model to collect reward-labeled rollouts, filter them, and fine-tune a student with TRL’s SFTTrainer as a warm-start for GRPO. | Yes | |
| Training a Real Coding Agent (Pi, deprecated) | Deprecated, removed in OpenEnv 0.9.0: use Harbor with harness="pi". Train the actual Pi agent (black-box, loop-owning) with TRL’s AsyncGRPOTrainer: a transparent proxy captures each turn’s token ids and logprobs while the agent runs its own tool loop. | Yes | — |