Environment Specifications¶
Environment sections use the same schema classes as Arena, re-exported here.
Construction lives beside them as functions: make_env()
dispatches to make_gym_env(),
make_pz_env(),
make_bandit_env(), or
make_llm_env(). PettingZoo runs use
GymEnvSpec with make_pz_env().
- class agilerl.arena.models.env.GymEnvSpec(*, num_envs: Annotated[int, Ge(ge=1)] = 32, version: str | int | None = None, env_type: Literal['gym'] = 'gym', name: Annotated[str, MinLen(min_length=1)], custom: bool = False, default_type: str = 'float32', entrypoint: str | None = None, factory: str | None = None, path: str | None = None, env_config: dict[str, Any] | str | None = None, env_wrappers: list[str | tuple[str, dict[str, Any]]] | None = None, sync: bool = False)¶
A gymnasium / PettingZoo environment, or a custom one from an entrypoint.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.env.OfflineEnvSpec(*, num_envs: Annotated[int, Ge(ge=1)] = 32, version: str | int | None = None, env_type: Literal['offline'] = 'offline', name: Annotated[str, MinLen(min_length=1)], custom: bool = False, default_type: str = 'float32', entrypoint: str | None = None, factory: str | None = None, path: str | None = None, env_config: dict[str, Any] | str | None = None, env_wrappers: list[str | tuple[str, dict[str, Any]]] | None = None, sync: bool = False, minari_dataset_id: str | None = None, dataset_path: str | None = None, remote: bool = False)¶
A gym environment whose transitions come from a fixed dataset.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.env.LLMEnvSpec(*, num_envs: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, version: str | int | None = None, env_type: ~typing.Literal['rollout', 'dataset'], objective: ~typing.Literal['preference', 'sft'] | None = None, dataset: str | None = None, dataset_path: str | None = None, hf_dataset_id: str | None = None, columns: dict[str, str] | None = None, response_column: str | None = None, prompt_template: dict[str, ~typing.Any] | None = None, rubric_file_path: str | None = None, rubric_name: ~types.Annotated[str | None, ~annotated_types.MinLen(min_length=1)] = None, max_reward: float | None = None, train_test_split: ~types.Annotated[float | None, ~annotated_types.Ge(ge=0.0), ~annotated_types.Le(le=1.0)] = None, entrypoint: str | None = None, factory: str | None = None, env_config: dict[str, ~typing.Any] | None = None, name: str | None = None, env_packages: dict[str, ~typing.Any] | None = None, max_turns: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, segment_prompt_tokens: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, segment_max_images: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, restart_keep_turns: ~typing.Annotated[int, ~annotated_types.Ge(ge=0)] = 0, action_error_field: str = '', strict_chat_template_boundary: bool | None = None, observation_field: str | None = None, observation_processor: str | None = None, env_url: str | list[str] | None = None, mcp_tool: str | None = None, action_field: str | None = None, request_timeout_s: ~types.Annotated[float | None, ~annotated_types.Ge(ge=0.0)] = None, max_concurrent_resets: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, adaptive_task_sampling: bool = False, env_hosts: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, env_image: str | None = None, env_port: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1), ~annotated_types.Le(le=65535)] = None, env_sessions_per_host: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, env_session_ports: dict[str, int] = <factory>, cpus_per_env_host: ~types.Annotated[float | None, ~annotated_types.Gt(gt=0.0)] = None, env_host_memory_bytes: ~types.Annotated[int | None, ~annotated_types.Gt(gt=0)] = None, env_host_memory_limit_bytes: ~types.Annotated[int | None, ~annotated_types.Gt(gt=0)] = None, env_host_ready_timeout_s: ~types.Annotated[float | None, ~annotated_types.Gt(gt=0.0)] = None, env_services: list[~agilerl.arena.models.env.EnvServiceSpec] = <factory>, env_host_resource: str | None = None, env_vars: dict[str, str] = <factory>, chat_template_kwargs: dict[str, ~typing.Any] = <factory>, chat_template_path: str | None = None, apply_chat_template: bool = True)¶
The dataset or interactive environment an LLM algorithm fine-tunes against.
A
rolloutenv is generative and comes from exactly one source:dataset rows plus a rubric — a single-turn env over labelled rows;
entrypoint— a callable returning a text env, run where training runs (or on env hosts under the hosting fields, which need the Ray runtime);env_url— an OpenEnv service someone else hosts;env_image— a prebuilt env image the Ray runtime runs as its own Pod.
A
datasetenv is teacher-forced over the rows;objectivepicks the loss.- property dataset_backed_rollout: bool¶
Whether this rollout env serves labelled rows scored by a rubric.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.env.BanditEnvSpec(*, num_envs: Annotated[int, Ge(ge=1)] = 1, version: str | int | None = None, env_type: Literal['bandit'] = 'bandit', name: str = 'BanditEnv', features: str | None = None, targets: str | None = None, entrypoint: str | None = None, path: str | None = None, env_config: dict[str, Any] | None = None)¶
A tabular dataset wrapped as a contextual-bandit environment.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- agilerl.models.env.make_env(spec: EnvSpec, *, multi_agent: bool = False, extra_wrappers: list[type] | None = None, tokenizer: PreTrainedTokenizerBase | None = None, features: pd.DataFrame | str | Path | None = None, targets: pd.DataFrame | str | Path | None = None, seed: int | None = None) GymEnvType | AsyncPettingZooVecEnv | BanditEnvProtocol | DatasetEnv¶
Build the live environment described by spec.
- agilerl.models.env.make_gym_env(spec: GymEnvSpec, extra_wrappers: list[type] | None = None, wrappers: Sequence[tuple[Any, dict[str, Any]] | str | Callable[[...], Any]] | None = None) AsyncVectorEnv | SyncVectorEnv¶
Instantiate the vectorized gym environment.
- agilerl.models.env.make_pz_env(spec: GymEnvSpec, extra_wrappers: list[type] | None = None, wrappers: Sequence[tuple[Any, dict[str, Any]] | str | Callable[[...], Any]] | None = None) AsyncPettingZooVecEnv¶
Instantiate vectorized PettingZoo environments.
- agilerl.models.env.make_bandit_env(spec: BanditEnvSpec, *, features: DataFrame | str | Path | None = None, targets: DataFrame | str | Path | None = None) BanditEnvProtocol¶
Construct a bandit environment from a spec, or from in-memory tables.
- agilerl.models.env.make_llm_env(spec: LLMEnvSpec, tokenizer: PreTrainedTokenizerBase, *, data_batch_size_per_gpu: int = 8, max_context_length: int | None = None, seed: int | None = None, rank: int = 0, world_size: int = 1) DatasetEnv¶
Build the teacher-forced dataset environment for an LLM spec.
Rollout envs are built per-trajectory instead: use
make_rollout_env_factory().- Parameters:
spec (LLMEnvSpec) – The env spec, with
env_type="dataset".tokenizer (PreTrainedTokenizerBase) – The tokenizer.
data_batch_size_per_gpu (int) – Rows each rank draws per step.
max_context_length (int | None) – Token budget a row is truncated to.
seed (int | None) – Seed for the dataset split and the env’s row order.
rank (int) – This process’s data-parallel shard index.
world_size (int) – Number of data-parallel shards.
- Returns:
The dataset environment.
- Return type:
- agilerl.models.env.make_rollout_env_factory(spec: LLMEnvSpec, tokenizer: PreTrainedTokenizerBase, *, max_model_len: int | None = None, max_output_tokens: int | None = None, seed: int | None = None) tuple[Callable[[], RolloutHarness], int]¶
Build a factory that creates fresh
RolloutHarnessinstances.Each call to the returned factory creates an independent env, so concurrent trajectories never share state. Which of the three builders below runs is decided by the one source the spec names — dataset rows, an
env_url, or anentrypoint.env_imageneeds an orchestrator that can place Pods, which this process is not.The resolved turn budget rides along because only the builder can know it: an entrypoint env that does not declare
max_turnsin the manifest is probed for one.- Parameters:
spec (LLMEnvSpec) – The env spec, with
env_type="rollout".tokenizer (PreTrainedTokenizerBase) – The tokenizer (shared across all instances).
max_model_len (int | None) – Maximum model context length for prompt truncation.
max_output_tokens (int | None) – Maximum newly generated tokens per turn.
seed (int | None) – Seed for a dataset-backed rollout’s train/test split.
- Returns:
A zero-argument callable that creates a
RolloutHarness, and the rollout’s turn budget.- Return type:
tuple[Callable[[], RolloutHarness], int]
- agilerl.models.env.make_single_env(spec: GymEnvSpec, *, multi_agent: bool = False) Env | ParallelEnv¶
Create a single (non-vectorized) environment instance.