Training & Replay Buffer Specs¶
TrainingSpec and
ReplayBufferSpec are the same schema classes
Arena uses. Buffer construction lives beside them as
init_buffer() and
init_n_step_buffer().
- class agilerl.arena.models.training.TrainingSpec(*, name: str | None = None, max_steps: Annotated[int, Ge(ge=1)] = 1000000, evo_steps: Annotated[int | None, Ge(ge=1)] = None, pop_size: Annotated[int, Ge(ge=1)] = 1, eval_steps: Annotated[int | None, Ge(ge=1)] = None, eval_loop: Annotated[int, Ge(ge=1)] = 1, hpo: bool = False, target_score: float | None = None, learning_delay: Annotated[int | None, Ge(ge=0)] = None, eps_start: float | None = None, eps_end: float | None = None, eps_decay: float | None = None, checkpoint_steps: Annotated[int | None, Ge(ge=1)] = None, checkpoint_path: str | None = None, overwrite_checkpoints: bool = False, checkpoint_optimizer: bool = True, evaluation_interval: Annotated[int | None, Ge(ge=1)] = None, eval_samples_per_task: Annotated[int, Ge(ge=0)] = 0, eval_greedy: bool = True, eval_max_concurrent_episodes: Annotated[int | None, Ge(ge=1)] = None, num_epochs: Annotated[int | None, Ge(ge=1)] = None, evo_epochs: Annotated[int | None, Ge(ge=1)] = None, max_wall_seconds: Annotated[float | None, Gt(gt=0)] = None, episode_steps: Annotated[int | None, Ge(ge=1)] = None, sum_scores: bool | None = None, reporting_interval: Annotated[int, Ge(ge=1)] = 1024, experience_sharing: bool | None = None, env_resource_ratio: Annotated[float | None, Gt(gt=0.0), Le(le=1.0)] = None, rollout_mode: Literal['colocated', 'async'] = 'colocated', rollout_batch_size: Annotated[int | None, Ge(ge=1)] = None, rollout_engines_per_agent: int | Literal['auto'] = 0, training_gpus_per_agent: Annotated[int, Ge(ge=1)] = 1, weight_sync_backend: Literal['filesystem', 'nccl'] | None = None, reuse_prefix_cache_across_syncs: bool = False, weight_sync_interval: Annotated[int | None, Ge(ge=1)] = None, rollout_version_stamp: Literal['publish', 'oldest_turn'] | None = None, checkpoint_export: CheckpointExportSpec | None = None)¶
The manifest’s
trainingsection.- effective_checkpoint_export() CheckpointExportSpec¶
LLM checkpoint export defaults when the training section omits it.
- effective_rollout_version_stamp() Literal['publish', 'oldest_turn'] | None¶
Async rollout stamping when the training section omits it.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.training.ReplayBufferSpec(*, kind: ~typing.Literal['classic'] = 'classic', name: str | None = None, max_size: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 100000, standard_buffer: bool = True, n_step_buffer: bool = False, n_step_buffer_args: ~agilerl.arena.models.training.NStepBufferArgs = <factory>, per_buffer: bool = False, per_buffer_args: ~agilerl.arena.models.training.PerBufferArgs = <factory>)¶
Classic branch of the
replay_buffersection.- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- agilerl.models.training.init_buffer(spec: ReplayBufferSpec, algo_spec: AlgoSpec, device: str | torch.device = 'cpu') BufferType¶
Initialize the replay buffer described by spec.
- agilerl.models.training.init_n_step_buffer(spec: ReplayBufferSpec, algo_spec: AlgoSpec, device: str | torch.device = 'cpu') BufferType | None¶
Initialize the n-step replay buffer for combined PER + n-step setups.
Returns
Noneunless bothper_bufferandn_step_bufferareTrue.