Training & Replay Buffer Specs

TrainingSpec and ReplayBufferSpec are the same schema classes Arena uses. Buffer construction lives beside them as init_buffer() and init_n_step_buffer().

class agilerl.arena.models.training.TrainingSpec(*, name: str | None = None, max_steps: Annotated[int, Ge(ge=1)] = 1000000, evo_steps: Annotated[int | None, Ge(ge=1)] = None, pop_size: Annotated[int, Ge(ge=1)] = 1, eval_steps: Annotated[int | None, Ge(ge=1)] = None, eval_loop: Annotated[int, Ge(ge=1)] = 1, hpo: bool = False, target_score: float | None = None, learning_delay: Annotated[int | None, Ge(ge=0)] = None, eps_start: float | None = None, eps_end: float | None = None, eps_decay: float | None = None, checkpoint_steps: Annotated[int | None, Ge(ge=1)] = None, checkpoint_path: str | None = None, overwrite_checkpoints: bool = False, checkpoint_optimizer: bool = True, evaluation_interval: Annotated[int | None, Ge(ge=1)] = None, eval_samples_per_task: Annotated[int, Ge(ge=0)] = 0, eval_greedy: bool = True, eval_max_concurrent_episodes: Annotated[int | None, Ge(ge=1)] = None, num_epochs: Annotated[int | None, Ge(ge=1)] = None, evo_epochs: Annotated[int | None, Ge(ge=1)] = None, max_wall_seconds: Annotated[float | None, Gt(gt=0)] = None, episode_steps: Annotated[int | None, Ge(ge=1)] = None, sum_scores: bool | None = None, reporting_interval: Annotated[int, Ge(ge=1)] = 1024, experience_sharing: bool | None = None, env_resource_ratio: Annotated[float | None, Gt(gt=0.0), Le(le=1.0)] = None, rollout_mode: Literal['colocated', 'async'] = 'colocated', rollout_batch_size: Annotated[int | None, Ge(ge=1)] = None, rollout_engines_per_agent: int | Literal['auto'] = 0, training_gpus_per_agent: Annotated[int, Ge(ge=1)] = 1, weight_sync_backend: Literal['filesystem', 'nccl'] | None = None, reuse_prefix_cache_across_syncs: bool = False, weight_sync_interval: Annotated[int | None, Ge(ge=1)] = None, rollout_version_stamp: Literal['publish', 'oldest_turn'] | None = None, checkpoint_export: CheckpointExportSpec | None = None)

The manifest’s training section.

property async_rollout: bool

Whether generation runs on rollout engines rather than the trainer.

effective_checkpoint_export() → CheckpointExportSpec

LLM checkpoint export defaults when the training section omits it.

effective_rollout_version_stamp() → Literal['publish', 'oldest_turn'] | None

Async rollout stamping when the training section omits it.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class agilerl.arena.models.training.ReplayBufferSpec(*, kind: ~typing.Literal['classic'] = 'classic', name: str | None = None, max_size: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 100000, standard_buffer: bool = True, n_step_buffer: bool = False, n_step_buffer_args: ~agilerl.arena.models.training.NStepBufferArgs = <factory>, per_buffer: bool = False, per_buffer_args: ~agilerl.arena.models.training.PerBufferArgs = <factory>)

Classic branch of the replay_buffer section.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

agilerl.models.training.init_buffer(spec: ReplayBufferSpec, algo_spec: AlgoSpec, device: str | torch.device = 'cpu') → BufferType

Initialize the replay buffer described by spec.

agilerl.models.training.init_n_step_buffer(spec: ReplayBufferSpec, algo_spec: AlgoSpec, device: str | torch.device = 'cpu') → BufferType | None

Initialize the n-step replay buffer for combined PER + n-step setups.

Returns None unless both per_buffer and n_step_buffer are True.