Algorithm Specifications

A training manifest is a YAML (or JSON) file that describes a full AgileRL run: which algorithm to train, on which environment, with which network and hyperparameters. You write it once. The trainer validates it before anything starts, so a typo or an unknown field fails immediately instead of an hour into training. The same file can run locally or on Arena. See Manifest Formulation for the file layout.

The algorithm section of that file is checked against a spec: a Pydantic model that lists the fields PPO, DQN, GRPO, and the others accept. Those models live in agilerl-arena. The framework uses them as-is: agilerl.models.PPOSpec is agilerl.arena.models.algorithms.PPOSpec. A manifest that validates on your laptop therefore validates when you submit it to Arena.

A spec is data. It does not construct a torch module or pick a training loop. agilerl.builders turn a spec into a live algorithm; agilerl.strategies pick the loop that trains it.

Base Specs

class agilerl.arena.models.algorithms.AlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64)

Fields every algorithm’s manifest section carries.

Runtime-only state (torch modules, resolved hyperparameter configs) is not declared here. Builders and strategies sit beside this spec; they do not subclass it to add construction.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

property name: str

The registry key this spec was resolved from.

class agilerl.arena.models.algorithms.SingleAgentAlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64, learn_step: Annotated[int, Ge(ge=1)] = 5, gamma: Annotated[float, Ge(ge=0.0), Le(le=1.0)] = 0.99, normalize_images: bool = True, net_config: NetworkSpec | None = None)

Single-agent reinforcement learning algorithms.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class agilerl.arena.models.algorithms.MultiAgentAlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64, learn_step: Annotated[int, Ge(ge=1)] = 2048, gamma: Annotated[float, Ge(ge=0.0), Le(le=1.0)] = 0.99, normalize_images: bool = True, torch_compiler: str | None = None)

Multi-agent reinforcement learning algorithms.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class agilerl.arena.models.algorithms.LLMAlgorithmSpec(*, batch_size: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 16, beta: ~typing.Annotated[float, ~annotated_types.Ge(ge=0.0), ~annotated_types.Le(le=1.0)] = 0.001, max_grad_norm: ~typing.Annotated[float, ~annotated_types.Ge(ge=0.0)] = 0.1, update_epochs: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 1, use_separate_reference_adapter: bool = True, calc_position_embeddings: bool = True, gradient_checkpointing: bool = True, use_liger_loss: bool = True, use_liger_kernel: bool | None = None, use_mamba_kernels: bool | None = None, cast_logprobs_to_fp32: bool = True, seed: int = 42, quantization: str | dict[str, ~typing.Any] | None = None, activation_offload: bool = False, moe_lora_recompute: bool | None = None, use_sequence_packing: bool = False, lora_target_scope: str | None = None, chunk_rows: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, attn_implementation: str | None = None, micro_batch_size_per_gpu: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, mini_batch_size: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, hf_generate_chunk_size: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, thinking_token_budget: ~types.Annotated[int | None, ~annotated_types.Ge(ge=0)] = None, constrain_answer_pattern: str | None = None, answer_continuation: bool = False, fsdp: ~types.Annotated[~agilerl.arena.models.fsdp.FSDPConfig | None, ~pydantic.json_schema.WithJsonSchema(json_schema={'anyOf': [{'type': 'null'}, {'type': 'boolean'}, {'additionalProperties': False, 'description': 'Settings for sharding the LLM actor with PyTorch FSDP2 (``fully_shard``).', 'properties': {'reshard_after_forward': {'default': True, 'description': "Free gathered parameters after each module's forward. False keeps them gathered: faster, more VRAM.", 'title': 'Reshard After Forward', 'type': 'boolean'}, 'cpu_offload': {'default': False, 'description': 'Offload sharded parameters and gradients to CPU. Needs colocated vLLM. Mutually exclusive with optim_cpu_offload.', 'title': 'Cpu Offload', 'type': 'boolean'}, 'optim_cpu_offload': {'default': True, 'description': 'Keep Adam state on CPU and move it to GPU only for step(). Parameters and gradients stay on GPU.', 'title': 'Optim Cpu Offload', 'type': 'boolean'}, 'defer_grad_sync': {'default': True, 'description': 'Reduce-scatter gradients only on the last micro-batch of an optimizer step. Saves communication; holds unsharded grads in between.', 'title': 'Defer Grad Sync', 'type': 'boolean'}, 'param_dtype': {'default': 'bfloat16', 'description': "Mixed-precision parameter dtype, e.g. 'bfloat16'. A torch.dtype is accepted and stored by name.", 'title': 'Param Dtype', 'type': 'string'}, 'reduce_dtype': {'default': 'float32', 'description': 'Dtype for gradient reduce-scatter and all-reduce.', 'title': 'Reduce Dtype', 'type': 'string'}, 'prefetch_units': {'default': 1, 'description': 'Neighbouring FSDP units to all-gather ahead during forward. 1 gathers unit i+1 while unit i runs.', 'minimum': 1, 'title': 'Prefetch Units', 'type': 'integer'}, 'backward_prefetch_units': {'default': 1, 'description': 'Neighbouring FSDP units to all-gather ahead during backward. 1 gathers unit i-1 while unit i runs backward.', 'minimum': 1, 'title': 'Backward Prefetch Units', 'type': 'integer'}, 'checkpoint_skip_layer_types': {'default': [], 'description': "Transformer block kinds left out of activation checkpointing. A block's kind is its block_type (hybrid models, e.g. 'linear_attention', 'full_attention', 'moe') or else its class name. Empty checkpoints every block. Applies when gradient_checkpointing is on.", 'items': {'type': 'string'}, 'title': 'Checkpoint Skip Layer Types', 'type': 'array'}, 'checkpoint_every_n_blocks': {'default': 1, 'description': 'Checkpoint one in every n blocks not skipped by checkpoint_skip_layer_types. 1 checkpoints all of them.', 'minimum': 1, 'title': 'Checkpoint Every N Blocks', 'type': 'integer'}, 'wrap_every_n_blocks': {'default': 1, 'description': 'Consecutive transformer blocks per FSDP unit. Larger values mean fewer, bigger collectives.', 'minimum': 1, 'title': 'Wrap Every N Blocks', 'type': 'integer'}, 'param_persistence_threshold': {'default': 100000, 'description': 'Parameters with fewer elements than this stay unsharded. 0 shards every parameter.', 'minimum': 0, 'title': 'Param Persistence Threshold', 'type': 'integer'}, 'ep': {'default': 1, 'description': 'Expert-parallel degree: packed MoE experts per layer are split across this many GPUs. 1 keeps data parallel plus FSDP sharding with no expert split.', 'minimum': 1, 'title': 'Ep', 'type': 'integer'}, 'ep_token_blocks': {'default': 1, 'description': "Token blocks per routed MoE layer under expert parallel. Each block runs dispatch, experts, and combine; on GPU the next block's all-to-all overlaps this block's experts. 1 moves every token in one all-to-all.", 'minimum': 1, 'title': 'Ep Token Blocks', 'type': 'integer'}, 'tp': {'default': 1, 'description': 'Tensor-parallel degree for dense layers. Ranks in one TP group share a batch shard. 1 gives every rank its own batch.', 'minimum': 1, 'title': 'Tp', 'type': 'integer'}, 'shard_group_size': {'anyOf': [{'minimum': 1, 'type': 'integer'}, {'type': 'null'}], 'default': None, 'description': 'Ranks that shard weights between them; groups replicate (HSDP). Set to GPUs per node so weight gathers stay inside a node. None shards across all trainer ranks.', 'title': 'Shard Group Size'}, 'compile_blocks': {'default': False, 'description': 'torch.compile the dense submodules of each transformer block (norms, MLPs) in place. MoE experts, routers, attention and Mamba mixers stay eager.', 'title': 'Compile Blocks', 'type': 'boolean'}, 'compile_backend': {'default': 'inductor', 'description': 'torch.compile backend for compile_blocks.', 'title': 'Compile Backend', 'type': 'string'}}, 'title': 'FSDPConfig', 'type': 'object'}]}, mode=None)] = None, vllm_engine_args: dict[str, ~typing.Any] = <factory>, vllm_lora_export_dir: str | None = None, pretrained_model_name_or_path: ~types.Annotated[str | None, ~annotated_types.MinLen(min_length=1)] = None, max_model_len: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 1024, lora_config: ~agilerl.arena.models.networks.LoraConfigDict | None = None)

LLM fine-tuning algorithms.

pretrained_model_name_or_path, max_model_len and lora_config are written under the manifest’s network section and lifted onto the algorithm when the manifest is validated.

model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

Registry

One registry maps names (for example "DQN") to spec classes. The local trainer, the manifest, and Arena all read it, so algorithm.name means the same thing everywhere.

class agilerl.arena.models.registry.AlgorithmRegistry

Maps the manifest’s algorithm.name to the spec class that validates it.

add(name: str, spec_cls: type[AlgoSpec]) → None

Register spec_cls under name.

Parameters:
  • name (str) – Algorithm name (e.g. "DQN").

  • spec_cls (type[AlgoSpec]) – The spec class to register.

create(name: str, /, **fields: Any) → AlgoSpec

Build a spec from the registry name and field values.

Parameters:
  • name (str) – Algorithm name.

  • fields – Field values for the spec.

Returns:

The spec instance.

Return type:

AlgoSpec

Raises:

KeyError – If name is not registered.

get(name: str) → type[AlgoSpec]

Look up a spec class by algorithm name.

Parameters:

name (str) – Algorithm name.

Returns:

The registered spec class.

Return type:

type[AlgoSpec]

Raises:

KeyError – If name is not registered.

items() → list[tuple[str, type[AlgoSpec]]]

Return every (name, spec class) pair, sorted by name.

Returns:

The registered entries.

Return type:

list[tuple[str, type[AlgoSpec]]]

names() → list[str]

Return every registered algorithm name, sorted.

Returns:

The registered algorithm names.

Return type:

list[str]

agilerl.arena.models.registry.MANIFEST_REGISTRY = <agilerl.arena.models.registry.AlgorithmRegistry object>

Maps the manifest’s algorithm.name to the spec class that validates it.

Builders and training strategies

A builder turns a spec into an algorithm object, including values the YAML cannot hold (a peft LoraConfig, the vLLM dataclass). A strategy chooses which training loop that paradigm uses and the keyword arguments the loop takes. Runtime-only inputs (a pre-built network, a resolved hyperparameter config) are arguments to build(), passed through from LocalTrainer. Both are looked up from the spec’s paradigm rather than declared on it, so the schema does not depend on framework classes.

Builders

agilerl.builders.select_builder(spec: AlgorithmSpec) → type[AlgorithmBuilder]

Return the builder class for spec’s paradigm.

Parameters:

spec (AlgorithmSpec) – The algorithm spec.

Returns:

The paradigm’s builder class.

Return type:

type[AlgorithmBuilder]

Raises:

TypeError – If spec is not one of the contract’s algorithm specs.

class agilerl.builders.AlgorithmBuilder

Paradigm-keyed factory that builds a live algorithm from a spec.

Concrete builders own the paradigm-specific build signature; the callers dispatch on paradigm and call the concrete class directly.

classmethod algo_class(spec: AlgorithmSpec) → type[EvolvableAlgorithm]

Resolve the algorithm class from agilerl.algorithms.

Naming convention: <Name>Spec -> <Name>. Walks the spec’s MRO so a user subclass still maps to the parent algorithm.

Parameters:

spec (AlgorithmSpec) – The algorithm spec.

Returns:

The algorithm class.

Return type:

type[EvolvableAlgorithm]

Raises:

AttributeError – If no algorithm matches the spec’s name.

class agilerl.builders.SingleAgentBuilder

Single-agent reinforcement learning.

classmethod build(spec: AlgorithmSpec, observation_space: spaces.Space | None = None, action_space: spaces.Space | None = None, *, runtime: AlgorithmBuildRuntime | None = None, **networks: Any) → SingleAgentAlgorithm

Build a single-agent algorithm.

Parameters:
  • spec (AlgorithmSpec) – The algorithm spec.

  • observation_space (spaces.Space | None) – Observation space.

  • action_space (spaces.Space | None) – Action space.

  • runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint.

  • networks (EvolvableModule) – Pre-built modules to hand the constructor, e.g. actor_network and critic_network. Only pass the ones the algorithm takes.

Returns:

Single-agent algorithm instance.

Return type:

SingleAgentAlgorithm

Raises:

ValueError – If observation_space, action_space, or index is None.

class agilerl.builders.MultiAgentBuilder

Multi-agent reinforcement learning.

classmethod build(spec: AlgorithmSpec, observation_spaces: dict[str, spaces.Space] | None = None, action_spaces: dict[str, spaces.Space] | None = None, *, runtime: AlgorithmBuildRuntime | None = None, **networks: Any) → MultiAgentAlgorithm

Build a multi-agent algorithm.

Parameters:
  • spec (AlgorithmSpec) – The algorithm spec.

  • observation_spaces (dict[str, spaces.Space] | None) – Per-agent observation spaces.

  • action_spaces (dict[str, spaces.Space] | None) – Per-agent action spaces.

  • runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint.

  • networks (ModuleDict) – Pre-built modules to hand the constructor, e.g. actor_networks and critic_networks.

Returns:

Multi-agent algorithm instance.

Return type:

MultiAgentAlgorithm

Raises:

ValueError – If observation_spaces, action_spaces, or index is None.

class agilerl.builders.LLMBuilder

LLM fine-tuning.

classmethod build(spec: AlgorithmSpec, *, tokenizer: PreTrainedTokenizerBase | None = None, runtime: AlgorithmBuildRuntime | None = None, actor_network: PreTrainedModel | PeftModel | None = None, rollout_mode: str = '') → LLMAlgorithm

Build an LLM algorithm.

Parameters:
  • spec (AlgorithmSpec) – The algorithm spec.

  • tokenizer (PreTrainedTokenizerBase | None) – A HuggingFace AutoTokenizer instance.

  • runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint. index defaults to 0 when omitted.

  • actor_network (PreTrainedModel | PeftModel | None) – Pre-built or cloned actor. When provided it is handed to the constructor instead of loading the model from pretrained_model_name_or_path.

  • rollout_mode (str) – training.rollout_mode. Colocated runs vLLM in the trainer process.

Returns:

LLM algorithm instance.

Return type:

LLMAlgorithm

Raises:

ValueError – If tokenizer is None.

Strategies

agilerl.strategies.select_strategy(spec: SingleAgentAlgorithmSpec | MultiAgentAlgorithmSpec | LLMAlgorithmSpec) → TrainingStrategy

Return the strategy that trains spec, from its paradigm flags.

The contract declares off_policy / offline / bandit on the RL specs and env_type on the LLM specs; a spec subclassed elsewhere inherits them, so it trains like its parent.

Parameters:

spec (AlgoSpec) – The algorithm spec.

Returns:

The paradigm’s strategy.

Return type:

TrainingStrategy

Raises:
  • TypeError – If spec is not one of the contract’s algorithm specs.

  • KeyError – If an LLM spec’s env_type has no strategy.

class agilerl.strategies.TrainingStrategy

Paradigm-keyed run-time orchestration for an algorithm spec.

Subclasses set default_loop and implement get_trainer_kwargs(); a paradigm with more than one loop overrides get_training_loop().

abstract get_trainer_kwargs(spec: AlgoSpec, *, training: TrainingSpec, env_spec: EnvSpecType, memory: BufferType | None = None, n_step_memory: BufferType | None = None) → dict[str, Any]

Return the extra keyword arguments the training loop takes.

Parameters:
  • spec (AlgoSpec) – The algorithm spec.

  • training (TrainingSpec) – Training specification.

  • env_spec (EnvSpecType) – Environment specification.

  • memory (BufferType | None) – Replay buffer instance.

  • n_step_memory (BufferType | None) – N-step replay buffer for combined PER + n-step setups.

Returns:

Extra keyword arguments for the training function.

Return type:

dict[str, Any]

get_training_loop(spec: AlgoSpec) → TrainingLoop

Select the training loop for spec.

Parameters:

spec (AlgoSpec) – The algorithm spec.

Returns:

The training function.

Return type:

TrainingLoop

Raises:

NotImplementedError – If the strategy names no loop.

class agilerl.strategies.SingleAgentOnPolicyStrategy

On-policy single-agent training (PPO).

class agilerl.strategies.SingleAgentOffPolicyStrategy

Off-policy single-agent training (DQN, Rainbow DQN, DDPG, TD3).

class agilerl.strategies.OfflineStrategy

Offline training from a fixed dataset (CQN).

class agilerl.strategies.BanditStrategy

Contextual bandit training (NeuralTS, NeuralUCB).

class agilerl.strategies.MultiAgentOnPolicyStrategy

On-policy multi-agent training (IPPO).

class agilerl.strategies.MultiAgentOffPolicyStrategy

Off-policy multi-agent training (MADDPG, MATD3).

class agilerl.strategies.LLMStrategy

Shared orchestration for the LLM fine-tuning loops.

class agilerl.strategies.LLMRolloutStrategy

Generative rollout fine-tuning (GRPO family, LLM PPO, LLM REINFORCE).

One loop for every rollout regime: single-turn reasoning is max_turns=1.

class agilerl.strategies.LLMDatasetStrategy

Teacher-forced fine-tuning over dataset rows (DPO and SFT).

One loop for both objectives; the env’s objective picks the loss.