Algorithm Specifications¶
A training manifest is a YAML (or JSON) file that describes a full AgileRL run: which algorithm to train, on which environment, with which network and hyperparameters. You write it once. The trainer validates it before anything starts, so a typo or an unknown field fails immediately instead of an hour into training. The same file can run locally or on Arena. See Manifest Formulation for the file layout.
The algorithm section of that file is checked against a spec: a
Pydantic model that lists the fields PPO, DQN, GRPO, and the others accept.
Those models live in agilerl-arena. The framework uses them as-is:
agilerl.models.PPOSpec is
agilerl.arena.models.algorithms.PPOSpec. A manifest that validates on
your laptop therefore validates when you submit it to Arena.
A spec is data. It does not construct a torch module or pick a training loop.
agilerl.builders turn a spec into a live algorithm;
agilerl.strategies pick the loop that trains it.
Base Specs¶
- class agilerl.arena.models.algorithms.AlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64)¶
Fields every algorithm’s manifest section carries.
Runtime-only state (torch modules, resolved hyperparameter configs) is not declared here. Builders and strategies sit beside this spec; they do not subclass it to add construction.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.algorithms.SingleAgentAlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64, learn_step: Annotated[int, Ge(ge=1)] = 5, gamma: Annotated[float, Ge(ge=0.0), Le(le=1.0)] = 0.99, normalize_images: bool = True, net_config: NetworkSpec | None = None)¶
Single-agent reinforcement learning algorithms.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.algorithms.MultiAgentAlgorithmSpec(*, batch_size: Annotated[int, Ge(ge=1)] = 64, learn_step: Annotated[int, Ge(ge=1)] = 2048, gamma: Annotated[float, Ge(ge=0.0), Le(le=1.0)] = 0.99, normalize_images: bool = True, torch_compiler: str | None = None)¶
Multi-agent reinforcement learning algorithms.
- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class agilerl.arena.models.algorithms.LLMAlgorithmSpec(*, batch_size: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 16, beta: ~typing.Annotated[float, ~annotated_types.Ge(ge=0.0), ~annotated_types.Le(le=1.0)] = 0.001, max_grad_norm: ~typing.Annotated[float, ~annotated_types.Ge(ge=0.0)] = 0.1, update_epochs: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 1, use_separate_reference_adapter: bool = True, calc_position_embeddings: bool = True, gradient_checkpointing: bool = True, use_liger_loss: bool = True, use_liger_kernel: bool | None = None, use_mamba_kernels: bool | None = None, cast_logprobs_to_fp32: bool = True, seed: int = 42, quantization: str | dict[str, ~typing.Any] | None = None, activation_offload: bool = False, moe_lora_recompute: bool | None = None, use_sequence_packing: bool = False, lora_target_scope: str | None = None, chunk_rows: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, attn_implementation: str | None = None, micro_batch_size_per_gpu: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, mini_batch_size: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, hf_generate_chunk_size: ~types.Annotated[int | None, ~annotated_types.Ge(ge=1)] = None, thinking_token_budget: ~types.Annotated[int | None, ~annotated_types.Ge(ge=0)] = None, constrain_answer_pattern: str | None = None, answer_continuation: bool = False, fsdp: ~types.Annotated[~agilerl.arena.models.fsdp.FSDPConfig | None, ~pydantic.json_schema.WithJsonSchema(json_schema={'anyOf': [{'type': 'null'}, {'type': 'boolean'}, {'additionalProperties': False, 'description': 'Settings for sharding the LLM actor with PyTorch FSDP2 (``fully_shard``).', 'properties': {'reshard_after_forward': {'default': True, 'description': "Free gathered parameters after each module's forward. False keeps them gathered: faster, more VRAM.", 'title': 'Reshard After Forward', 'type': 'boolean'}, 'cpu_offload': {'default': False, 'description': 'Offload sharded parameters and gradients to CPU. Needs colocated vLLM. Mutually exclusive with optim_cpu_offload.', 'title': 'Cpu Offload', 'type': 'boolean'}, 'optim_cpu_offload': {'default': True, 'description': 'Keep Adam state on CPU and move it to GPU only for step(). Parameters and gradients stay on GPU.', 'title': 'Optim Cpu Offload', 'type': 'boolean'}, 'defer_grad_sync': {'default': True, 'description': 'Reduce-scatter gradients only on the last micro-batch of an optimizer step. Saves communication; holds unsharded grads in between.', 'title': 'Defer Grad Sync', 'type': 'boolean'}, 'param_dtype': {'default': 'bfloat16', 'description': "Mixed-precision parameter dtype, e.g. 'bfloat16'. A torch.dtype is accepted and stored by name.", 'title': 'Param Dtype', 'type': 'string'}, 'reduce_dtype': {'default': 'float32', 'description': 'Dtype for gradient reduce-scatter and all-reduce.', 'title': 'Reduce Dtype', 'type': 'string'}, 'prefetch_units': {'default': 1, 'description': 'Neighbouring FSDP units to all-gather ahead during forward. 1 gathers unit i+1 while unit i runs.', 'minimum': 1, 'title': 'Prefetch Units', 'type': 'integer'}, 'backward_prefetch_units': {'default': 1, 'description': 'Neighbouring FSDP units to all-gather ahead during backward. 1 gathers unit i-1 while unit i runs backward.', 'minimum': 1, 'title': 'Backward Prefetch Units', 'type': 'integer'}, 'checkpoint_skip_layer_types': {'default': [], 'description': "Transformer block kinds left out of activation checkpointing. A block's kind is its block_type (hybrid models, e.g. 'linear_attention', 'full_attention', 'moe') or else its class name. Empty checkpoints every block. Applies when gradient_checkpointing is on.", 'items': {'type': 'string'}, 'title': 'Checkpoint Skip Layer Types', 'type': 'array'}, 'checkpoint_every_n_blocks': {'default': 1, 'description': 'Checkpoint one in every n blocks not skipped by checkpoint_skip_layer_types. 1 checkpoints all of them.', 'minimum': 1, 'title': 'Checkpoint Every N Blocks', 'type': 'integer'}, 'wrap_every_n_blocks': {'default': 1, 'description': 'Consecutive transformer blocks per FSDP unit. Larger values mean fewer, bigger collectives.', 'minimum': 1, 'title': 'Wrap Every N Blocks', 'type': 'integer'}, 'param_persistence_threshold': {'default': 100000, 'description': 'Parameters with fewer elements than this stay unsharded. 0 shards every parameter.', 'minimum': 0, 'title': 'Param Persistence Threshold', 'type': 'integer'}, 'ep': {'default': 1, 'description': 'Expert-parallel degree: packed MoE experts per layer are split across this many GPUs. 1 keeps data parallel plus FSDP sharding with no expert split.', 'minimum': 1, 'title': 'Ep', 'type': 'integer'}, 'ep_token_blocks': {'default': 1, 'description': "Token blocks per routed MoE layer under expert parallel. Each block runs dispatch, experts, and combine; on GPU the next block's all-to-all overlaps this block's experts. 1 moves every token in one all-to-all.", 'minimum': 1, 'title': 'Ep Token Blocks', 'type': 'integer'}, 'tp': {'default': 1, 'description': 'Tensor-parallel degree for dense layers. Ranks in one TP group share a batch shard. 1 gives every rank its own batch.', 'minimum': 1, 'title': 'Tp', 'type': 'integer'}, 'shard_group_size': {'anyOf': [{'minimum': 1, 'type': 'integer'}, {'type': 'null'}], 'default': None, 'description': 'Ranks that shard weights between them; groups replicate (HSDP). Set to GPUs per node so weight gathers stay inside a node. None shards across all trainer ranks.', 'title': 'Shard Group Size'}, 'compile_blocks': {'default': False, 'description': 'torch.compile the dense submodules of each transformer block (norms, MLPs) in place. MoE experts, routers, attention and Mamba mixers stay eager.', 'title': 'Compile Blocks', 'type': 'boolean'}, 'compile_backend': {'default': 'inductor', 'description': 'torch.compile backend for compile_blocks.', 'title': 'Compile Backend', 'type': 'string'}}, 'title': 'FSDPConfig', 'type': 'object'}]}, mode=None)] = None, vllm_engine_args: dict[str, ~typing.Any] = <factory>, vllm_lora_export_dir: str | None = None, pretrained_model_name_or_path: ~types.Annotated[str | None, ~annotated_types.MinLen(min_length=1)] = None, max_model_len: ~typing.Annotated[int, ~annotated_types.Ge(ge=1)] = 1024, lora_config: ~agilerl.arena.models.networks.LoraConfigDict | None = None)¶
LLM fine-tuning algorithms.
pretrained_model_name_or_path,max_model_lenandlora_configare written under the manifest’snetworksection and lifted onto the algorithm when the manifest is validated.- model_config: ClassVar[ConfigDict] = {'extra': 'forbid'}¶
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
Registry¶
One registry maps names (for example "DQN") to spec classes. The local
trainer, the manifest, and Arena all read it, so algorithm.name means the
same thing everywhere.
- class agilerl.arena.models.registry.AlgorithmRegistry¶
Maps the manifest’s
algorithm.nameto the spec class that validates it.
- agilerl.arena.models.registry.MANIFEST_REGISTRY = <agilerl.arena.models.registry.AlgorithmRegistry object>¶
Maps the manifest’s
algorithm.nameto the spec class that validates it.
Builders and training strategies¶
A builder turns a spec into an algorithm object, including values the YAML
cannot hold (a peft LoraConfig, the vLLM dataclass). A strategy
chooses which training loop that paradigm uses and the keyword arguments the
loop takes. Runtime-only inputs (a pre-built network, a resolved
hyperparameter config) are arguments to
build(), passed through from
LocalTrainer. Both are looked up from the
spec’s paradigm rather than declared on it, so the schema does not depend on
framework classes.
Builders¶
- agilerl.builders.select_builder(spec: AlgorithmSpec) type[AlgorithmBuilder]¶
Return the builder class for spec’s paradigm.
- Parameters:
spec (AlgorithmSpec) – The algorithm spec.
- Returns:
The paradigm’s builder class.
- Return type:
- Raises:
TypeError – If spec is not one of the contract’s algorithm specs.
- class agilerl.builders.AlgorithmBuilder¶
Paradigm-keyed factory that builds a live algorithm from a spec.
Concrete builders own the paradigm-specific
buildsignature; the callers dispatch on paradigm and call the concrete class directly.- classmethod algo_class(spec: AlgorithmSpec) type[EvolvableAlgorithm]¶
Resolve the algorithm class from
agilerl.algorithms.Naming convention:
<Name>Spec-><Name>. Walks the spec’s MRO so a user subclass still maps to the parent algorithm.- Parameters:
spec (AlgorithmSpec) – The algorithm spec.
- Returns:
The algorithm class.
- Return type:
- Raises:
AttributeError – If no algorithm matches the spec’s name.
- class agilerl.builders.SingleAgentBuilder¶
Single-agent reinforcement learning.
- classmethod build(spec: AlgorithmSpec, observation_space: spaces.Space | None = None, action_space: spaces.Space | None = None, *, runtime: AlgorithmBuildRuntime | None = None, **networks: Any) SingleAgentAlgorithm¶
Build a single-agent algorithm.
- Parameters:
spec (AlgorithmSpec) – The algorithm spec.
observation_space (spaces.Space | None) – Observation space.
action_space (spaces.Space | None) – Action space.
runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint.
networks (EvolvableModule) – Pre-built modules to hand the constructor, e.g.
actor_networkandcritic_network. Only pass the ones the algorithm takes.
- Returns:
Single-agent algorithm instance.
- Return type:
- Raises:
ValueError – If observation_space, action_space, or index is None.
- class agilerl.builders.MultiAgentBuilder¶
Multi-agent reinforcement learning.
- classmethod build(spec: AlgorithmSpec, observation_spaces: dict[str, spaces.Space] | None = None, action_spaces: dict[str, spaces.Space] | None = None, *, runtime: AlgorithmBuildRuntime | None = None, **networks: Any) MultiAgentAlgorithm¶
Build a multi-agent algorithm.
- Parameters:
spec (AlgorithmSpec) – The algorithm spec.
observation_spaces (dict[str, spaces.Space] | None) – Per-agent observation spaces.
action_spaces (dict[str, spaces.Space] | None) – Per-agent action spaces.
runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint.
networks (ModuleDict) – Pre-built modules to hand the constructor, e.g.
actor_networksandcritic_networks.
- Returns:
Multi-agent algorithm instance.
- Return type:
- Raises:
ValueError – If observation_spaces, action_spaces, or index is None.
- class agilerl.builders.LLMBuilder¶
LLM fine-tuning.
- classmethod build(spec: AlgorithmSpec, *, tokenizer: PreTrainedTokenizerBase | None = None, runtime: AlgorithmBuildRuntime | None = None, actor_network: PreTrainedModel | PeftModel | None = None, rollout_mode: str = '') LLMAlgorithm¶
Build an LLM algorithm.
- Parameters:
spec (AlgorithmSpec) – The algorithm spec.
tokenizer (PreTrainedTokenizerBase | None) – A HuggingFace
AutoTokenizerinstance.runtime (AlgorithmBuildRuntime | None) – Population slot, device, HPO, and optional checkpoint.
indexdefaults to 0 when omitted.actor_network (PreTrainedModel | PeftModel | None) – Pre-built or cloned actor. When provided it is handed to the constructor instead of loading the model from
pretrained_model_name_or_path.rollout_mode (str) –
training.rollout_mode. Colocated runs vLLM in the trainer process.
- Returns:
LLM algorithm instance.
- Return type:
LLMAlgorithm
- Raises:
ValueError – If tokenizer is None.
Strategies¶
- agilerl.strategies.select_strategy(spec: SingleAgentAlgorithmSpec | MultiAgentAlgorithmSpec | LLMAlgorithmSpec) TrainingStrategy¶
Return the strategy that trains spec, from its paradigm flags.
The contract declares
off_policy/offline/banditon the RL specs andenv_typeon the LLM specs; a spec subclassed elsewhere inherits them, so it trains like its parent.- Parameters:
spec (AlgoSpec) – The algorithm spec.
- Returns:
The paradigm’s strategy.
- Return type:
- Raises:
- class agilerl.strategies.TrainingStrategy¶
Paradigm-keyed run-time orchestration for an algorithm spec.
Subclasses set
default_loopand implementget_trainer_kwargs(); a paradigm with more than one loop overridesget_training_loop().- abstract get_trainer_kwargs(spec: AlgoSpec, *, training: TrainingSpec, env_spec: EnvSpecType, memory: BufferType | None = None, n_step_memory: BufferType | None = None) dict[str, Any]¶
Return the extra keyword arguments the training loop takes.
- Parameters:
spec (AlgoSpec) – The algorithm spec.
training (TrainingSpec) – Training specification.
env_spec (EnvSpecType) – Environment specification.
memory (BufferType | None) – Replay buffer instance.
n_step_memory (BufferType | None) – N-step replay buffer for combined PER + n-step setups.
- Returns:
Extra keyword arguments for the training function.
- Return type:
- get_training_loop(spec: AlgoSpec) TrainingLoop¶
Select the training loop for spec.
- Parameters:
spec (AlgoSpec) – The algorithm spec.
- Returns:
The training function.
- Return type:
TrainingLoop
- Raises:
NotImplementedError – If the strategy names no loop.
- class agilerl.strategies.SingleAgentOnPolicyStrategy¶
On-policy single-agent training (PPO).
- class agilerl.strategies.SingleAgentOffPolicyStrategy¶
Off-policy single-agent training (DQN, Rainbow DQN, DDPG, TD3).
- class agilerl.strategies.OfflineStrategy¶
Offline training from a fixed dataset (CQN).
- class agilerl.strategies.BanditStrategy¶
Contextual bandit training (NeuralTS, NeuralUCB).
- class agilerl.strategies.MultiAgentOnPolicyStrategy¶
On-policy multi-agent training (IPPO).
- class agilerl.strategies.MultiAgentOffPolicyStrategy¶
Off-policy multi-agent training (MADDPG, MATD3).
- class agilerl.strategies.LLMStrategy¶
Shared orchestration for the LLM fine-tuning loops.
- class agilerl.strategies.LLMRolloutStrategy¶
Generative rollout fine-tuning (GRPO family, LLM PPO, LLM REINFORCE).
One loop for every rollout regime: single-turn reasoning is
max_turns=1.
- class agilerl.strategies.LLMDatasetStrategy¶
Teacher-forced fine-tuning over dataset rows (DPO and SFT).
One loop for both objectives; the env’s
objectivepicks the loss.