Releases¶
v2.53.0: default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages¶
Released on 2026-10-10 - GitHub - PyPI
Breaking Changes
-
default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages
GRPO, GSPO and CISPO now default to
loss_norm: accumulation_window,kl_clamp: 10.0(also LLM PPO),off_policy_token_mask_bounds: [0.5, 5.0],off_policy_sequence_mask_threshold: 0.03,use_bias_correction_kl: true,adv_norm: mean_only, andtop_p: 1.0. Agents built or resumed without those keys pick up the new defaults. Opt out withkl_clamp: null,loss_norm: micro_batch,null/false/mean_std/0.95.kl_clampis a one-sided per-token bound on the K3 KL penalty: tokens further below the reference get no KL gradient. Off-policy token and sequence masks drop the policy and KL gradient of out-of-band tokens and drifted negative-advantage rows. Bias-corrected KL weights K3 by the un-clamped ratio to the old policy.The fused Liger GRPO-family loss now uses the same per-token weights as the PyTorch path under every
loss_norm. GSPO withuse_liger_loss=Trueruns (Liger sequence level) instead of raising. CISPO + Liger configs that trained atmicro_batchnow setloss_norm: accumulation_window.Arena
GRPOSpec/CISPOSpec/GSPOSpec/LLMPPOSpecexpose the new fields with those defaults.RolloutLLMSpec.top_pdefaults to 1.0.
Adaptive task sampling can pool rows into families:TaskAssigner(families=..., family_prior_strength=...)starts each row at its family's informative rate. Arena env specs exposetask_family_field(off by default) andtask_family_prior_strength(1.0).
ArenaTrainingSpec.reuse_prefix_cache_across_syncs(off by default) keeps the served LoRA name, and so vLLM's prefix cache, formax_rollout_version_lagweight versions; it requires an LLMreplay_bufferwithmax_rollout_version_lag >= 1.
Full Changelog: v2.52.0...v2.53.0
agilerl-arena/v1.24.0: agilerl-arena v1.24.0: default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages¶
Released on 2026-10-10 - GitHub - PyPI
Breaking Changes
-
default GRPO-family off-policy masks, kl_clamp 10, accumulation_window loss_norm and mean-only advantages
GRPO, GSPO and CISPO now default to
loss_norm: accumulation_window,kl_clamp: 10.0(also LLM PPO),off_policy_token_mask_bounds: [0.5, 5.0],off_policy_sequence_mask_threshold: 0.03,use_bias_correction_kl: true,adv_norm: mean_only, andtop_p: 1.0. Agents built or resumed without those keys pick up the new defaults. Opt out withkl_clamp: null,loss_norm: micro_batch,null/false/mean_std/0.95.kl_clampis a one-sided per-token bound on the K3 KL penalty: tokens further below the reference get no KL gradient. Off-policy token and sequence masks drop the policy and KL gradient of out-of-band tokens and drifted negative-advantage rows. Bias-corrected KL weights K3 by the un-clamped ratio to the old policy.The fused Liger GRPO-family loss now uses the same per-token weights as the PyTorch path under every
loss_norm. GSPO withuse_liger_loss=Trueruns (Liger sequence level) instead of raising. CISPO + Liger configs that trained atmicro_batchnow setloss_norm: accumulation_window.Arena
GRPOSpec/CISPOSpec/GSPOSpec/LLMPPOSpecexpose the new fields with those defaults.RolloutLLMSpec.top_pdefaults to 1.0.
Adaptive task sampling can pool rows into families:TaskAssigner(families=..., family_prior_strength=...)starts each row at its family's informative rate. Arena env specs exposetask_family_field(off by default) andtask_family_prior_strength(1.0).
ArenaTrainingSpec.reuse_prefix_cache_across_syncs(off by default) keeps the served LoRA name, and so vLLM's prefix cache, formax_rollout_version_lagweight versions; it requires an LLMreplay_bufferwithmax_rollout_version_lag >= 1.
Full Changelog: agilerl-arena/v1.23.0...agilerl-arena/v1.24.0
v2.52.0: speed up LLM learn with row balance, routed-expert chunks, and memory auto-pick¶
Released on 2026-10-10 - GitHub - PyPI
Features
-
speed up LLM learn with row balance, routed-expert chunks, and memory auto-pick
LLM algorithms share a faster learn loop: data-parallel ranks exchange segment rows so each rank runs about the mean count; packed-row layout, mixer scan, and micro-batch metrics stay on device until the end of learn; loss finiteness is checked on device and read once per optimizer step.
Routed-expert LoRA chunks by fixed row ranges with device-side offsets (
FSDPConfig.routed_expert_chunk_mib). Unset chunk size andoptim_cpu_offloadare resolved from the agilerl-arena estimate (largest fitting 64/128/256/512 MiB, then GPU fused AdamW only if the estimate still fits).optim_cpu_offloaddefaults to None. EP LoRA grads keep their parameter strides so fused AdamW accepts them. Frozen Mamba2out_projruns after the fused scan.torch._grouped_mmruns only on sm90/sm100 bf16.The arena estimator adds a routed-chunk term, GPU Adam state, packed expert LoRA as stacked tensors, a host breakdown, and shards weights across one
shard_group_sizegroup. Segmented rollouts keep prompts withinsegment_prompt_tokensand end withprompt_limitwhen a restart does not fit;max_row_tokenssizes the estimate. Optionalrestart_older_obs_field/restart_older_imagesshrink restart context (off by default). Frozen vision tower outputs are reused within a GRPO, PPO, or REINFORCE learn step.VLLMConfig.limit_mm_per_promptis a typed field.TaskAssignershards rows by stride. Checkpoints can snapshot on the host (snapshot_checkpoint) and store optimizer state (training.checkpoint_optimizer). Hugging Facetrust_remote_codeloads copy checkpoint code under a cross-process file lock; weight loads run outside that lock.Breaking:
balanced_row_plantakesranks_per_group;balance_rows_across_rankstakesshard_group_sizeandrow_values=, and returns(rows, values).pad_row_advantagesispad_row_values.materialize_fsdp2_from_cpu_stateneedsrouted_expert_chunk_mibset;FSDPRuntime.prepare_actorneedsoptim_cpu_offloadset (wrap_modelsresolves both).fsdp.compile_blocksandcompile_backendare removed; a manifest that sets them fails validation.DPO.learnreturns cross-rank means of loss and implicit rewards. Distributed LoRA init uses the shared seed on every rank.
Full Changelog: v2.51.0...v2.52.0
agilerl-arena/v1.23.0: agilerl-arena v1.23.0: speed up LLM learn with row balance, routed-expert chunks, and memory auto-pick¶
Released on 2026-10-10 - GitHub - PyPI
Features
-
speed up LLM learn with row balance, routed-expert chunks, and memory auto-pick
LLM algorithms share a faster learn loop: data-parallel ranks exchange segment rows so each rank runs about the mean count; packed-row layout, mixer scan, and micro-batch metrics stay on device until the end of learn; loss finiteness is checked on device and read once per optimizer step.
Routed-expert LoRA chunks by fixed row ranges with device-side offsets (
FSDPConfig.routed_expert_chunk_mib). Unset chunk size andoptim_cpu_offloadare resolved from the agilerl-arena estimate (largest fitting 64/128/256/512 MiB, then GPU fused AdamW only if the estimate still fits).optim_cpu_offloaddefaults to None. EP LoRA grads keep their parameter strides so fused AdamW accepts them. Frozen Mamba2out_projruns after the fused scan.torch._grouped_mmruns only on sm90/sm100 bf16.The arena estimator adds a routed-chunk term, GPU Adam state, packed expert LoRA as stacked tensors, a host breakdown, and shards weights across one
shard_group_sizegroup. Segmented rollouts keep prompts withinsegment_prompt_tokensand end withprompt_limitwhen a restart does not fit;max_row_tokenssizes the estimate. Optionalrestart_older_obs_field/restart_older_imagesshrink restart context (off by default). Frozen vision tower outputs are reused within a GRPO, PPO, or REINFORCE learn step.VLLMConfig.limit_mm_per_promptis a typed field.TaskAssignershards rows by stride. Checkpoints can snapshot on the host (snapshot_checkpoint) and store optimizer state (training.checkpoint_optimizer). Hugging Facetrust_remote_codeloads copy checkpoint code under a cross-process file lock; weight loads run outside that lock.Breaking:
balanced_row_plantakesranks_per_group;balance_rows_across_rankstakesshard_group_sizeandrow_values=, and returns(rows, values).pad_row_advantagesispad_row_values.materialize_fsdp2_from_cpu_stateneedsrouted_expert_chunk_mibset;FSDPRuntime.prepare_actorneedsoptim_cpu_offloadset (wrap_modelsresolves both).fsdp.compile_blocksandcompile_backendare removed; a manifest that sets them fails validation.DPO.learnreturns cross-rank means of loss and implicit rewards. Distributed LoRA init uses the shared seed on every rank.
Fixes
-
scale MoE expert LoRA on rank-r rows; load LLM checkpoints with extra LoRA target modules
low_rank_deltainagilerl.lora.moe.adaptersapplies the LoRA scaling to the rank-r intermediate, so no multiply runs over the full[rows, out]output in forward or backward. The result is unchanged.LLMAlgorithmcheckpoint loading accepts a checkpoint whose LoRA config adapts a strict superset of the live config'starget_modulesand otherwise matches. It warns and skips the extra modules' adapter weights. Other LoRA config mismatches still raise.
Full Changelog: agilerl-arena/v1.22.0...agilerl-arena/v1.23.0
v2.51.0: invert one memory setting and pick the cheapest fitting GPU tier¶
Released on 2026-10-09 - GitHub - PyPI
Features
-
invert one memory setting and pick the cheapest fitting GPU tier
arena memory solve FIELDholds every other input fixed and returns the largest value that still fits. Training uses the same underprediction buffer asestimate. Fields:max_model_len,max_num_seqs.--inferencesizes a dedicated serving GPU.arena memory estimatewithout--gpuor--device-gbpicks the cheapest Arena resource tier (credits per node-hour) whose node fits the manifest. Training fit uses the estimator's underprediction buffer; a tier that only fits on the point estimate is not recommended.
Fixes
-
scale MoE expert LoRA on rank-r rows; load LLM checkpoints with extra LoRA target modules
low_rank_deltainagilerl.lora.moe.adaptersapplies the LoRA scaling to the rank-r intermediate, so no multiply runs over the full[rows, out]output in forward or backward. The result is unchanged.LLMAlgorithmcheckpoint loading accepts a checkpoint whose LoRA config adapts a strict superset of the live config'starget_modulesand otherwise matches. It warns and skips the extra modules' adapter weights. Other LoRA config mismatches still raise.
Full Changelog: v2.50.0...v2.51.0
agilerl-arena/v1.22.0: agilerl-arena v1.22.0: invert one memory setting and pick the cheapest fitting GPU tier¶
Released on 2026-10-09 - GitHub - PyPI
Features
-
invert one memory setting and pick the cheapest fitting GPU tier
arena memory solve FIELDholds every other input fixed and returns the largest value that still fits. Training uses the same underprediction buffer asestimate. Fields:max_model_len,max_num_seqs.--inferencesizes a dedicated serving GPU.arena memory estimatewithout--gpuor--device-gbpicks the cheapest Arena resource tier (credits per node-hour) whose node fits the manifest. Training fit uses the estimator's underprediction buffer; a tier that only fits on the point estimate is not recommended.
Full Changelog: agilerl-arena/v1.21.0...agilerl-arena/v1.22.0
v2.50.0: add CNN→LSTM encoder for image observations (cnn_lstm)¶
Released on 2026-10-08 - GitHub - PyPI
Features
-
add CNN→LSTM encoder for image observations (cnn_lstm)
Add
EvolvableCnnLstm(CNN trunk + LSTM head) and select it when the observation is a 3D image Box andrecurrent=True. Hidden-state keys follow LSTM ({name}_h/{name}_c). Architecture mutation on the composite encoder is disabled.input_shapeis channel-first. Sequence forward keeps the time axis like LSTM. The module takesnet_config(CnnLstmNetConfigor a dict).Arena exposes
CnnLstmSpecandarch: cnn_lstm.infer_encoder_archreturnscnn_lstmfor image observations withrecurrent=True. Sample configs:ppo_cnn_lstm.yamlandppo_image_recurrent.yaml. Module defaults matchCnnLstmNetConfig.
Full Changelog: v2.49.1...v2.50.0
agilerl-arena/v1.21.0: agilerl-arena v1.21.0: add CNN→LSTM encoder for image observations (cnn_lstm)¶
Released on 2026-10-08 - GitHub - PyPI
Features
-
add CNN→LSTM encoder for image observations (cnn_lstm)
Add
EvolvableCnnLstm(CNN trunk + LSTM head) and select it when the observation is a 3D image Box andrecurrent=True. Hidden-state keys follow LSTM ({name}_h/{name}_c). Architecture mutation on the composite encoder is disabled.input_shapeis channel-first. Sequence forward keeps the time axis like LSTM. The module takesnet_config(CnnLstmNetConfigor a dict).Arena exposes
CnnLstmSpecandarch: cnn_lstm.infer_encoder_archreturnscnn_lstmfor image observations withrecurrent=True. Sample configs:ppo_cnn_lstm.yamlandppo_image_recurrent.yaml. Module defaults matchCnnLstmNetConfig.
Full Changelog: agilerl-arena/v1.20.0...agilerl-arena/v1.21.0
v2.49.1: pre-submission GPU memory estimate gate¶
Released on 2026-10-08 - GitHub - PyPI
Features
-
pre-submission GPU memory estimate gate
arena memory estimatesizes a training manifest against a GPU before submit. Exit 0 if both phases fit, 3 if either is over budget. Pass--configto stay offline.
Full Changelog: v2.49.0...v2.49.1
agilerl-arena/v1.20.0: agilerl-arena v1.20.0: pre-submission GPU memory estimate gate¶
Released on 2026-10-08 - GitHub - PyPI
Features
-
pre-submission GPU memory estimate gate
arena memory estimatesizes a training manifest against a GPU before submit. Exit 0 if both phases fit, 3 if either is over budget. Pass--configto stay offline.
Full Changelog: agilerl-arena/v1.19.0...agilerl-arena/v1.20.0
v2.49.0: Show action errors and instruction images in LLM env prompts, keep recent turns across context restarts, add PPO critic warmup and per-group gradient clipping¶
Released on 2026-10-07 - GitHub - PyPI
Other
-
Show action errors and instruction images in LLM env prompts, keep recent turns across context restarts, add PPO critic warmup and per-group gradient clipping
RolloutHarness(action_error_field=...)lists an observation field's error text after its action when a context restart summarizes past actions.RolloutHarnessputs the reset observation'sgoal_imagesafter the first prompt and every restarted prompt, letterboxed to the observation image's size so every image in an episode stacks into onepixel_valuestensor.RolloutHarness(restart_keep_turns=k)repeats the last k turns verbatim after a context restart.EnvResponsetakes gymnasiumterminatedandtruncated;doneis derived from them.TaskAssigner.record_outcometakes a group success outcome, andTaskRowStatsreports tied-failure / mixed / tied-success counts.- LLMPPO clips actor and critic gradients by their own norms (
share_grad_clip=Trueuses one coefficient from the combined norm).OptimizerStep.clip_coefsreplacesclip_coef. - LLMPPO
critic_warmup_stepstrains only the critic for the first N learn steps. - FSDP resumes value-head
lora_onlycheckpoints, and the critic LoRA stays trainable after a LoRA load. sync_grads/all_reduce_gradsskip params whose grad isNoneon every rank; ranks that disagree still raise.ModelArch.from_hf_configreads configs that nest the text model underllm_configand give depth only as a per-layer type list.GroupReplayStorekeeps per-task trajectories by return so tied rollout groups can be given a different-return member.CosineLRScheduleConfigsteps once per learn call:num_epochsis renamednum_steps, with newmin_lr_ratioandactor_start_step, and onelr_multiplier(step, start_step)method. The LLMPPO actor's schedule starts aftercritic_warmup_steps.LLMAlgorithm.current_lr_criticgives the critic rate the next learn trains with.
Full Changelog: v2.48.1...v2.49.0
agilerl-arena/v1.19.0: agilerl-arena v1.19.0: size the fused-pass check from the checkpoint’s config.json¶
Released on 2026-10-07 - GitHub - PyPI
Fixes
-
size the fused-pass check from the checkpoint's config.json
The auto fused-pass check now builds
ModelArchfrom the checkpoint's rawconfig.json(PretrainedConfig.get_config_dict) instead ofactor.config.to_dict(). transformers 5's NemotronHConfig dropsnum_hidden_layers, renames hybrid layer types, and can emit MoE defaults for dense checkpoints, which crashed withKeyError: 'num_hidden_layers'.CUDA fuse tests write that
config.json; a Nemotron Nano 4B config is parsed as 42 layers, 4 attention, 21 Mamba, dense.
Other
-
Show action errors and instruction images in LLM env prompts, keep recent turns across context restarts, add PPO critic warmup and per-group gradient clipping
RolloutHarness(action_error_field=...)lists an observation field's error text after its action when a context restart summarizes past actions.RolloutHarnessputs the reset observation'sgoal_imagesafter the first prompt and every restarted prompt, letterboxed to the observation image's size so every image in an episode stacks into onepixel_valuestensor.RolloutHarness(restart_keep_turns=k)repeats the last k turns verbatim after a context restart.EnvResponsetakes gymnasiumterminatedandtruncated;doneis derived from them.TaskAssigner.record_outcometakes a group success outcome, andTaskRowStatsreports tied-failure / mixed / tied-success counts.- LLMPPO clips actor and critic gradients by their own norms (
share_grad_clip=Trueuses one coefficient from the combined norm).OptimizerStep.clip_coefsreplacesclip_coef. - LLMPPO
critic_warmup_stepstrains only the critic for the first N learn steps. - FSDP resumes value-head
lora_onlycheckpoints, and the critic LoRA stays trainable after a LoRA load. sync_grads/all_reduce_gradsskip params whose grad isNoneon every rank; ranks that disagree still raise.ModelArch.from_hf_configreads configs that nest the text model underllm_configand give depth only as a per-layer type list.GroupReplayStorekeeps per-task trajectories by return so tied rollout groups can be given a different-return member.CosineLRScheduleConfigsteps once per learn call:num_epochsis renamednum_steps, with newmin_lr_ratioandactor_start_step, and onelr_multiplier(step, start_step)method. The LLMPPO actor's schedule starts aftercritic_warmup_steps.LLMAlgorithm.current_lr_criticgives the critic rate the next learn trains with.
Full Changelog: agilerl-arena/v1.18.0...agilerl-arena/v1.19.0
v2.48.1: size the fused-pass check from the checkpoint’s config.json¶
Released on 2026-10-07 - GitHub - PyPI
Fixes
-
size the fused-pass check from the checkpoint's config.json
The auto fused-pass check now builds
ModelArchfrom the checkpoint's rawconfig.json(PretrainedConfig.get_config_dict) instead ofactor.config.to_dict(). transformers 5's NemotronHConfig dropsnum_hidden_layers, renames hybrid layer types, and can emit MoE defaults for dense checkpoints, which crashed withKeyError: 'num_hidden_layers'.CUDA fuse tests write that
config.json; a Nemotron Nano 4B config is parsed as 42 layers, 4 attention, 21 Mamba, dense.
Full Changelog: v2.48.0...v2.48.1
v2.48.0: run LLMPPO actor and critic as separate or fused passes, picked by memory estimate¶
Released on 2026-10-07 - GitHub - PyPI
Features
-
run LLMPPO actor and critic as separate or fused passes, picked by memory estimate
LLMPPOruns the actor (policy and KL loss) and the critic (value loss) as separate forward/backward passes per micro-batch, so peak activation memory is one batch of rows, not two. This covers both the Liger and unfused loss paths.- New
fuse_actor_critic_pass: bool | None = NoneonLLMPPOandLLMPPOSpec.Trueruns one fused pass,Falseruns two.Nonefuses when theagilerl.arena.memoryestimate of the fused pass fits the GPU, and fuses on non-CUDA devices. The choice is re-resolved on checkpoint load. - On the Liger path the critic pass runs under the critic adapter, so the value loss trains the critic LoRA.
- The critic adapter stays trainable after FSDP2 replaces the model's parameters.
- With
activation_offload=True, the split Liger actor pass also offloads saved activations to CPU. - The fused-pass memory estimate uses the packed-expert LoRA path the model actually runs.
- PPO
learnreports actor and critic gradient norms before and after clipping. agilerl-arenamemory estimator:TrainingSettingsgainsfuse_actor_critic_pass(PPO only) andmicro_batch_size.agilerlnow requiresagilerl-arena>=1.18.0.
Full Changelog: v2.47.0...v2.48.0
agilerl-arena/v1.18.0: agilerl-arena v1.18.0: run LLMPPO actor and critic as separate or fused passes, picked by memory estimate¶
Released on 2026-10-07 - GitHub - PyPI
Features
-
run LLMPPO actor and critic as separate or fused passes, picked by memory estimate
LLMPPOruns the actor (policy and KL loss) and the critic (value loss) as separate forward/backward passes per micro-batch, so peak activation memory is one batch of rows, not two. This covers both the Liger and unfused loss paths.- New
fuse_actor_critic_pass: bool | None = NoneonLLMPPOandLLMPPOSpec.Trueruns one fused pass,Falseruns two.Nonefuses when theagilerl.arena.memoryestimate of the fused pass fits the GPU, and fuses on non-CUDA devices. The choice is re-resolved on checkpoint load. - On the Liger path the critic pass runs under the critic adapter, so the value loss trains the critic LoRA.
- The critic adapter stays trainable after FSDP2 replaces the model's parameters.
- With
activation_offload=True, the split Liger actor pass also offloads saved activations to CPU. - The fused-pass memory estimate uses the packed-expert LoRA path the model actually runs.
- PPO
learnreports actor and critic gradient norms before and after clipping. agilerl-arenamemory estimator:TrainingSettingsgainsfuse_actor_critic_pass(PPO only) andmicro_batch_size.agilerlnow requiresagilerl-arena>=1.18.0.
Full Changelog: agilerl-arena/v1.17.0...agilerl-arena/v1.18.0
v2.47.0: Nemotron 3.5 Super-VL training with expert and tensor parallel LoRA¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
Nemotron 3.5 Super-VL training with expert and tensor parallel LoRA
Adds the Nemotron-H / Super-VL architecture, expert-parallel, tensor-parallel and node-local sharded FSDP training, and MoE expert LoRA with grouped GEMM and recompute. Also adds sequence packing that resets Mamba state at document boundaries, old log-probs from rollouts, an episode-balanced
loss_normfor GRPO-family learners, and learn profiling. Image episodes record the processor inputs they keep, so a trainer can rebuildpixel_valuesfrom the screenshots. Moves LoRA intoagilerl.lora.
Full Changelog: v2.46.1...v2.47.0
agilerl-arena/v1.17.0: agilerl-arena v1.17.0: Nemotron 3.5 Super-VL training with expert and tensor parallel LoRA¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
Nemotron 3.5 Super-VL training with expert and tensor parallel LoRA
Adds the Nemotron-H / Super-VL architecture, expert-parallel, tensor-parallel and node-local sharded FSDP training, and MoE expert LoRA with grouped GEMM and recompute. Also adds sequence packing that resets Mamba state at document boundaries, old log-probs from rollouts, an episode-balanced
loss_normfor GRPO-family learners, and learn profiling. Image episodes record the processor inputs they keep, so a trainer can rebuildpixel_valuesfrom the screenshots. Moves LoRA intoagilerl.lora.
Full Changelog: agilerl-arena/v1.16.0...agilerl-arena/v1.17.0
v2.46.1: closed-form GPU memory estimator¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
closed-form GPU memory estimator
Peak GPU memory for LLM RL from model geometry, device, and training/generation settings. Two independent phase bars (training and generation never peak at once). No profiling, no fitted correction, no weight download.
PhaseTimer CPU tests drive a fake clock so they do not depend on Windows sleep resolution.
Full Changelog: v2.46.0...v2.46.1
agilerl-arena/v1.16.0: agilerl-arena v1.16.0: expose MF-PBT evolution planning and hyperparameter reset¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
expose MF-PBT evolution planning and hyperparameter reset
MultiFrequencySelection.plan_evolutionreturns one generation as an index plan without cloning agents.apply_hp_resetis public so a migrant can take a destination elite's mutable hyperparameters. -
closed-form GPU memory estimator
Peak GPU memory for LLM RL from model geometry, device, and training/generation settings. Two independent phase bars (training and generation never peak at once). No profiling, no fitted correction, no weight download.
PhaseTimer CPU tests drive a fake clock so they do not depend on Windows sleep resolution.
Full Changelog: agilerl-arena/v1.15.0...agilerl-arena/v1.16.0
v2.46.0: add release status to supported models¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
add release status to supported models
ModelInfo.statusmarks each supported modellive,deprecated, or
preview, andarena models supportedincludes it.ArenaClient.submit_experiment
warns when the manifest uses a deprecated model and rejects a preview one.
Manifest validation still does not consult status.The Super-VL bundled inspected LoRA data includes language-tower dims
(vision and connector targets stay names-only), so that id has derived LoRA
ranks and drops latent projections from the allowlist. -
expose MF-PBT evolution planning and hyperparameter reset
MultiFrequencySelection.plan_evolutionreturns one generation as an index plan without cloning agents.apply_hp_resetis public so a migrant can take a destination elite's mutable hyperparameters.
Full Changelog: v2.45.0...v2.46.0
agilerl-arena/v1.15.0: agilerl-arena v1.15.0: load every parquet split under a dataset directory¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
load every parquet split under a dataset directory
A local parquet directory whose files live in immediate child directories now loads every split. When a train split exists alongside another split, training uses every train row and evaluation uses every other split. A file, a flat shard directory, a train-only directory, or split directories with no train directory still use a random holdout. Empty directories still raise.
-
add release status to supported models
ModelInfo.statusmarks each supported modellive,deprecated, or
preview, andarena models supportedincludes it.ArenaClient.submit_experiment
warns when the manifest uses a deprecated model and rejects a preview one.
Manifest validation still does not consult status.The Super-VL bundled inspected LoRA data includes language-tower dims
(vision and connector targets stay names-only), so that id has derived LoRA
ranks and drops latent projections from the allowlist.
Full Changelog: agilerl-arena/v1.14.0...agilerl-arena/v1.15.0
v2.45.0: load every parquet split under a dataset directory¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
load every parquet split under a dataset directory
A local parquet directory whose files live in immediate child directories now loads every split. When a train split exists alongside another split, training uses every train row and evaluation uses every other split. A file, a flat shard directory, a train-only directory, or split directories with no train directory still use a random holdout. Empty directories still raise.
Full Changelog: v2.44.0...v2.45.0
v2.44.0: let LLMEnvSpec declare env sessions and shared services¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
let LLMEnvSpec declare env sessions and shared services
LLMEnvSpecnow acceptsenv_sessions_per_host,env_session_ports,env_host_memory_limit_bytes,env_host_ready_timeout_s, andenv_servicesso oneenv_imagecan run several sessions plus shared backing services. More than one session per host requiresenv_session_ports. Withenv_image,env_sessions_per_hostdefaults to 1 andenv_host_ready_timeout_sto 600.EnvServiceSpecandEnvContainerSpecare public models (agilerl.arena.models) and reject unknown keys. Those host settings still requireenv_image, and dataset environments reject them. Theenv_imagedefault forcpus_per_env_hostis0.01. -
VisualWebArena env client and multi-turn vision segment learning
Adds image observations and a vision transcript to the OpenEnv harness, so vision-language models can train on VisualWebArena. Long episodes restart as new segments when the prompt outgrows a token or image budget. GRPO, PPO and REINFORCE for LLMs learn from those episode segments and pass pixel values through the fused log-prob forward and loss. Every LLM algorithm reports learn-phase timings. Adds held-out evaluation fields and adaptive task sampling to the arena env and training models.
Full Changelog: v2.43.0...v2.44.0
agilerl-arena/v1.14.0: agilerl-arena v1.14.0: VisualWebArena env client and multi-turn vision segment learning¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
VisualWebArena env client and multi-turn vision segment learning
Adds image observations and a vision transcript to the OpenEnv harness, so vision-language models can train on VisualWebArena. Long episodes restart as new segments when the prompt outgrows a token or image budget. GRPO, PPO and REINFORCE for LLMs learn from those episode segments and pass pixel values through the fused log-prob forward and loss. Every LLM algorithm reports learn-phase timings. Adds held-out evaluation fields and adaptive task sampling to the arena env and training models.
Full Changelog: agilerl-arena/v1.13.0...agilerl-arena/v1.14.0
agilerl-arena/v1.13.0: agilerl-arena v1.13.0: let LLMEnvSpec declare env sessions and shared services¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
let LLMEnvSpec declare env sessions and shared services
LLMEnvSpecnow acceptsenv_sessions_per_host,env_session_ports,env_host_memory_limit_bytes,env_host_ready_timeout_s, andenv_servicesso oneenv_imagecan run several sessions plus shared backing services. More than one session per host requiresenv_session_ports. Withenv_image,env_sessions_per_hostdefaults to 1 andenv_host_ready_timeout_sto 600.EnvServiceSpecandEnvContainerSpecare public models (agilerl.arena.models) and reject unknown keys. Those host settings still requireenv_image, and dataset environments reject them. Theenv_imagedefault forcpus_per_env_hostis0.01.
Full Changelog: agilerl-arena/v1.12.0...agilerl-arena/v1.13.0
v2.43.0: support Qwen/Qwen3.8-27B in agilerl-arena¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
support Qwen/Qwen3.8-27B in agilerl-arena
Adds
Qwen/Qwen3.8-27BtoSUPPORTED_MODEL_INFOwith its LoRA targets, allowed ranks, and bundledconfig.json.
Full Changelog: v2.42.5...v2.43.0
agilerl-arena/v1.12.0: agilerl-arena v1.12.0: support Qwen/Qwen3.8-27B in agilerl-arena¶
Released on 2026-10-06 - GitHub - PyPI
Features
-
support Qwen/Qwen3.8-27B in agilerl-arena
Adds
Qwen/Qwen3.8-27BtoSUPPORTED_MODEL_INFOwith its LoRA targets, allowed ranks, and bundledconfig.json.
Fixes
-
map FSDP checkpoint keys forward through one-to-one renames
FSDP shard load applies Hugging Face one-to-one weight renames forward over safetensors keys, then looks live parameter names up in that map, matching
from_pretrained. Checkpoints that only need part of a family's mapping now load; Nemotron-H storesbackbone.embeddings.weight, and undoing bothbackbone.→model.andembedding.weight→embeddings.weightsearched forbackbone.embedding.weight. Packed expert shards stacked into one parameter are found under the renamed keys. Split packing converters still resolve through the reverse path. Two checkpoint keys that rename to the same parameter raise an error.
Full Changelog: agilerl-arena/v1.11.1...agilerl-arena/v1.12.0
v2.42.5: point arena CLI login at arena-auth.agilerl.com¶
Released on 2026-10-05 - GitHub - PyPI
Fixes
-
point arena CLI login at arena-auth.agilerl.com
The default Keycloak URL for
arena loginwashttps://auth.arena.agilerl.com, which returns 404 and does not serve the Arena realm. It is nowhttps://arena-auth.agilerl.com.--keycloak-urlandARENA_KEYCLOAK_URLstill override the default. -
map FSDP checkpoint keys forward through one-to-one renames
FSDP shard load applies Hugging Face one-to-one weight renames forward over safetensors keys, then looks live parameter names up in that map, matching
from_pretrained. Checkpoints that only need part of a family's mapping now load; Nemotron-H storesbackbone.embeddings.weight, and undoing bothbackbone.→model.andembedding.weight→embeddings.weightsearched forbackbone.embedding.weight. Packed expert shards stacked into one parameter are found under the renamed keys. Split packing converters still resolve through the reverse path. Two checkpoint keys that rename to the same parameter raise an error.
Full Changelog: v2.42.4...v2.42.5
agilerl-arena/v1.11.1: agilerl-arena v1.11.1: point arena CLI login at arena-auth.agilerl.com¶
Released on 2026-10-05 - GitHub - PyPI
Fixes
-
point arena CLI login at arena-auth.agilerl.com
The default Keycloak URL for
arena loginwashttps://auth.arena.agilerl.com, which returns 404 and does not serve the Arena realm. It is nowhttps://arena-auth.agilerl.com.--keycloak-urlandARENA_KEYCLOAK_URLstill override the default.
Full Changelog: agilerl-arena/v1.11.0...agilerl-arena/v1.11.1
agilerl-arena/v1.11.0: agilerl-arena v1.11.0: fold vLLM into agilerl[llm] and lead README with LLM post-training¶
Released on 2026-10-05 - GitHub - PyPI
Features
-
fold vLLM into agilerl[llm] and lead README with LLM post-training
agilerl[llm]now includes vLLM on Linux. New extraagilerl[cpu-llm]is the Hugging Face, PEFT and datasets stack without vLLM.HAS_LLM_DEPENDENCIESfollows[cpu-llm]; vLLM isHAS_VLLM.Rewrite the README to lead with LLM post-training: SFT, DPO and RL (GRPO, CISPO, GSPO, REINFORCE, PPO), multi-turn OpenEnv environments, LoRA, FSDP2, memory and model-specific optimizations, and vLLM rollouts. Add sections on training at scale with the Arena CLI (including use from coding agents) and on local LLM training. Shorten the classic RL section, document
cpu-llmin the install extras table, and add SFT to the algorithm tables. -
ship supported model info and configs in agilerl-arena
agilerl.arena.models.SUPPORTED_MODEL_INFOlists supported Hugging Face ids
with their architecture, LoRAtarget_modulesnames, packed-expert
target_parameterspaths, per-target LoRA dims (lora_info), parameter
counts, and allowed LoRA ranks. Each id bundles one JSON file with itsconfig.jsonand inspected info;
everything is read from those.arena models supported
prints the list as JSON without a server call.Manifests for listed ids now reject LoRA ranks above the model's cap and any
target_modulesortarget_parametersnot in the stored lists, with the valid
options in the error.target_modulesmust beall-linearor exact module
names; regex strings are rejected. On MoE models an unsettarget_parameters
adapts every expert and zeroslora_dropout([]adapts none). Unlisted ids skip these checks. Gemma 4
example configs use bare projection names. Supported ids resolve their family from the bundled config with no Hub call.agilerlnow requiresagilerl-arena1.11.0 or newer.
Fixes
-
re-tie output head and recompute RoPE after meta load
to_emptyon the FSDP meta-device load path splits a tiedlm_head/embed_outfromembed_tokens. Tied checkpoints omit the head tensor, so the head stays empty and LoRA gradients are zero. Calltie_weights()on the module whose registered children include the head (walking through PEFT and value-head wrappers) and raise if a tied key is still a separate parameter. Non-persistent RoPEinv_freqbuffers are not stored in safetensors; recompute them the same way Hugging Face_init_weightsdoes.Pin
COVERAGE_FILEto an absolute path so pytest's session chdir cannot drop coverage hits, and combine parallel coverage DBs from Linux 3.13 shards. -
Gemma 4 vLLM startup, FSDP scalar buffers, L4 flex tiles
Gemma 4
text_confighas noarchitectureslist. Stripping multimodal towers now maps that nested config toGemma4ForCausalLMso vLLM can load the language engine. Tower-connector LoRA stays off for Gemma 4; only families that implement encoder token counts enable it. FSDP shard load usesget_tensorfor full buffer copies so 0-dim buffers (clipped-linear clamp bounds) load. Flex attention on L4-class GPUs uses smaller backward tiles to fit their ~99 KB shared memory. -
load Nemotron-H trainers with the transformers class
Nemotron-H trainer runtime no longer sets
trust_remote_code. Hugging Face transformers then uses its ownnemotron_hclass instead of the checkpointauto_mapcode, which only allows eager attention.nemotron_h_omnistill setstrust_remote_codebecause transformers does not ship that model type. vLLM still setstrust_remote_codefor both families.
Full Changelog: agilerl-arena/v1.10.0...agilerl-arena/v1.11.0
v2.42.3: load Nemotron-H trainers with the transformers class¶
Released on 2026-10-05 - GitHub - PyPI
Fixes
-
load Nemotron-H trainers with the transformers class
Nemotron-H trainer runtime no longer sets
trust_remote_code. Hugging Face transformers then uses its ownnemotron_hclass instead of the checkpointauto_mapcode, which only allows eager attention.nemotron_h_omnistill setstrust_remote_codebecause transformers does not ship that model type. vLLM still setstrust_remote_codefor both families.
Full Changelog: v2.42.2...v2.42.3
v2.42.2: Gemma 4 vLLM startup, FSDP scalar buffers, L4 flex tiles¶
Released on 2026-10-05 - GitHub - PyPI
Fixes
-
Gemma 4 vLLM startup, FSDP scalar buffers, L4 flex tiles
Gemma 4
text_confighas noarchitectureslist. Stripping multimodal towers now maps that nested config toGemma4ForCausalLMso vLLM can load the language engine. Tower-connector LoRA stays off for Gemma 4; only families that implement encoder token counts enable it. FSDP shard load usesget_tensorfor full buffer copies so 0-dim buffers (clipped-linear clamp bounds) load. Flex attention on L4-class GPUs uses smaller backward tiles to fit their ~99 KB shared memory.
Full Changelog: v2.42.1...v2.42.2
v2.42.1: re-tie output head and recompute RoPE after meta load¶
Released on 2026-10-02 - GitHub - PyPI
Fixes
-
re-tie output head and recompute RoPE after meta load
to_emptyon the FSDP meta-device load path splits a tiedlm_head/embed_outfromembed_tokens. Tied checkpoints omit the head tensor, so the head stays empty and LoRA gradients are zero. Calltie_weights()on the module whose registered children include the head (walking through PEFT and value-head wrappers) and raise if a tied key is still a separate parameter. Non-persistent RoPEinv_freqbuffers are not stored in safetensors; recompute them the same way Hugging Face_init_weightsdoes.Pin
COVERAGE_FILEto an absolute path so pytest's session chdir cannot drop coverage hits, and combine parallel coverage DBs from Linux 3.13 shards.
Full Changelog: v2.42.0...v2.42.1
v2.42.0: fold vLLM into agilerl[llm] and lead README with LLM post-training¶
Released on 2026-10-02 - GitHub - PyPI
Features
-
fold vLLM into agilerl[llm] and lead README with LLM post-training
agilerl[llm]now includes vLLM on Linux. New extraagilerl[cpu-llm]is the Hugging Face, PEFT and datasets stack without vLLM.HAS_LLM_DEPENDENCIESfollows[cpu-llm]; vLLM isHAS_VLLM.Rewrite the README to lead with LLM post-training: SFT, DPO and RL (GRPO, CISPO, GSPO, REINFORCE, PPO), multi-turn OpenEnv environments, LoRA, FSDP2, memory and model-specific optimizations, and vLLM rollouts. Add sections on training at scale with the Arena CLI (including use from coding agents) and on local LLM training. Shorten the classic RL section, document
cpu-llmin the install extras table, and add SFT to the algorithm tables.
Fixes
-
emit wizard schema defaults and drop unused form fields
Served
/manifest/schemanow includes numericevo_stepsdefaults (LLM 20, classic 160000, CNN/MultiInput 320000), inlineslora_configRank/Alpha/Dropout (1/32/0.05), omitslearning_delayfrom on-policy training, strips unusedanswer_pattern, and gates Ornstein-Uhlenbecktheta/dtwith if/then.
Full Changelog: v2.41.1...v2.42.0
agilerl-arena/v1.10.0: agilerl-arena v1.10.0: onboard gpt-oss-20b with flex attention and packed-expert LoRA¶
Released on 2026-10-01 - GitHub - PyPI
Features
-
onboard gpt-oss-20b with flex attention and packed-expert LoRA
Family trainer attention defaults (Gemma SWA / gpt-oss Flex Attention) beat YAML and spec values; ATTN_IMPLEMENTATION still overrides. Packed-expert LoRA (target_parameters) forces lora_dropout=0 because PEFT cannot factor dropout out of the parameter-level product. GptOssExperts uses a split-LoRA forward for its matmul layout, biases, and interleaved SwiGLU. The grouped-GEMM capability probe no longer adds saved tensors to an activation checkpoint and no longer reports unsupported when first called under no_grad.
-
load parquet files or shard directories into one Dataset
Add
load_parquet_datasetfor a local.parquet/.pqfile or a directory of*.parquetshards. Shards concatenate as Hugging Face datasets in filename order; mismatched schemas raise.make_llm_envuses this for parquet files and directories that contain*.parquet; other paths still load as a Hub id or local Hugging Face dataset. -
log sampled rollout completions during training
train_llm_rollouttakes a newcompletion_logging=CompletionLoggingConfig(...)argument. When set, each iteration samples prompt groups from the rollout batch and decodes prompt, completion, reward, turn count and completion-token count. Everyintervaliterations the samples from every population member go in one write to the console (verbose), a W&Bcompletionstable (wb) and an optional JSONL file (jsonl_path). Text fields are capped atmax_chars, keeping head and tail. The latest samples are also logged at error level if training raises.The building blocks (
CompletionLogger,CompletionRecord,build_completion_recordand the writers) live inagilerl.training.llm.completion_loggingfor custom training loops. -
sample trajectories per group and mark truncated logs
CompletionLoggingConfiggainssamples_per_group: when set, each sampled prompt group logs that many random trajectories instead of the whole group (unset keeps the previous whole-group behaviour). Over-long prompt/completion text is still head/tail truncated, now with a<LOG TRUNCATED: N chars>seam marker.
Fixes
-
release device memory after evolutionary mutation
Add
release_device_memory()and call it fromMutations.mutationand from
clean_up()so MPS/CUDA caches from cloning and architecture changes do not
accumulate across generations. Document architecture-mutation RSS on CPU and
custom encoderget_init_dictrequirements in the mutation guide. -
emit wizard schema defaults and drop unused form fields
Served
/manifest/schemanow includes numericevo_stepsdefaults (LLM 20, classic 160000, CNN/MultiInput 320000), inlineslora_configRank/Alpha/Dropout (1/32/0.05), omitslearning_delayfrom on-policy training, strips unusedanswer_pattern, and gates Ornstein-Uhlenbecktheta/dtwith if/then.
Other
-
save duration maps on the Linux test runners
CI durations writes the Linux CPU timing map on ubuntu-24.04 and the GPU map on the self-hosted scale-set, matching the jobs that restore them.
-
run a single macOS job on Python 3.13
GitHub-hosted macOS CI runs one job (macos-26, Python 3.13) instead of four Python versions.
-
bump GitHub Actions off Node 20 and CodeQL Action v3
Put the CodeQL job steps in
.github/workflows/codeql.ymlso that file has
on.push(CI still calls it as a merge gate; push to main and the weekly
schedule still feed the Security tab). Useactions/checkout@v5and
github/codeql-actionv4. -
shard macOS and Windows pytest with duration maps
macOS runs three duration-balanced shards and Windows two. Timing maps are stored on the default branch and restored on the same runner family that wrote them. Before a map exists, shards are packed with equal-weight LPT rather than contiguous slices.
-
refresh Python lockfile
Regenerates the Python lockfile after dependency updates.
Applies Liger Kernel patches to the module returned by PEFT's
get_base_model().
Full Changelog: agilerl-arena/v1.9.4...agilerl-arena/v1.10.0
v2.41.1: refresh Python lockfile¶
Released on 2026-09-30 - GitHub - PyPI
Other
-
refresh Python lockfile
Regenerates the Python lockfile after dependency updates.
Applies Liger Kernel patches to the module returned by PEFT's
get_base_model().
Full Changelog: v2.41.0...v2.41.1
v2.41.0: sample trajectories per group and mark truncated logs¶
Released on 2026-09-30 - GitHub - PyPI
Features
-
sample trajectories per group and mark truncated logs
CompletionLoggingConfiggainssamples_per_group: when set, each sampled prompt group logs that many random trajectories instead of the whole group (unset keeps the previous whole-group behaviour). Over-long prompt/completion text is still head/tail truncated, now with a<LOG TRUNCATED: N chars>seam marker.
Full Changelog: v2.40.1...v2.41.0
v2.40.1: shard macOS and Windows pytest with duration maps¶
Released on 2026-09-30 - GitHub - PyPI
Other
-
shard macOS and Windows pytest with duration maps
macOS runs three duration-balanced shards and Windows two. Timing maps are stored on the default branch and restored on the same runner family that wrote them. Before a map exists, shards are packed with equal-weight LPT rather than contiguous slices.
Full Changelog: v2.40.0...v2.40.1
v2.40.0: log sampled rollout completions during training¶
Released on 2026-09-30 - GitHub - PyPI
Features
-
log sampled rollout completions during training
train_llm_rollouttakes a newcompletion_logging=CompletionLoggingConfig(...)argument. When set, each iteration samples prompt groups from the rollout batch and decodes prompt, completion, reward, turn count and completion-token count. Everyintervaliterations the samples from every population member go in one write to the console (verbose), a W&Bcompletionstable (wb) and an optional JSONL file (jsonl_path). Text fields are capped atmax_chars, keeping head and tail. The latest samples are also logged at error level if training raises.The building blocks (
CompletionLogger,CompletionRecord,build_completion_recordand the writers) live inagilerl.training.llm.completion_loggingfor custom training loops.
Other
-
bump GitHub Actions off Node 20 and CodeQL Action v3
Put the CodeQL job steps in
.github/workflows/codeql.ymlso that file has
on.push(CI still calls it as a merge gate; push to main and the weekly
schedule still feed the Security tab). Useactions/checkout@v5and
github/codeql-actionv4.
Full Changelog: v2.39.0...v2.40.0
v2.39.0: load parquet files or shard directories into one Dataset¶
Released on 2026-09-30 - GitHub - PyPI
Features
-
load parquet files or shard directories into one Dataset
Add
load_parquet_datasetfor a local.parquet/.pqfile or a directory of*.parquetshards. Shards concatenate as Hugging Face datasets in filename order; mismatched schemas raise.make_llm_envuses this for parquet files and directories that contain*.parquet; other paths still load as a Hub id or local Hugging Face dataset.
Full Changelog: v2.38.1...v2.39.0
v2.38.1: release device memory after evolutionary mutation¶
Released on 2026-09-30 - GitHub - PyPI
Fixes
-
release device memory after evolutionary mutation
Add
release_device_memory()and call it fromMutations.mutationand from
clean_up()so MPS/CUDA caches from cloning and architecture changes do not
accumulate across generations. Document architecture-mutation RSS on CPU and
custom encoderget_init_dictrequirements in the mutation guide.
Full Changelog: v2.38.0...v2.38.1
v2.38.0: onboard gpt-oss-20b with flex attention and packed-expert LoRA¶
Released on 2026-09-29 - GitHub - PyPI
Features
-
onboard gpt-oss-20b with flex attention and packed-expert LoRA
Family trainer attention defaults (Gemma SWA / gpt-oss Flex Attention) beat YAML and spec values; ATTN_IMPLEMENTATION still overrides. Packed-expert LoRA (target_parameters) forces lora_dropout=0 because PEFT cannot factor dropout out of the parameter-level product. GptOssExperts uses a split-LoRA forward for its matmul layout, biases, and interleaved SwiGLU. The grouped-GEMM capability probe no longer adds saved tensors to an activation checkpoint and no longer reports unsupported when first called under no_grad.
Other
-
save duration maps on the Linux test runners
CI durations writes the Linux CPU timing map on ubuntu-24.04 and the GPU map on the self-hosted scale-set, matching the jobs that restore them.
-
run a single macOS job on Python 3.13
GitHub-hosted macOS CI runs one job (macos-26, Python 3.13) instead of four Python versions.
Full Changelog: v2.37.0...v2.38.0
v2.37.0: publish training_gpus_per_agent only on LLM training schemas¶
Released on 2026-09-29 - GitHub - PyPI
Fixes
-
publish training_gpus_per_agent only on LLM training schemas
The published agilerl-arena training-manifest schema now defines
training_gpus_per_agentonly on the LLM training spec. Classic RL and multi-agent gym algorithms no longer expose that key.to_payload()omits it for non-LLM runs, and for LLM runs that leave the default unset.
Other
-
merge per-root junit as well-formed XML
Two-root pytest shards no longer concatenate
<testsuites>wrappers with string slicing. The merged junit report is built with ElementTree so duration promote can parse it.Cover stripping non-form algorithm fields from JSON Schema required lists.
-
run Linux CPU, ty, and coverage on GitHub-hosted Ubuntu
Linux CPU pytest, ty, and coverage combine run on ubuntu-24.04 without CUDA torch or vLLM. GPU pytest stays on self-hosted gha-runner-scale-set runners. The llm extra is Hugging Face transformers/PEFT/datasets plus Liger and bitsandbytes on Linux. vLLM lives in the vllm extra. The cpu extra selects CPU-only PyTorch.
Full Changelog: v2.36.2...v2.37.0
agilerl-arena/v1.9.4: agilerl-arena v1.9.4: run Linux CPU, ty, and coverage on GitHub-hosted Ubuntu¶
Released on 2026-09-29 - GitHub - PyPI
Other
-
run Linux CPU, ty, and coverage on GitHub-hosted Ubuntu
Linux CPU pytest, ty, and coverage combine run on ubuntu-24.04 without CUDA torch or vLLM. GPU pytest stays on self-hosted gha-runner-scale-set runners. The llm extra is Hugging Face transformers/PEFT/datasets plus Liger and bitsandbytes on Linux. vLLM lives in the vllm extra. The cpu extra selects CPU-only PyTorch.
Full Changelog: agilerl-arena/v1.9.3...agilerl-arena/v1.9.4
agilerl-arena/v1.9.3: agilerl-arena v1.9.3: publish training_gpus_per_agent only on LLM training schemas¶
Released on 2026-09-29 - GitHub - PyPI
Fixes
-
publish training_gpus_per_agent only on LLM training schemas
The published agilerl-arena training-manifest schema now defines
training_gpus_per_agentonly on the LLM training spec. Classic RL and multi-agent gym algorithms no longer expose that key.to_payload()omits it for non-LLM runs, and for LLM runs that leave the default unset.
Full Changelog: agilerl-arena/v1.9.2...agilerl-arena/v1.9.3
agilerl-arena/v1.9.2: agilerl-arena v1.9.2: merge per-root junit as well-formed XML¶
Released on 2026-09-29 - GitHub - PyPI
Other
-
merge per-root junit as well-formed XML
Two-root pytest shards no longer concatenate
<testsuites>wrappers with string slicing. The merged junit report is built with ElementTree so duration promote can parse it.Cover stripping non-form algorithm fields from JSON Schema required lists.
Full Changelog: agilerl-arena/v1.9.1...agilerl-arena/v1.9.2
v2.36.2: publish dataset, epsilon, and SFT beta defaults only for runs that use them¶
Released on 2026-09-28 - GitHub - PyPI
Fixes
-
publish dataset, epsilon, and SFT beta defaults only for runs that use them
LLM environment fields
train_test_split,response_column,rubric_name,strict_chat_template_boundary, andnum_envsare optional onLLMEnvSpec. Dataset-backed rollouts still get the train/test split andrubric_name. SFT still gets the split andresponse_column. GEM-style rollouts still getnum_envsandstrict_chat_template_boundary.make_llm_envrequiresresponse_column; dataset loaders requiretrain_test_split; an unset rollout chat-template boundary is treated as strict. Rainbow DQN no longer publishes epsilon-greedy training defaults. The SFT algorithm schema no longer includesbeta; DPO still defaults it to 0.1.
Other
-
fail closed when duration promote has no junit
CI durationsnow fails if a successfulCIrun has no 3.13 junit XML or an empty duration map, instead of skipping the cache save and still going green.merge-junitfinds XML under the download directory itself.
Full Changelog: v2.36.1...v2.36.2
agilerl-arena/v1.9.1: agilerl-arena v1.9.1: publish dataset, epsilon, and SFT beta defaults only for runs that use them¶
Released on 2026-09-28 - GitHub - PyPI
Fixes
-
publish dataset, epsilon, and SFT beta defaults only for runs that use them
LLM environment fields
train_test_split,response_column,rubric_name,strict_chat_template_boundary, andnum_envsare optional onLLMEnvSpec. Dataset-backed rollouts still get the train/test split andrubric_name. SFT still gets the split andresponse_column. GEM-style rollouts still getnum_envsandstrict_chat_template_boundary.make_llm_envrequiresresponse_column; dataset loaders requiretrain_test_split; an unset rollout chat-template boundary is treated as strict. Rainbow DQN no longer publishes epsilon-greedy training defaults. The SFT algorithm schema no longer includesbeta; DPO still defaults it to 0.1.
Other
-
split Linux pytest into duration-balanced CPU and GPU shards
Linux GitHub Actions runs CPU tests (
not gpu and not vllm, CUDA hidden) and GPU/vLLM tests as two duration-balanced shards each per Python 3.10–3.13 viascripts/pytest_shard.py. A greenCIrun promotes 3.13 junit timings onto default-branch Actions caches so the next PR or dispatch can pack from measured times. Each shard appends coverage acrosstests/andagilerl-arena/testsand merges their junit reports. A root with no matching tests is an empty shard, not a failure. Combine requires both 3.13 CPU shards and both GPU shards before writingframework-coverage.json. Coverage reports store paths relative to the repo root. macOS and Windows run the full suite and upload junit. Latent-encoder function-preserving checks allow float32 GEMM noise (rtol=1e-5). -
fail closed when duration promote has no junit
CI durationsnow fails if a successfulCIrun has no 3.13 junit XML or an empty duration map, instead of skipping the cache save and still going green.merge-junitfinds XML under the download directory itself.
Full Changelog: agilerl-arena/v1.9.0...agilerl-arena/v1.9.1
v2.36.1: split Linux pytest into duration-balanced CPU and GPU shards¶
Released on 2026-09-28 - GitHub - PyPI
Other
-
split Linux pytest into duration-balanced CPU and GPU shards
Linux GitHub Actions runs CPU tests (
not gpu and not vllm, CUDA hidden) and GPU/vLLM tests as two duration-balanced shards each per Python 3.10–3.13 viascripts/pytest_shard.py. A greenCIrun promotes 3.13 junit timings onto default-branch Actions caches so the next PR or dispatch can pack from measured times. Each shard appends coverage acrosstests/andagilerl-arena/testsand merges their junit reports. A root with no matching tests is an empty shard, not a failure. Combine requires both 3.13 CPU shards and both GPU shards before writingframework-coverage.json. Coverage reports store paths relative to the repo root. macOS and Windows run the full suite and upload junit. Latent-encoder function-preserving checks allow float32 GEMM noise (rtol=1e-5).
Full Changelog: v2.36.0...v2.36.1
v2.36.0: Train vision-language models with image observations and vision-tower LoRA¶
Released on 2026-09-27 - GitHub - PyPI
Other
-
Train vision-language models with image observations and vision-tower LoRA
Image turns pass pixel values through rollout and into GRPO. LoRA adapter keys for the vision encoder and projector are remapped onto the module names used for inference. Vision weights load from the checkpoint conversion registered on the vision module.
Full Changelog: v2.35.0...v2.36.0
agilerl-arena/v1.9.0: agilerl-arena v1.9.0: Train vision-language models with image observations and vision-tower LoRA¶
Released on 2026-09-27 - GitHub - PyPI
Other
-
Train vision-language models with image observations and vision-tower LoRA
Image turns pass pixel values through rollout and into GRPO. LoRA adapter keys for the vision encoder and projector are remapped onto the module names used for inference. Vision weights load from the checkpoint conversion registered on the vision module.
Full Changelog: agilerl-arena/v1.8.0...agilerl-arena/v1.9.0
v2.35.0: train Nemotron Super VL on the FSDP language tower¶
Released on 2026-09-27 - GitHub - PyPI
Features
-
train Nemotron Super VL on the FSDP language tower
Catalog
nemotron_h_omniand serve the language tower asNemotronHOmniLanguageForCausalLM. FSDP finds that tower underlanguage_model, including abackbonebody, and uses it for wrap units, gradient checkpointing, andlm_head.trust_remote_codecomes from the family catalog.model_typeis read fromconfig.json.
Full Changelog: v2.34.1...v2.35.0
agilerl-arena/v1.8.0: agilerl-arena v1.8.0: train Nemotron Super VL on the FSDP language tower¶
Released on 2026-09-27 - GitHub - PyPI
Features
-
train Nemotron Super VL on the FSDP language tower
Catalog
nemotron_h_omniand serve the language tower asNemotronHOmniLanguageForCausalLM. FSDP finds that tower underlanguage_model, including abackbonebody, and uses it for wrap units, gradient checkpointing, andlm_head.trust_remote_codecomes from the family catalog.model_typeis read fromconfig.json.
Full Changelog: agilerl-arena/v1.7.0...agilerl-arena/v1.8.0
v2.34.1: give Rainbow DQN and Recurrent PPO their own schema names and shortlist defaults¶
Released on 2026-09-27 - GitHub - PyPI
Features
-
give Rainbow DQN and Recurrent PPO their own schema names and shortlist defaults
Rainbow DQN and Recurrent PPO are distinct agilerl-arena algorithm specs. Recurrent PPO keeps recurrent on; a payload that turns it off is rejected.
Manifest defaults now match the public shortlist: DQN double, PPO and IPPO action_std_init 0.6, Rainbow value support [-10, 10], LLM use_liger_loss, LLMPPO gae_lambda 0.95, LLMPPO and LLMREINFORCE beta 0.001, gym num_envs 32, evolvable latent_dim 128, and episode_steps left unset.
Algorithm constructors use the same defaults, so an omitted field means the same locally and on Arena.
The published JSON Schema lists algorithm-conditional training defaults on the manifest root (if algorithm.name, then training.properties): DQN epsilon, LLM reporting_interval 1, LLMPPO/LLMREINFORCE max_steps 200, SFT/DPO num_epochs 1, and off-policy experience_sharing. A PPO document does not pick up the LLMPPO or off-policy defaults.
Bandit trainer kwargs omit unset episode_steps so the loop keeps its 500 default.
Full Changelog: v2.34.0...v2.34.1
agilerl-arena/v1.7.0: agilerl-arena v1.7.0: give Rainbow DQN and Recurrent PPO their own schema names and shortlist defaults¶
Released on 2026-09-27 - GitHub - PyPI
Features
-
give Rainbow DQN and Recurrent PPO their own schema names and shortlist defaults
Rainbow DQN and Recurrent PPO are distinct agilerl-arena algorithm specs. Recurrent PPO keeps recurrent on; a payload that turns it off is rejected.
Manifest defaults now match the public shortlist: DQN double, PPO and IPPO action_std_init 0.6, Rainbow value support [-10, 10], LLM use_liger_loss, LLMPPO gae_lambda 0.95, LLMPPO and LLMREINFORCE beta 0.001, gym num_envs 32, evolvable latent_dim 128, and episode_steps left unset.
Algorithm constructors use the same defaults, so an omitted field means the same locally and on Arena.
The published JSON Schema lists algorithm-conditional training defaults on the manifest root (if algorithm.name, then training.properties): DQN epsilon, LLM reporting_interval 1, LLMPPO/LLMREINFORCE max_steps 200, SFT/DPO num_epochs 1, and off-policy experience_sharing. A PPO document does not pick up the LLMPPO or off-policy defaults.
Bandit trainer kwargs omit unset episode_steps so the loop keeps its 500 default.
Full Changelog: agilerl-arena/v1.6.0...agilerl-arena/v1.7.0
v2.34.0: split env factory from entrypoint in the training manifest¶
Released on 2026-09-25 - GitHub - PyPI
Features
-
split env factory from entrypoint in the training manifest
environment.entrypointis the env to build. Optionalenvironment.factoryis the callable that receives that entrypoint.env_configis leftover constructor kwargs only;env_idinside it is rejected.GymEnvSpec and LLMEnvSpec gain
factory. Construction isfactory(entrypoint, **env_config)when factory is set, otherwiseentrypoint(**env_config). LLMnameis a catalog label (use.labelfor display), not a dataset source. Gym still acceptsname: CartPole-v1with implied gymnasium.make.agilerl.llm_envs.env_specsis nowagilerl.llm_envs.env_sources.
Full Changelog: v2.33.0...v2.34.0
agilerl-arena/v1.6.0: agilerl-arena v1.6.0: split env factory from entrypoint in the training manifest¶
Released on 2026-09-25 - GitHub - PyPI
Features
-
split env factory from entrypoint in the training manifest
environment.entrypointis the env to build. Optionalenvironment.factoryis the callable that receives that entrypoint.env_configis leftover constructor kwargs only;env_idinside it is rejected.GymEnvSpec and LLMEnvSpec gain
factory. Construction isfactory(entrypoint, **env_config)when factory is set, otherwiseentrypoint(**env_config). LLMnameis a catalog label (use.labelfor display), not a dataset source. Gym still acceptsname: CartPole-v1with implied gymnasium.make.agilerl.llm_envs.env_specsis nowagilerl.llm_envs.env_sources.
Full Changelog: agilerl-arena/v1.5.0...agilerl-arena/v1.6.0
v2.33.0: shard LLM trainers with PyTorch FSDP2¶
Released on 2026-09-25 - GitHub - PyPI
Features
-
shard LLM trainers with PyTorch FSDP2
LLM trainers shard the actor with
torch.distributedand PyTorch FSDP2 (fully_shard). Omitalgorithm.fsdp/fsdp_configfor flat data-parallel replicas.Rank, device, and collective helpers live in
agilerl.distributed. Importgather_tensor,aggregate_metrics_across_gpus, andaggregate_metrics_dictfrom there; they are no longer inagilerl.utils.llm_utils, andsafe_aggregate_metricsis removed.LLM algorithms use
DPRuntime(full replica per rank) withoutfsdp_configandFSDPRuntimewith it. Wrap units are outermost HuggingFace_no_split_modules, then an untiedlm_head, then the root. LoRA adapters stay replicated. Optimizer states can live on CPU viaoptim_cpu_offloadorcpu_offload.
FSDPConfiglives inagilerl.arena.models.fsdpsoagilerl-arenainstalls without torch;agilerl.distributed.FSDPConfigstill works.
agilerl.utils.llm_utils.get_state_dictis removed; use the algorithm'sshard_runtime.export_model_state.
FSDPConfigfields carry descriptions, and the manifest JSON schema foralgorithm.fsdpis generated from them.
agilerl train --use-acceleratoris removed; accelerate launch is detected automatically.
replay_buffer.buffer_occupancy_multiplieris removed from LLM rollout buffer specs; it had no effect.
use_memory_efficient_params(INIT_HPUSE_MEMORY_EFFICIENT_PARAMS) is renamed tooffload_trainer_during_rollout(OFFLOAD_TRAINER_DURING_ROLLOUT).
Fused logprob helpers takehead_w/head_binstead oflm_head_weight/lm_head_bias.
AutoModelForCausalLMWithValueHead.state_dict()now returns the standard module state dict (pretrained_model.*andv_head.*) for PEFT and non-PEFT backbones.save_pretrainedoutput is unchanged.
BaseRuntime.backwardreturns anOptimizerStep(pre/post-clip gradient norms and the new learning rate) on step boundaries, elseNone.
.population()copies LoRA agents withclone=Trueso adapters are not re-attached onto a PeftModel.
Full Changelog: v2.32.0...v2.33.0
agilerl-arena/v1.5.0: agilerl-arena v1.5.0: per-learn LLM telemetry¶
Released on 2026-09-25 - GitHub - PyPI
Features
-
per-learn LLM telemetry
GRPO learn now reports per-learn advantage stats, policy entropy, K3 KL vs reference and vs rollout policy, importance-ratio tails with advantage-signed clip fractions, and grad norm pre/post clip; vLLM sampling-mismatch stats reach the metrics tracker (also fixed in PPO/REINFORCE). The fused-kernel aux is split into fixed
kl(NaN at beta=0) andclipfrackeys, replacingaux_metric_name. Two correctness fixes: the K3 helper had its arguments flipped vs Schulman/TRL/Liger (Liger runs were unaffected — the kernel computes its own KL), and standard-path CISPO now clamps from above only like the kernel. -
shard LLM trainers with PyTorch FSDP2
LLM trainers shard the actor with
torch.distributedand PyTorch FSDP2 (fully_shard). Omitalgorithm.fsdp/fsdp_configfor flat data-parallel replicas.Rank, device, and collective helpers live in
agilerl.distributed. Importgather_tensor,aggregate_metrics_across_gpus, andaggregate_metrics_dictfrom there; they are no longer inagilerl.utils.llm_utils, andsafe_aggregate_metricsis removed.LLM algorithms use
DPRuntime(full replica per rank) withoutfsdp_configandFSDPRuntimewith it. Wrap units are outermost HuggingFace_no_split_modules, then an untiedlm_head, then the root. LoRA adapters stay replicated. Optimizer states can live on CPU viaoptim_cpu_offloadorcpu_offload.
FSDPConfiglives inagilerl.arena.models.fsdpsoagilerl-arenainstalls without torch;agilerl.distributed.FSDPConfigstill works.
agilerl.utils.llm_utils.get_state_dictis removed; use the algorithm'sshard_runtime.export_model_state.
FSDPConfigfields carry descriptions, and the manifest JSON schema foralgorithm.fsdpis generated from them.
agilerl train --use-acceleratoris removed; accelerate launch is detected automatically.
replay_buffer.buffer_occupancy_multiplieris removed from LLM rollout buffer specs; it had no effect.
use_memory_efficient_params(INIT_HPUSE_MEMORY_EFFICIENT_PARAMS) is renamed tooffload_trainer_during_rollout(OFFLOAD_TRAINER_DURING_ROLLOUT).
Fused logprob helpers takehead_w/head_binstead oflm_head_weight/lm_head_bias.
AutoModelForCausalLMWithValueHead.state_dict()now returns the standard module state dict (pretrained_model.*andv_head.*) for PEFT and non-PEFT backbones.save_pretrainedoutput is unchanged.
BaseRuntime.backwardreturns anOptimizerStep(pre/post-clip gradient norms and the new learning rate) on step boundaries, elseNone.
.population()copies LoRA agents withclone=Trueso adapters are not re-attached onto a PeftModel.
Fixes
-
close non-LLM training resources and stop clone memory leaks
Training loops now finish loggers, progress bars, and environments when training returns.
BanditEnvandDatasetEnvimplementclose().Tournament selection and MF-PBT free evicted agents.
clone()skips rollout buffers, copies GraMa scores, and restores the parent Accelerate wrap whenwrap=True. Bandit clones trim regret history.minari_to_agile_datasetwrites HDF5 and returns the file path. Offline training closes HDF5 handles after loading transitions; in-memory array mappings are unchanged.test()puts networks back in training mode after a successful evaluation. -
pin vllm extra to 0.25.1
Pin the Linux llm extra to vLLM 0.25.1. vLLM 0.26 and later crash on Nemotron-H hybrid models when CUDA graphs are enabled.
Full Changelog: agilerl-arena/v1.4.0...agilerl-arena/v1.5.0
v2.32.0: per-learn LLM telemetry¶
Released on 2026-09-24 - GitHub - PyPI
Features
-
per-learn LLM telemetry
GRPO learn now reports per-learn advantage stats, policy entropy, K3 KL vs reference and vs rollout policy, importance-ratio tails with advantage-signed clip fractions, and grad norm pre/post clip; vLLM sampling-mismatch stats reach the metrics tracker (also fixed in PPO/REINFORCE). The fused-kernel aux is split into fixed
kl(NaN at beta=0) andclipfrackeys, replacingaux_metric_name. Two correctness fixes: the K3 helper had its arguments flipped vs Schulman/TRL/Liger (Liger runs were unaffected — the kernel computes its own KL), and standard-path CISPO now clamps from above only like the kernel.
Full Changelog: v2.31.0...v2.32.0
v2.31.0: pin vllm extra to 0.25.1¶
Released on 2026-09-22 - GitHub - PyPI
Fixes
-
pin vllm extra to 0.25.1
Pin the Linux llm extra to vLLM 0.25.1. vLLM 0.26 and later crash on Nemotron-H hybrid models when CUDA graphs are enabled.
Full Changelog: v2.30.0...v2.31.0
v2.30.0: close non-LLM training resources and stop clone memory leaks¶
Released on 2026-09-18 - GitHub - PyPI
Fixes
-
close non-LLM training resources and stop clone memory leaks
Training loops now finish loggers, progress bars, and environments when training returns.
BanditEnvandDatasetEnvimplementclose().Tournament selection and MF-PBT free evicted agents.
clone()skips rollout buffers, copies GraMa scores, and restores the parent Accelerate wrap whenwrap=True. Bandit clones trim regret history.minari_to_agile_datasetwrites HDF5 and returns the file path. Offline training closes HDF5 handles after loading transitions; in-memory array mappings are unchanged.test()puts networks back in training mode after a successful evaluation.
Full Changelog: v2.29.0...v2.30.0
v2.29.0: reject unknown fields on live pydantic models¶
Released on 2026-09-18 - GitHub - PyPI
Fixes
-
reject unknown fields on live pydantic models
Unknown keys on framework runtime configs now raise a validation error instead of being dropped (
VllmRuntimeConfig,TrainerRuntimeConfig,MambaPatchConfig,PatchRuntimeConfig,ModelRuntimeConfig). Arena inference request and status DTOs (LLMParams,AgentInfo,StatusResponse,LLMResults,SessionMessage) do the same.PredictResultandSessionInfoignore extra keys because they parse a subset of the inference JSON (resultson/predict;title/created_byon/sessions).
Full Changelog: v2.28.0...v2.29.0
agilerl-arena/v1.4.0: agilerl-arena v1.4.0: reject unknown fields on live pydantic models¶
Released on 2026-09-18 - GitHub - PyPI
Fixes
-
reject unknown fields on live pydantic models
Unknown keys on framework runtime configs now raise a validation error instead of being dropped (
VllmRuntimeConfig,TrainerRuntimeConfig,MambaPatchConfig,PatchRuntimeConfig,ModelRuntimeConfig). Arena inference request and status DTOs (LLMParams,AgentInfo,StatusResponse,LLMResults,SessionMessage) do the same.PredictResultandSessionInfoignore extra keys because they parse a subset of the inference JSON (resultson/predict;title/created_byon/sessions).
Full Changelog: agilerl-arena/v1.3.0...agilerl-arena/v1.4.0
v2.28.0: drop use_vllm from the training spec¶
Released on 2026-09-17 - GitHub - PyPI
Features
-
drop use_vllm from the training spec
Training YAML and the algorithm constructor no longer have
use_vllm. Colocated rollout (training.rollout_mode) fills a defaultvllm_configwhen it is unset; that config is what starts the in-process engine. Example GRPO / PPO / REINFORCE / GSPO / CISPO configs and the remote env tutorial drop the flag.
Full Changelog: v2.27.1...v2.28.0
agilerl-arena/v1.3.0: agilerl-arena v1.3.0: resolve trainer attention in one helper¶
Released on 2026-09-17 - GitHub - PyPI
Features
-
resolve trainer attention in one helper
resolve_attn_implementationnow picks the trainer attention backend in one place. Order: an explicit value, thenATTN_IMPLEMENTATION/AGILERL_ATTN_IMPLEMENTATION, then the family trainer default (Gemma sliding-windowflex_attention, Nemotron-Hflash_attention_2), thenflash_attention_2ifflash_attnis installed otherwisesdpa.LLMBuilderandcreate_model_from_name_or_pathboth go through that helper. Nemotron-H's catalog trainer config setsattn_implementationtoflash_attention_2. -
attach reference and critic adapters with packed-expert LoRA
PEFT hosts multiple target_parameters adapters on the same expert weights. GRPO/CISPO can keep a frozen reference adapter for KL, and a value head can attach a trainable critic. Mixed fused routing works on both routed (Nemotron-H) and sorted (JetMoE) packed-expert modules.
-
drop use_vllm from the training spec
Training YAML and the algorithm constructor no longer have
use_vllm. Colocated rollout (training.rollout_mode) fills a defaultvllm_configwhen it is unset; that config is what starts the in-process engine. Example GRPO / PPO / REINFORCE / GSPO / CISPO configs and the remote env tutorial drop the flag.
Fixes
-
pin deepspeed to 0.19.2
Pin the llm extra to DeepSpeed 0.19.2. 0.19.3+ (#8148) leaves ZeRO-3 frozen params gathered after activation-checkpoint recompute.
-
cap vLLM max_num_batched_tokens at seqs times context length
Family and explicit
max_num_batched_tokensvalues are capped atmax_num_seqs * max_model_len. vLLM cannot schedule more tokens than that product.
Full Changelog: agilerl-arena/v1.2.0...agilerl-arena/v1.3.0
v2.27.1: cap vLLM max_num_batched_tokens at seqs times context length¶
Released on 2026-09-16 - GitHub - PyPI
Fixes
-
cap vLLM max_num_batched_tokens at seqs times context length
Family and explicit
max_num_batched_tokensvalues are capped atmax_num_seqs * max_model_len. vLLM cannot schedule more tokens than that product.
Full Changelog: v2.27.0...v2.27.1
v2.26.1: pin deepspeed to 0.19.2¶
Released on 2026-09-16 - GitHub - PyPI
Fixes
-
pin deepspeed to 0.19.2
Pin the llm extra to DeepSpeed 0.19.2. 0.19.3+ (#8148) leaves ZeRO-3 frozen params gathered after activation-checkpoint recompute.
Full Changelog: v2.26.0...v2.26.1
v2.26.0: resolve trainer attention in one helper¶
Released on 2026-09-15 - GitHub - PyPI
Features
-
resolve trainer attention in one helper
resolve_attn_implementationnow picks the trainer attention backend in one place. Order: an explicit value, thenATTN_IMPLEMENTATION/AGILERL_ATTN_IMPLEMENTATION, then the family trainer default (Gemma sliding-windowflex_attention, Nemotron-Hflash_attention_2), thenflash_attention_2ifflash_attnis installed otherwisesdpa.LLMBuilderandcreate_model_from_name_or_pathboth go through that helper. Nemotron-H's catalog trainer config setsattn_implementationtoflash_attention_2.
Full Changelog: v2.25.0...v2.26.0
v2.24.0: load family trainer and vLLM defaults by Hugging Face model_type¶
Released on 2026-09-14 - GitHub - PyPI
Features
-
load family trainer and vLLM defaults by Hugging Face model_type
Callers pass a checkpoint id or path to
family_runtime, which readsAutoConfig.model_typeand applies catalog defaults withsetdefault. Missingconfig.jsonraises. Family patches use a loadedconfig.model_typewhen present, otherwise the Hugging Face id.nemotron_hgets vLLMmamba_cache_mode=align,max_num_batched_tokens=8192,reasoning_parser=nemotron_v3, andenable_prefix_caching=True. Gemma 3 and 4 types get trainerattn_implementation=flex_attention. Explicit yaml /vllm_configvalues still win.
Full Changelog: v2.23.0...v2.24.0
v2.23.0: rename RLAlgorithm to SingleAgentAlgorithm¶
Released on 2026-09-14 - GitHub - PyPI
Refactoring
-
rename RLAlgorithm to SingleAgentAlgorithm
Rename
RLAlgorithmtoSingleAgentAlgorithmandMultiAgentRLAlgorithmtoMultiAgentAlgorithm. Spec bases follow:RLAlgorithmSpec/MultiAgentRLAlgorithmSpecbecomeSingleAgentAlgorithmSpec/MultiAgentAlgorithmSpec. The old algorithm class names remain importable fromagilerl.algorithms.core.
Full Changelog: v2.22.0...v2.23.0
agilerl-arena/v1.1.0: agilerl-arena v1.1.0: rename RLAlgorithm to SingleAgentAlgorithm¶
Released on 2026-09-14 - GitHub - PyPI
Refactoring
-
rename RLAlgorithm to SingleAgentAlgorithm
Rename
RLAlgorithmtoSingleAgentAlgorithmandMultiAgentRLAlgorithmtoMultiAgentAlgorithm. Spec bases follow:RLAlgorithmSpec/MultiAgentRLAlgorithmSpecbecomeSingleAgentAlgorithmSpec/MultiAgentAlgorithmSpec. The old algorithm class names remain importable fromagilerl.algorithms.core.
Full Changelog: agilerl-arena/v1.0.0...agilerl-arena/v1.1.0
v2.22.0: define training specs only in agilerl-arena¶
Released on 2026-09-14 - GitHub - PyPI
Features
-
define training specs only in agilerl-arena
The training manifest lives in
agilerl.arena.models. The framework imports those classes as the specs; it does not subclass them to addmake_env/init_buffer/build. Builders and strategies sit beside the specs. Unknown keys are rejected. Defaults match the algorithm constructors.evo_stepsis optional.LocalTrainertakes networks and HPO as arguments. New CLI:arena manifest validateandarena manifest schema.agilerlnow depends onagilerl-arena>=1.0.0,<2.0.
Full Changelog: v2.21.1...v2.22.0
agilerl-arena/v1.0.0: agilerl-arena v1.0.0: build algorithms from specs via paradigm builders¶
Released on 2026-09-11 - GitHub - PyPI
Features
-
build algorithms from specs via paradigm builders
spec.build_algorithm()delegates toagilerl.builders. Specs remain arena field subclasses with construction wrappers. Training loops still come from the spec. Builderbuild()takes anAlgorithmBuildRuntimefor the population slot, device, HPO, and checkpoint. -
dispatch local training through paradigm strategies
Training loops are selected by
agilerl.strategies.select_strategyfrom the spec's paradigm flags (off_policy,offline,bandit,env_type).LocalTraineruses that layer. Specs still exposeget_training_fnandget_training_kwargs. Multi-agent fitness logs take a per-agent dict. -
define training specs only in agilerl-arena
The training manifest lives in
agilerl.arena.models. The framework imports those classes as the specs; it does not subclass them to addmake_env/init_buffer/build. Builders and strategies sit beside the specs. Unknown keys are rejected. Defaults match the algorithm constructors.evo_stepsis optional.LocalTrainertakes networks and HPO as arguments. New CLI:arena manifest validateandarena manifest schema.agilerlnow depends onagilerl-arena>=1.0.0,<2.0.
Fixes
-
isolate dummy algorithm specs from the global registry
A unit test no longer leaves a dummy spec on the global algorithm registry, which made later tests that walk every registered spec fail depending on collection order. Training strategy types now include LLM and bandit envs and match the fitness values the loops return.
Other
-
bump peft to 0.20.0 and liger-kernel to 0.8.2
PEFT 0.20 rejects LoRA on Mamba mixer out_proj and conv1d. adapt_lora_config_for_model excludes those modules. hydra-core stays on 1.3.x. LLMAlgorithm backward stays under AMP so fp16 checkpoint recompute matches LoRA dtypes.
Full Changelog: agilerl-arena/v0.9.0...agilerl-arena/v1.0.0
v2.21.1: isolate dummy algorithm specs from the global registry¶
Released on 2026-09-10 - GitHub - PyPI
Fixes
-
isolate dummy algorithm specs from the global registry
A unit test no longer leaves a dummy spec on the global algorithm registry, which made later tests that walk every registered spec fail depending on collection order. Training strategy types now include LLM and bandit envs and match the fitness values the loops return.
Full Changelog: v2.21.0...v2.21.1
v2.21.0: dispatch local training through paradigm strategies¶
Released on 2026-09-10 - GitHub - PyPI
Features
-
dispatch local training through paradigm strategies
Training loops are selected by
agilerl.strategies.select_strategyfrom the spec's paradigm flags (off_policy,offline,bandit,env_type).LocalTraineruses that layer. Specs still exposeget_training_fnandget_training_kwargs. Multi-agent fitness logs take a per-agent dict.
Full Changelog: v2.20.0...v2.21.0
agilerl-arena/v0.9.0: agilerl-arena v0.9.0: clamp generation to remaining context per turn¶
Released on 2026-09-08 - GitHub - PyPI
Features
-
subclass framework algorithm specs from agilerl-arena field models
Framework algorithm specs now subclass the agilerl-arena pydantic field models and keep construction (
build_algorithm,get_training_fn). Network encoder specs are the arena classes. Installing agilerl requires agilerl-arena>=0.9.0,<1.0. Builders, strategies, and LocalTrainer are unchanged.
Fixes
-
publish to PyPI on tag push and wait for agilerl-arena
Pushing a
v*oragilerl-arena/v*tag now runs Publish release for that tag,
so every released version reaches PyPI without a second manual step. Running
the workflow by hand from a tag still works, for backfilling a version PyPI is
missing or retrying a failed run.Before uploading
agilerl, the publish job reads theagilerl-arenarange
from the wheel's ownRequires-Distand waits for a matching version to appear
on PyPI. The two tags are pushed together and the index takes a moment to serve
a new release, so theagilerlupload waits foragilerl-arenainstead of
failing. Anagilerl-arenawheel declares no such requirement and never waits.This closes a real gap:
agilerl2.16.1 shipped declaring
agilerl-arena>=0.7.0,<0.8when no such version was on PyPI, leaving its
arenaextra uninstallable. -
bump the arena extra from the agilerl-arena version kind
The committed agilerl-arena extra range must cover the version this release will tag. That version now follows the arena package's own bump kind, not a shared stack bump.
Breaking Changes
-
clamp generation to remaining context per turn
Each turn generates min(configured max_output_tokens, remaining context), or remaining context when the cap is unset. Setup no longer rejects a per-turn cap larger than the window. min_new_tokens clamps to that budget so it cannot exceed max_new_tokens.
RolloutHarness and make_rollout_env_factory no longer take max_output_tokens. The harness only truncates when the prompt fills max_model_len. max_prompt_tokens_for_model_len now takes only max_model_len. validate_llm_context_lengths is removed.
Full Changelog: agilerl-arena/v0.8.1...agilerl-arena/v0.9.0
v2.18.2: clamp generation to remaining context per turn¶
Released on 2026-09-08 - GitHub - PyPI
Breaking Changes
-
clamp generation to remaining context per turn
Each turn generates min(configured max_output_tokens, remaining context), or remaining context when the cap is unset. Setup no longer rejects a per-turn cap larger than the window. min_new_tokens clamps to that budget so it cannot exceed max_new_tokens.
RolloutHarness and make_rollout_env_factory no longer take max_output_tokens. The harness only truncates when the prompt fills max_model_len. max_prompt_tokens_for_model_len now takes only max_model_len. validate_llm_context_lengths is removed.
Full Changelog: v2.18.1...v2.18.2
v2.18.1: align ArenaClient parquet config names with prefix stripping¶
Released on 2026-09-07 - GitHub - PyPI
Fixes
-
align ArenaClient parquet config names with prefix stripping
ArenaClient strips shared directories from parquet folder uploads while every remaining path still has more than one component. Split directories such as train and test are config names. A flat parquet tree uses config default. Tabular dataset uploads stay rejected.
-
publish to PyPI on tag push and wait for agilerl-arena
Pushing a
v*oragilerl-arena/v*tag now runs Publish release for that tag,
so every released version reaches PyPI without a second manual step. Running
the workflow by hand from a tag still works, for backfilling a version PyPI is
missing or retrying a failed run.Before uploading
agilerl, the publish job reads theagilerl-arenarange
from the wheel's ownRequires-Distand waits for a matching version to appear
on PyPI. The two tags are pushed together and the index takes a moment to serve
a new release, so theagilerlupload waits foragilerl-arenainstead of
failing. Anagilerl-arenawheel declares no such requirement and never waits.This closes a real gap:
agilerl2.16.1 shipped declaring
agilerl-arena>=0.7.0,<0.8when no such version was on PyPI, leaving its
arenaextra uninstallable. -
bump the arena extra from the agilerl-arena version kind
The committed agilerl-arena extra range must cover the version this release will tag. That version now follows the arena package's own bump kind, not a shared stack bump.
Full Changelog: v2.18.0...v2.18.1
agilerl-arena/v0.8.0: agilerl-arena v0.8.0: write Arena completeness files after merged LoRA export¶
Released on 2026-09-07 - GitHub - PyPI
Features
-
write Arena completeness files after merged LoRA export
Successful
export_merged_pretrainedclearsoutput_dir, then writes
arena_artifact_manifest.json(formathf_merged, sha256 and byte size
per file) and.complete.adapter_config.jsonis omitted fromfiles[].
ZeRO layer-wise gather is unchanged. -
strip Hugging Face parquet prefixes and add organisation-key auth
ArenaClient.create_dataset strips shared Hugging Face folder prefixes before
choosing a parquet config. Partner servers can authenticate with an organisation
key plus X-External-User-Id. Inference agents opened from an org-key client send
the same headers. PAT and Keycloak device login are unchanged. The
agilerl[arena] extra is agilerl-arena>=0.8.0,<0.9.
Other
-
gate agilerl PyPI publish on a resolvable agilerl-arena pin
Refuse to upload an agilerl wheel whose agilerl-arena requirement is not yet on PyPI. Retryable publish skips artifacts already on the index.
Full Changelog: agilerl-arena/v0.7.1...agilerl-arena/v0.8.0
v2.17.0: write Arena completeness files after merged LoRA export¶
Released on 2026-09-07 - GitHub - PyPI
Features
-
write Arena completeness files after merged LoRA export
Successful
export_merged_pretrainedclearsoutput_dir, then writes
arena_artifact_manifest.json(formathf_merged, sha256 and byte size
per file) and.complete.adapter_config.jsonis omitted fromfiles[].
ZeRO layer-wise gather is unchanged.
Tests
-
close Rich Live in arena output tests
Stop a leftover Live renderer after the check-event test, and use dummy Live objects in select_row tests so they do not start a real Rich Live.
Other
-
gate agilerl PyPI publish on a resolvable agilerl-arena pin
Refuse to upload an agilerl wheel whose agilerl-arena requirement is not yet on PyPI. Retryable publish skips artifacts already on the index.
Full Changelog: v2.16.1...v2.17.0
v2.16.1: add ZeRO-aware LoRA merge to Hugging Face format¶
Released on 2026-09-06 - GitHub - PyPI
Features
-
add ZeRO-aware LoRA merge to Hugging Face format
Add
export_merged_pretrained, a public helper that folds a LoRA adapter
into base weights and writes a Hugging Face directory (config, bf16
safetensors, optional tokenizer and generation config) without gathering
the full model onto one rank.Works from a live PEFT module (DeepSpeed ZeRO-3 gathers one module at a
time) or from a storedactor/adapter plus a base model path. Replay
loads a meta skeleton and streams base weights from safetensors. The live
module is not mutated. A failed collective or write leaves every rank
able to continue training.
Full Changelog: v2.16.0...v2.16.1
agilerl-arena/v0.7.0: agilerl-arena v0.7.0: upload local parquet files and HF shard folders¶
Released on 2026-09-07 - GitHub - PyPI
Features
-
upload local parquet files and HF shard folders
ArenaClient.create_dataset and
arena datasets create --filenow upload a.parquetfile with the parquet content type, or walk a Hugging Face layout folder and send shards as repeated multipartfileparts with relative paths. Passconfig=/--configwhen the folder has more than one parquet config. CSV uploads are unchanged.The
agilerl[arena]extra isagilerl-arena>=0.7.0,<0.8.
Full Changelog: agilerl-arena/v0.6.0...agilerl-arena/v0.7.0
v2.16.0: upload local parquet files and HF shard folders¶
Released on 2026-09-06 - GitHub - PyPI
Features
-
upload local parquet files and HF shard folders
ArenaClient.create_dataset and
arena datasets create --filenow upload a.parquetfile with the parquet content type, or walk a Hugging Face layout folder and send shards as repeated multipartfileparts with relative paths. Passconfig=/--configwhen the folder has more than one parquet config. CSV uploads are unchanged.The
agilerl[arena]extra isagilerl-arena>=0.7.0,<0.8.
Full Changelog: v2.15.0...v2.16.0
agilerl-arena/v0.6.0: agilerl-arena v0.6.0: install architecture patches through one family installer¶
Released on 2026-09-07 - GitHub - PyPI
Features
-
install architecture patches through one family installer
detect_model_familyreturns a single family key (or None).FAMILY_PATCHESmaps each family to one callable. Nemotron-H usesinstall_nemotron_h_patches, which applies the Mamba fused-path and stream-ordering workarounds on every ZeRO stage.install_family_patchesreturns that family key or None. -
add model catalog and default RunSpec discovery to ArenaClient
ArenaClient can list the HuggingFace model catalog, fetch model info (LoRA modules and context length), and load a default RunSpec from the CLI API.
The
arenaextra isagilerl-arena>=0.6.0,<0.7.
Fixes
-
publish from the tag ref and open a GitHub Release after PyPI
Run Publish release from the version tag (Use workflow from). After PyPI
accepts the upload, the workflow opens a GitHub Release with generated notes
and the dist files attached. -
build GitHub Release notes from spoke commits
Publish release opens a GitHub Release with a
vX.Y.Z: headlinetitle and
notes grouped from the commits since the last tag (Features, Fixes, and so on),
not GitHub’s PR-only generate-notes button.
Tests
-
cover encoder layer_norm disable path
Initialize TD3 with
encoder_config.layer_norm=Trueso the warning that disables it is covered.
Other
-
publish PyPI from workflow_dispatch only
Tag push no longer publishes. Create a GitHub Release from an existing
v*oragilerl-arena/v*tag withworkflow_dispatch. Minting those tags from hub does not upload to PyPI. -
ReGraMa & Amplified-Gaussian / Random-Reset Parameter Mutations Switches
-
Function-Preserving Node Addition & Layer Addition Architecture Mutations
-
pin Node 24 Actions and skip uv cache on publish
Publish release uses Node 24 action pins. The publish job does not check
out the repo, so uv cache is off. -
Chore(deps): Bump pygame-ce from 2.5.7 to 2.5.8
-
stop GitHub Dependabot on the public clone
Dependency updates for Python packages are handled on the internal hub. This clone no longer ships a
.github/dependabot.ymlconfig.
Full Changelog: agilerl-arena/v0.5.0...agilerl-arena/v0.6.0
v2.15.0: add model catalog and default RunSpec discovery to ArenaClient¶
Released on 2026-09-06 - GitHub - PyPI
Features
-
add model catalog and default RunSpec discovery to ArenaClient
ArenaClient can list the HuggingFace model catalog, fetch model info (LoRA modules and context length), and load a default RunSpec from the CLI API.
The
arenaextra isagilerl-arena>=0.6.0,<0.7.
Other
-
stop GitHub Dependabot on the public clone
Dependency updates for Python packages are handled on the internal hub. This clone no longer ships a
.github/dependabot.ymlconfig.
Full Changelog: v2.14.3...v2.15.0
v2.14.3: install architecture patches through one family installer¶
Released on 2026-09-06 - GitHub - PyPI
Features
-
install architecture patches through one family installer
detect_model_familyreturns a single family key (or None).FAMILY_PATCHESmaps each family to one callable. Nemotron-H usesinstall_nemotron_h_patches, which applies the Mamba fused-path and stream-ordering workarounds on every ZeRO stage.install_family_patchesreturns that family key or None.
Full Changelog: v2.14.2...v2.14.3
v2.14.2: Chore(deps): Bump pygame-ce from 2.5.7 to 2.5.8¶
Released on 2026-09-03 - GitHub - PyPI
Fixes
-
build GitHub Release notes from spoke commits
Publish release opens a GitHub Release with a
vX.Y.Z: headlinetitle and
notes grouped from the commits since the last tag (Features, Fixes, and so on),
not GitHub’s PR-only generate-notes button.
Other
-
pin Node 24 Actions and skip uv cache on publish
Publish release uses Node 24 action pins. The publish job does not check
out the repo, so uv cache is off. -
Chore(deps): Bump pygame-ce from 2.5.7 to 2.5.8
Full Changelog: v2.14.1...v2.14.2
v2.14.0: Function-preserving node & layer additions¶
Released on 2026-09-01 - GitHub - PyPI
Architecture mutations that add capacity now keep the network’s function when the architecture allows it, so a
widened or deepened agent is not immediately a different policy.
Features
- Widening (
add_node,add_channel,add_latent_node): new incoming weights stay as the original operator
set them; outgoing weights into the next layer are initialised to small noise (~2% of that layer’s existing column
scale). The consumer’s output is then unchanged whatever the activation does, and the new units still receive
gradient. - Deepening (
add_layer): the inserted layer is initialised to the identity (Net2Net), which is exact for ReLU and Identity. - No new flag. Preservation is applied automatically when it can be, and the original random init is used when it
cannot. Removals (remove_node,remove_channel,remove_latent_node,remove_layer) are unchanged. - Preservation stands down when a norm layer sits between the new units and their consumer, when the activation
mixes units (Softmax,LogSoftmax,Softmin,GumbelSoftmax), or when the layer is inside an RNN core,
multi-input encoder, residual, or SimBa block.add_layeralso needs a square MLP under ReLU or Identity. Each
reason is warned once perMutationsinstance. - Example manifests (layer norm off, activation mutation off) ship for PPO, DQN, CQN, NeuralUCB, IPPO, and MADDPG
underconfigs/training/**/*_func_preserving.yaml.
Breaking Changes
None.
What's Changed
- Function-Preserving Node Addition & Layer Addition Architecture Mutations by @agilerl-hub-sync in
#690
Full Changelog: v2.13.0...v2.14.0
v2.12.0¶
Released on 2026-09-01 - GitHub - PyPI
What's Changed
- fix: restore Nemotron-H Mamba out_proj LoRA gradients on pre-built models by @agilerl-hub-sync[bot] in #697
- test: deterministic env-resolution and beam-search tests by @agilerl-hub-sync[bot] in #698
- ci: run Linux, macOS, Windows, and type checks on pull requests only by @agilerl-hub-sync[bot] in #699
- ci: run Linux, macOS, Windows, ty, and CodeQL from one CI workflow by @agilerl-hub-sync[bot] in #700
- chore: raise llm extra vLLM floor to 0.26 by @agilerl-hub-sync[bot] in #701
Full Changelog: v2.11.0...v2.12.0
v2.13.0: ReGraMa dormant-neuron resets¶
Released on 2026-09-01 - GitHub - PyPI
Parameter mutation now resets dormant neurons before the Gaussian pass, using ReGraMa (gradient-magnitude
scoring). The old amplified (“super”) Gaussian band is gone.
Features
- ReGraMa runs as the first stage of every parameter mutation. It scores each neuron with the GraMa metric of
Liu et al., “Measure gradients, not activations!” — mean absolute gradient of
the loss w.r.t. the pre-activation, normalised by that layer’s mean. Neurons at or belowdormant_threshold
(default0.01) are reset: Xavier-uniform incoming weights, zero bias, small non-zero outgoing weights, and any
adjacent norm entry restored to the identity. Output layers of heads are never reset. Target / shared networks are
re-synced afterwards. - Capture rides the existing
init_training_step/finalize_training_steppair, so on-policy, off-policy,
multi-agent, bandit, and offline trainers all get scores with no extra forward/backward pass. LLM algorithms still
skip parameter mutation. - Sensitivity is one field on the existing
mutationblock:The same argument exists onmutation: dormant_threshold: 0.01
Mutations(...). Existing manifests and constructor calls keep working.
Changes
- The Gaussian pass no longer has a “super” band (10× noise on ~5% of sampled weights). Of the 10% of weights
sampled for mutation, 95% get ordinary noise scaled bymutation_sdand the weight’s own magnitude; 5%
are redrawn fromN(0, 1). The split is fixed.
CI
- Pushing a
v*oragilerl-arena/v*tag no longer publishes to PyPI. Create a GitHub Release from an existing
tag withworkflow_dispatchon thePublish releaseworkflow.
What's Changed
- ReGraMa & Amplified-Gaussian / Random-Reset Parameter Mutations Switches by @agilerl-hub-sync in
#685 - ci: publish PyPI from workflow_dispatch only
Full Changelog: v2.12.0...v2.13.0
v2.11.0: LLM training environments use OpenEnv 🤗¶
Released on 2026-08-25 - GitHub - PyPI
This release is one PR, #696: LLM training environments all move onto OpenEnv.
Features
- Every LLM training environment now works the same way: text in, text out, through the OpenEnv API, which ships with the
[llm]extra. An environment is any object withreset(seed) -> (prompt_text, info)andstep(action) -> (observation_text, reward, terminated, truncated, info). That is the whole interface. - You declare
env_type. It is never inferred from the algorithm.rolloutmeans the model generates and the environment scores it, and a single-turn scored task is justmax_turns: 1rather than a type of its own.datasetmeans teacher forcing, and wantsobjective: sftorobjective: preference. - An env can come from dataset rows, a Python
entrypoint, or anenv_urlpointing at a server you already have running. Over a URL one server fronts the whole batch: each rollout opens its own WebSocket session and gets a fresh env built for it. RolloutHarnessruns the env in the training process with no HTTP at all (RolloutHarness.local), or over a URL, behind the same interface either way.advantage_granularity: autoworks the grain out from the batch itself. PPO and REINFORCE send single-turn batches totokenand multi-turn toturn. GRPO has no token-level advantage, so it usestrajectoryandturninstead.- Rubric scoring for QA datasets, in
agilerl.llm_envs.rubrics, plus aTaskAssignerthat spreads seeds across concurrent rollouts. - An env can name the packages it needs in
env_packages. If the entrypoint will not import, those get installed and it tries again. - New
LLMRolloutDatacomponent for rollout trajectories, with a dev docs page. - Docs: an Environments (OpenEnv) page, and a tutorial for pointing training at your own OpenEnv server.
Optimizations
AsyncBatchCollectordrives the in-processRolloutCollectorfrom an asyncio loop. Per-episode calls are still synchronous, so they go to an executor sized to the slot count, and a tokenizer lock keeps tokenizer work serial.
Fixes
- Retired and unrecognized manifest keys fail at startup now instead of being quietly ignored, so a stale spec stops you rather than training something you did not ask for.
- Checkpoint resume, observation roles, token-observation tools and the LLM finetuning demos all have regression tests. None of them were covered in CI before.
Breaking Changes
None of these have a migration alias. Old specs are meant to fail rather than be quietly translated.
LLMEnvType.MULTITURNis nowROLLOUT. The oldsftandpreferenceenv types becomeenv_type: datasetplus anobjective.- GRPO
advantage_granularitytakesauto(the default),trajectoryorturn.tokenis not a GRPO value any more.action_granularitystill aliases the setting. - Trajectory tensors use
token_ids, notcompletion_ids. PromptDatasetEnvis nowQADatasetEnv.agilerl.training.train_llmis gone. There are two loops instead:train_llm_rolloutfor generative rollout RL (GRPO, PPO, REINFORCE) andtrain_llm_datasetfor teacher forcing (SFT, preference). The env type in your spec picks which one runs.- The
llmextra needsopenenv>=0.4.1,<0.5, andagilerl[arena]needsagilerl-arena>=0.4.0,<0.5.
What's Changed
- feat(llm)!: unify LLM environments on OpenEnv (
rolloutvsdataset) by @agilerl-hub-sync[bot] in #696
Full Changelog: v2.10.0...v2.11.0
v2.10.0: Tag-driven releases, PyPI Trusted Publishing, DPO Liger memory fix¶
Released on 2026-08-25 - GitHub - PyPI
CI and Versioning Changes
- Versions now come from git tags.
agilerlandagilerl-arenabuild withhatch-vcs, soproject.versionis gone from bothpyproject.tomlfiles and cutting the release tag is the version bump — no more "Update agilerl version to X" PRs (#683). - New
Publish releaseworkflow. Pushing (or dispatching from) av*oragilerl-arena/v*tag publishes the matching package to PyPI via Trusted Publishing (OIDC). It uploads the artifact already built for that tag rather than rebuilding, and uses short-lived OIDC credentials end to end — no long-lived PyPI or AWS keys. The read-role ARN moved to a repository secret so it is masked in logs and unavailable to fork pull requests (#668, #670, #692). arenaextra is a compatible range, not an exact pin.agilerl[arena]resolvesagilerl-arena>=0.3.0,<0.4, so an arena patch release no longer requires a newagilerlwheel.scripts/check-extras.pygained a--require-tagsmode plus a full test suite to keep the range and theallunion honest (#683, #692, #693).- pre-commit no longer fail-closes when
agilerl-arena/v*tags are absent from a clone, so a fresh fork or shallow clone can run the hooks. Ruff is pinned at 0.16.2 and pre-commit.ci autoupdates targetmain(#683). just check-extrasruns throughuv run python, picking the project interpreter on machines that only exposepython3(#693).pytest-timeoutadded to thedevextra (#668).
Optimizations
- DPO's Liger path no longer materializes a discarded
(batch, seq_len, vocab)logits tensor on every forward:_get_hiddenidentity-patches the LM head instead of capturing its input with a forward pre-hook. At production shapes that transient was 12+ GiB per forward (#678).
Fixes
- Distributed launch env vars (
WORLD_SIZE,RANK,LOCAL_RANK,MASTER_ADDR, …) leaked by DeepSpeed/LLM tests are now cleared around every test, so laterAccelerator()calls stop DDP-wrapping models on Linux or hanging on rendezvous on macOS/Windows. The process group is deliberately left intact — destroying it between consecutive GPU DeepSpeed tests surfacedGroup <ProcessGroup ...> is not registered(#668). - Minari tests handle both the 0.5.2 (
termination/truncation) and 0.5.3+ (terminated/truncated) episode-buffer keys (#668).
Breaking Changes
- The
nightlybranch is retired —mainis now the development branch.pip install git+https://github.com/AgileRL/AgileRL.git@nightlyno longer resolves; use@main. Contributions should branch from and targetmain(#683, #686, #687). agilerl[arena]now requiresagilerl-arena>=0.3.0,<0.4; the 0.2.x series no longer satisfies the extra (#692).- Contributors should leave their pull requests open rather than merging them: a maintainer applies the
hub-sync-importlabel to run internal validation, and automation completes the original PR so authorship and the review thread stay on GitHub (#686, #687).
What's Changed
- Minor test fixes + PyPi publish workflow by @agilerl-hub-sync[bot] in #668
- Harden security on PyPi publish script by @agilerl-hub-sync[bot] in #670
- sync: hub export b2ef6ce7b993 by @agilerl-hub-sync[bot] in #678
- chore: pin ruff 0.16.2 and stop requiring arena git tags in pre-commit by @agilerl-hub-sync[bot] in #683
- docs: keep GitHub PRs open until hub-sync-import completes them by @agilerl-hub-sync[bot] in #686
- docs: keep GitHub PRs open and apply review comments there by @agilerl-hub-sync[bot] in #687
- Chore(deps): Bump google-cloud-storage from 3.10.1 to 3.13.1 by @dependabot[bot] in #673
- ci: load the release-index role ARN from a GitHub secret by @agilerl-hub-sync[bot] in #692
- chore: run check-extras through uv run by @agilerl-hub-sync[bot] in #693
Full Changelog: v2.9.1...v2.10.0
v2.9.1: Arena chat sessions & PAT auth, MoE LoRA under ZeRO-3, GRPO deadlock fix¶
Released on 2026-08-12 - GitHub - PyPI
Features
- Arena chat sessions: LLM deployments now keep conversation state server-side.
Agent.list_sessions/
get_session/delete_session, an optionalsession_idongenerate()andgenerate_stream(), and a new
arena agent sessions list/get/resume/clear/deletecommand group. The CLI keeps one conversation going per
deployment, with--new-sessionto start fresh and--session-idfor a one-off;resumewith no id opens an
arrow-key picker built on termios and msvcrt, so it adds no dependency (#661). - PAT-authenticated inference:
Agenttakes a single credential, sent asAuthorization: Beareron every
route and falling back toARENA_API_KEY. It is never logged, shown inrepr(), or written to disk (#661). memory_scope:deploy_agentandarena agent deploytake an optional--memory-scope user|organization, omitted from the request unless set so a redeploy keeps the stored scope (#661).
Fixes
- Expert LoRA attaches under ZeRO-3 without all-gathering the packed experts.
get_peft_modeland
upgrade_moe_param_wrappersread only the targeted parameters' shapes, dtype and device, so partitioned params
now get zero-storageds_shapeviews. The gather OOMed at attach time on large MoEs, around 55 GB of experts per
rank for Nemotron-3.5-Lightning-30B on 80 GB A100s (#658). - ZeRO-3 persistent params stay resident while the fetch trace is incomplete. deepspeed 0.19.3 honours
ds_persistinrelease_sub_moduleonly once the trace completes, and leaf-module models never complete one, so
every sub-threshold parameter was re-gathered on each use despite being marked persistent (#658). - Multi-process GRPO no longer deadlocks on advantage filtering. Per-rank sample dropping desynchronized the
data-parallel collective schedule, parking a fully filtered rank at a barrier until the NCCL watchdog fired.
Multi-process runs now keep the full batch and zero the advantages of filtered samples; single-process runs keep
the drop and its compute saving (#660). - A stale deployment URL recovers instead of failing with a bare 404. Redeploying moves a deployment, so
open_inference_agentcatches the 404 at bind time, refetches the binding once, and retries, repairing the cache
for later commands. Narrow by design: only a 404, only when the cached URL was used, and only when the refetched
URL differs (#661). - Family-detection tests use the released
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16id in place of the dev
preview name. The Hub redirects the old id, so nothing was broken;detect_model_familiesmatches on the
nemotronsubstring and dispatch is unchanged (#663).
Breaking Changes
- Arena inference callers authenticate as themselves with a user PAT. Deployment API keys are gone from the platform, and
the deploymentapi_keyis gone from the binding cache:save_bindingpurges any key an older release left in
~/.arena/inference.json, so one write cleans every stale entry (#661).
Versions
agilerl-arena==0.2.0 released alongside, carrying chat sessions and PAT-based inference auth; the arena extra pins it.
What's Changed
- Attach expert LoRA under ZeRO-3 without gathering the packed experts by @micdoh in #658
- Add PAT-authenticated chat sessions for Arena LLM deployments by @jaimesabalbermudez in #661
- Mask zero-advantage samples instead of dropping them under multi-process GRPO by @micdoh in #660
- Use the released Nemotron 3.5 Lightning model id by @micdoh in #663
- v2.9.1: Arena chat sessions with PAT auth, ZeRO-3 MoE LoRA attach fix by @jaimesabalbermudez in #662
Full Changelog: v2.9.0...v2.9.1
v2.9.0: ZeRO-3 training + Nemotron support + Multi-Frequency Population-Based Training 🔱🤖🧬¶
Released on 2026-08-10 - GitHub - PyPI
Features
- DeepSpeed ZeRO-3 training: LLM post-training now runs under ZeRO-3, enabling LoRA fine-tuning of ~30B-parameter models with LoRA-scoped gathers, fused/Liger loss support, and robust checkpoint resume (#621).
- Nemotron support: new
agilerl.architecturespackage with Nemotron-H hybrid Mamba and Liger kernel support (#621). - Expert LoRA for MoE models: LoRA can now target packed 3D expert weights via PEFT
target_parameters, including under ZeRO-3 (#621). - Multi-Frequency Population-Based Training: new
MultiFrequencySelectionoperator (MF-PBT, Doulazmi et al.) as a drop-in alternative to tournament selection, with manifest wiring, docs, and a tutorial (#611). mini_batch_size: explicit optimizer-step sizing for LLM algorithms, plus a new batch-sizing docs page (#621).
Optimizations
- LoRA inputs skip PEFT's fp32 cast under bf16 autocast, cutting gradient-step peak memory by ~20–24% with bitwise-identical gradients (#623).
- Fused classic-RL hot paths (Polyak updates, CUDA Adam, batched double-Q) and a vectorized DQN curriculum tutorial: lesson-2 wall time 12.75 h → 2.48 h (#620).
- Colocated rollouts pay offload synchronization once per engine wake instead of once per turn (#638).
Fixes
- Manifest plumbing:
chunk_rows,micro_batch_size_per_gpu,max_wall_seconds, and wandb run naming all flow through correctly, and unrecognized keys are reported with their dotted paths (#638). - Fixed flex-decoding
NoValidChoicesErroron SM90+ GPUs (#621). VLLMConfigdocs examples updated — colocated vLLM always serves LoRA (#640).- CI now guards that the
allextra matches the union of every other extra (#641). agilerl-arena0.1.3 released alongside, carrying theselection_strategymanifest support; thearenaextra pins it (#654).
Breaking Changes
- Unrecognized manifest keys now raise a
ValueErrorat startup; usechunk_rowsin place of the formerfused_*_chunk_rowsfields (#638). tournament_selection_and_mutationis deprecated in favour ofrun_selection_and_mutation(#611).
What's Changed
- perf: fuse algorithm hot paths and vectorize the DQN curriculum tutorial by @micdoh in #620
- Stop paying for PEFT's fp32 LoRA input cast under autocast by @micdoh in #623
- Multiple-Frequencies Population-Based Training (MF-PBT) by @sgarcia56 in #611
- Chore(deps): Bump hydra-core from 1.3.2 to 1.3.4 by @dependabot in #625
- Chore(deps): Bump datasets from 5.0.0 to 5.0.1 by @dependabot in #627
- Fix manifest plumbing bugs and colocated rollout overhead by @micdoh in #638
- Chore(deps): Bump deepspeed from 0.19.2 to 0.19.3 by @dependabot in #629
- [pre-commit.ci] pre-commit autoupdate by @pre-commit-ci in #634
- Chore(deps-dev): Bump pytest from 9.0.3 to 9.1.1 by @dependabot in #628
- Drop dead enable_lora kwarg from VLLMConfig docs by @micdoh in #640
- Guard the all extra alongside the arena dep pin by @micdoh in #641
- Chore(deps): Bump pettingzoo from 1.24.4 to 1.26.1 by @dependabot in #626
- Chore(deps): Bump redis from 8.0.1 to 8.1.0 by @dependabot in #644
- Chore(deps): Bump sphinx-toolbox from 4.2.0 to 4.3.0 by @dependabot in #647
- Chore(deps): Bump python-keycloak from 5.12.0 to 7.1.1 by @dependabot in #645
- Chore(deps): Bump click from 8.4.1 to 8.4.2 by @dependabot in #646
- Chore(deps): Bump pymunk from 7.2.0 to 7.3.0 by @dependabot in #648
- Scope ZeRO-3 gathers to LoRA params and fix adapter copy writes by @mikepratt1 in #621
- Update agilerl version to 2.8.5 by @micdoh in #653
- Update agilerl to 2.9.0 and agilerl-arena to 0.1.3 by @micdoh in #654
New Contributors
- @sgarcia56 made their first contribution in #611
Full Changelog: v2.8.4...v2.9.0
v2.8.4: Housekeeping + ensure alignment between agilerl and agilerl-arena versions¶
Released on 2026-07-29 - GitHub - PyPI
What's Changed
- chore: use
...for empty stub bodies and drop stub-body coverage tests by @micdoh in #619 - 2.8.4: Housekeeping + ensure alignment between agilerl and agilerl-arena versions by @nicku-a in #622
Full Changelog: v2.8.3...v2.8.4
v2.8.3: Static type checking, dependency updates, copyright headers¶
Released on 2026-07-28 - GitHub - PyPI
What's Changed
- Chore(deps): Bump the uv group across 1 directory with 9 updates by @dependabot[bot] in #615
- Fix LLM PPO/REINFORCE silently training on CPU by @micdoh in #617
- test: run the test session in a temporary working directory by @micdoh in #618
- 2.8.3: Static type checking, dependency updates, copyright headers by @nicku-a in #614
Full Changelog: v2.8.2...v2.8.3
v2.8.2: Arena manifest validation simplification¶
Released on 2026-07-22 - GitHub - PyPI
Arena manifest Pydantic validation simplification (#597)
Arena resolves the network encoder architecture server-side, but TrainingManifest._process_manifest was resolving it client-side and assigning the result to algorithm.net_config from arch (which we don't expect users to assign and the server rejects). All the arch normalisation existed only to support that.
- net_config is never set. With exclude_none=True it no longer reaches the server at all.
- _resolve_network is now a raw passthrough. Only
FinetuningNetworkSpecis still materialised client-side, since its fields feed the algorithm section for LLM finetuning. - Deleted _normalize_network_arch, _network_has_arch, _normalize_network_for_platform, _ARCH_TO_NAME, _MANIFEST_ENCODER_ARCHS, and the arch re-injection in _ensure_platform_run_spec_keys.
- The simba/recurrent guard now reads the top-level simba flag against recurrent on the algorithm spec, with no arch inspection.
Fused LoRA routing rewrite (#585, #598)
The fused multi-adapter path drove PEFT's _mixed_batch_forward through a forward pre-hook that injected adapter_names, plus a second hook cloning each base output so PEFT's in-place accumulation never touched a bitsandbytes view. fused_lora.py now owns the layer forward instead.
- Routing is always contiguous runs (["actor"] * B + ["critic"] * B), so the replacement forward walks same-adapter runs with itertools.groupby, slices rows with narrow(), and adds each delta out of place. Out-of-place is what makes it work unchanged on a quantized base, so the clone hook is gone.
- Validation moved to set-time and got stronger (merged-adapter, DoRA, and routing-length checks that PEFT's _check_forward_args used to do). Embedding adapters and aLoRA still delegate to PEFT.
- clear_fused_adapter_routing renamed to unset_fused_adapter_routing; the rest of the public API is unchanged.
- #598 fixes an fp16-checkpoint dtype mismatch: transformers 5.x honours the checkpoint's config dtype, so under bf16 autocast the final norm promotes hidden states to fp32 while the patched lm_head weight stays fp16, crashing the fused logprob matmul. Both operands are now promoted via torch.promote_types before the matmul.
Changes
- dad92d8 Bump supersuit from 3.10.0 to 3.11.0 (#594)
- f286102 Bump dill from 0.4.0 to 0.4.1 (#596)
- 0c6ce74 [pre-commit.ci] pre-commit autoupdate (#579)
- cd0477f Arena manifest Pydantic validation simplification (#597)
- 803b0f2 CI: cap Windows CPU-ISA dispatch at AVX2 to fix intermittent 0xc000001d worker crashes (#599)
- 3f0e4b0 Bump rich from 13.9.4 to 15.0.0 (#595)
- c82594b Fix fp16-checkpoint dtype mismatch in fused lm_head matmuls (#598)
- 639ff48 refactor(llm): compute fused LoRA routing with sliced views, drop PEFT-hook path (#585)
- 98dd736 Bump tqdm from 4.68.2 to 4.68.4 (#593)
- fd0d302 Bump wandb from 0.28.0 to 0.28.1 (#592)
Full Changelog: v2.8.1...v2.8.2
v2.8.1: Fix GRPO crash on first learn() after eval¶
Released on 2026-07-21 - GitHub - PyPI
What's Changed
- Fix GRPO crash on first learn() after eval: restore env batch state on eval_mode exit by @micdoh in #590
- v2.8.1: Fix GRPO post-eval rewards mismatch by @micdoh in #591
Full Changelog: v2.8.0...v2.8.1
v2.8.0: Arena Client & CLI, Trainers, Metrics Observability & More!¶
Released on 2026-07-15 - GitHub - PyPI
Features
Arena Client & CLI (#524, #576): agilerl.arena is the SDK for Arena, the RLOps platform from AgileRL. The goal is that users can do anything they can do in Arena directly from the IDE. ArenaClient provides:
- OAuth2 device-flow authentication with Arena through KeyCloak.
- Upload and validate custom environments and datasets, estimate their resource requirements (profiling), list available environments, and more.
- Project management, training-job submission through a training manifest, and metrics download.
- Agent deployment and inference requests, including streamed LLM completions.
The new agilerl-arena package ships the arena command for driving all of the above from the terminal. Install it directly, or through the AgileRL extra: pip install agilerl[arena].
Trainers (#524): the agilerl.training.trainer module makes it easier to define and iterate on arbitrarily complex RL pipelines, so you can move between local training for rapid development and remote clusters for heavy workloads.
Trainer: base class defining the API for all trainers. Training jobs are declared through Pydantic models representing the underlying training objects (algorithm, buffer, mutations, etc.), submitted viatrain(), and can be built from a dict / YAML / JSON withfrom_manifest().LocalTrainer: initializes training components through their Pydantic models, minimizing overhead when training locally.ArenaTrainer: sets up the same configuration and submits the job to Arena through anArenaClientinstance or an API key.
Agent metrics (#524): agilerl.metrics adds AgentMetrics and MultiAgentMetrics, initialized in all algorithms to abstract metrics logging away from the training loops and simplify them considerably.
Population wrapper (#524): agilerl.population implements Population, a wrapper around a list of individuals training simultaneously that aggregates population-level metrics and provides methods that rely on population-level information.
Flexible logging tools (#524): the agilerl.logger suite extracts gathered metrics in specific ways — StdOutLogger (Rich table), CSVLogger, WandbLogger, and TensorboardLogger (via torch.utils.tensorboard.SummaryWriter).
LLM chunking unified under chunk_rows (#565): LLMAlgorithm had two chunk-size knobs (FUSED_LOGPROBS_CHUNK_ROWS and FUSED_LOSS_CHUNK_ROWS) that were always set to the same value. They are collapsed into a single chunk_rows arg / CHUNK_ROWS INIT_HP key bounding the per-chunk logit workspace for both the standard fused-logprob and Liger fused-loss paths.
Colocated vLLM LoRA sync hardening (#565): new VLLMConfig.sleep_mode_level (1 or 2) passed through to llm.sleep(level=...), with the colocated engine sleeping and waking on every rank rather than only the main process; per-rank adapter staging via lora_staging_per_rank so each distributed rank writes and loads its own adapter; and a CUDA device guard on the colocated add_lora call.
Docs & tutorials (#524, #562, #576): new GRPO-on-GSM8K fine-tuning tutorial via Arena with example manifest and reward file; expanded PPO custom-env tutorial with a validation fail-then-fix walkthrough; new sections for Trainer, the Arena client, and metrics/logging; multi-turn LLM benchmark charts on the README and docs landing page; and sphinx-copybutton for copyable code snippets.
Also in this release: PPO action masking during policy evaluation, multi-agent TensorDict buffers, swap_channels moved inside algorithms with ImageTranspose, NetworkSpec resolution fixes, and EvolvableAlgorithm.population() made robust to LLM algorithms.
Breaking Changes
- Standardised common arguments for all training functions (
INIT_HP→init_hp,MUT_P→mut_p). (#524) MultiAgentReplayBufferhas been removed; the single-agentReplayBuffernow supports multi-agent transitions transparently. (#524)PPOno longer learns from an experiences tuple. It uses a rollout buffer stored on the algorithm;PPO.learn()takes no required arguments and optionally accepts a pre-collected rollout batch. (#524, #587)- Removed the
swap_channelsargument from all training loops - now handled under the hood in the baseEvolvableAlgorithm. (#524) - Removed
eval_loopfromTournamentSelection, since the average fitness across evaluation episodes is appended and only the last element is needed. (#524) - Removed the unused/redundant
perandn_steparguments fromtrain_off_policy. (#524) - The old LLM chunking names are hard-removed: passing
FUSED_LOGPROBS_CHUNK_ROWSorFUSED_LOSS_CHUNK_ROWSraises a clear error pointing tochunk_rows. (#565) pettingzoois now pinned to>=1.23.1,<1.25: the MPE environments moved out of PettingZoo into the separatempe2package as of 1.25.create_population()is deprecated in favour ofEvolvableAlgorithm.population(), which the documentation now uses throughout. (#524)
Bugs
- CISPO / Liger multi-GPU NCCL deadlocks (#586): distributed runs hung after the first learn/metrics step because ranks issued different collective sequences. Fixes three desync sources: cross-rank completion-length mismatch before
learn()(ranks now pad to the global max sequence length for Liger token-level importance sampling), main-process-onlyreport_metrics()(all ranks now report;StdOutLoggerprints only on main), and uneven multi-turn rollout loop lengths (ranks stay in lockstep, idle ranks run a dummy generation turn). - Multi-agent
RSNorm(#562): per-agent observations were routed through the wrongrmsshape; multi-agent paths now delegate per agent instead of inlining the normalization math. build_rms(#562): crashed whennorm_obs_keysfiltered aDictspace; dict spaces are now filtered viaspaces_mapwithout treating a plaindictas agymnasium.spaces.Dict.DummyEvolvable.to_evolvable()(#562): passed positional args in the wrong order; now constructs by keyword to match__init__.MATD3(#562): removed duplicated unreachable critic-set validation, aligning thecritics_listcheck withMADDPG.- Offline training loop (#524): was not using the TensorDict replay buffer.
- Bandit training loop (#524): context was not indexed by action correctly.
_prepare_vllm_for_training(#565): theuse_vllm=Falselearn path no longer dereferences aNonevllm_config.
What's Changed
- Raise unit test coverage, fix RSNorm, DummyEvolvable, and MATD3 validation, and add LLM benchmark graphs by @nicku-a in #562
- Enable Ruff linting on tests and fix violations by @nicku-a in #564
- Raise unit test coverage, fix RSNorm, DummyEvolvable, and MATD3 validation, and add LLM benchmark graphs by @nicku-a in #563
- ci: run test matrix on uv.lock changes by @micdoh in #581
- Bump accelerate from 1.13.0 to 1.14.0 by @dependabot[bot] in #538
- Bump deepspeed from 0.19.1 to 0.19.2 by @dependabot[bot] in #550
- Bump wandb from 0.27.0 to 0.28.0 by @dependabot[bot] in #566
- ci: drop container: for ops GPU runner image by @dougalrea in #583
- fix: resolve code-quality findings from PR #578 by @jaimesabalbermudez in #580
- fix: align learn/train/metrics signatures with base classes by @micdoh in #587
- Bugfix/cispo norm cross rank hang by @mikepratt1 in #586
- Bump redis from 8.0.0 to 8.0.1 by @dependabot[bot] in #567
- v2.8.0: Arena Client & CLI, Trainers, Metrics Observability & More by @jaimesabalbermudez in #578
Full Changelog: v2.7.1...v2.8.0