NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2487 most downloaded on PyPI
Train transformer language models with reinforcement learning.
Last release 5 days ago
29 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
103 releases · first in 2020
One column per quarter.
Deprecate & remove lmbda on/off-policy mixing and the off-policy training branch ( #6458 , #6460 ); GKD paper reproduction now points at GKDTrainer
Long-standing request finally landed. GRPO and RLOO rely on RepeatSampler to repeat each prompt num_generations times and group them across processes — but samplers can't attach to iterable datasets (no length, no indexing), and the old behavior silently bypassed the sampler, quietly corrupting advantage computation.
A new repeat_iterable_dataset generator reproduces RepeatSampler's exact ordering by transforming the stream itself instead of reordering indices. Trainers wrap iterable datasets via IterableDataset.from_generator, force dispatch_batches=False, and shard the stream per-process to preserve prompt groups after cross-process gathering.
from datasets import load_dataset
from trl import GRPOConfig, GRPOTrainer
dataset = load_dataset("trl-lib/DeepMath-103K", split="train", streaming=True)
trainer = GRPOTrainer(
model="Qwen/Qwen3-4B",
args=GRPOConfig(max_steps=1000), # required for iterable datasets
train_dataset=dataset,
reward_funcs=accuracy_reward,
)max_steps is required for iterable datasets (no length). A clear error also fires if dispatch_batches=True is combined with an iterable dataset.
by @albertvillanova in #6351
When an environment_factory is provided, the environment can now own the data. train_dataset becomes optional: reset() returns the prompt. No more fabricating a dummy dataset just to drive the loop.
class WordleEnv:
def reset(self, **kwargs) -> str: # returns the prompt
self._target = sample(words)
return f"Guess a {len(self._target)}-letter word."
def get_reward(self) -> float:
return 1.0 if self._solved else 0.0
def guess(self, word: str) -> str: # exposed tool
self._solved = word == self._target
return _feedback(word, self._target)
trainer = GRPOTrainer(
model="Qwen/Qwen3-4B",
args=GRPOConfig(max_steps=1000),
environment_factory=WordleEnv, # no train_dataset needed
)Applies to both GRPOTrainer and AsyncGRPOTrainer. A provided train_dataset still works exactly as before; multi-environment routing datasets can now be environment-only (no prompt column required).
by @qgallouedec in #6349
An opt-in way to build training rows from a multi-turn conversation. Message mode keeps the conversation as messages and re-tokenizes the whole thing each turn, then checks whether the fresh tokens still start with the tokens held so far. If yes → append the new part (same as token mode). If no → a rewrite happened → close the row and open a new one that matches what the model actually read.
AsyncGRPOConfig(
rollout_protocol="message", # "token" (default) | "message"
fork_threshold_tokens=1024, # message-mode only
)Per-turn drift is classified as CLEAN (append), REALIGN (last-answer tail wobble → overwrite same row), or FORK (real rewrite → new row). One advantage per conversation, stamped on every row it produced; under token-mean loss a fork is invisible. Sets up the plumbing for future tree-trajectory support — scoring is already branch-agnostic (groups by rollout_id).
by @AmineDiro in #6250
VLLMClient and decoupled weight syncA follow-up to v1.8's rollout worker split. A small VLLMClient becomes the single place that speaks HTTP to the vLLM server (named methods for wait_for_server_ready, get_max_model_len, pause/resume, weight-update endpoints — stateless, so picklable). Generation deliberately stays on its own async aiohttp session in the rollout child.
Weight sync is now built from config independent of rollout_worker (previously a custom rollout worker force-set it to None), and is injectable via a new WeightTransferProtocol. token_budget now waits for /health before defaulting to get_max_model_len(), so a slow-starting vLLM waits instead of failing the run.
by @AmineDiro in #6269
DistillationTrainer refactorA ~13-PR sweep that reshapes the experimental DistillationTrainer for future stability:
ServerDistillationTrainer for the teacher-server path (#6454)lmbda on/off-policy mixing and the off-policy training branch (#6458, #6460); GKD paper reproduction now points at GKDTrainer (#6459)prompt column alongside messages (#6461); deprecate messages-format datasets (#6474)num_items_in_batch to count generated completion tokens — the loss denominator was wrong for on-policy training (counted dataset completions, not generated ones) and NaN'd on prompt-only datasets. This changes loss values for anyone on the old path (#6478)All by @qgallouedec.
Importance-Weighted On-Policy Distillation lands as an optional objective on the experimental DistillationTrainer — detached IW-OPD advantages built from sampled-token teacher and rollout logprobs (with cached vLLM rollout logprobs when use_vllm=True).
DistillationConfig(
distillation_objective="iw_opd",
iw_opd_gamma=1.0,
iw_opd_epsilon=0.2,
lmbda=1.0, reverse_kl_top_1_mode="sampled", # required
)Now split into its own experimental trainer as part of the refactor above.
log_multimodal for GRPO / RLOONew config toggle to control whether images from multimodal completions are logged to trackers. Also applied to experimental DPPO and GRPO-with-replay-buffer.
by @apardyl in #5408 and by @albertvillanova in #6352
generate_rollout_completions by @sergiopaniego in #5870top_p in GKD on-policy generation by @sergiopaniego in #6337DatasetDict and IterableDatasetDict as eval_dataset in evaluate by @albertvillanova in #6326AsyncGRPOTrainer args handling with stable trainers by @qgallouedec in #6377auto_find_batch_size in GRPO and RLOO by @qgallouedec in #6411GMPOTrainer to the telemetry allowlist by @qgallouedec in #6410[GRPO] Treat explicit max_tool_calling_iterations=0 as zero, not unlimited by @Strongich in #6423hf CLI by @davanstrien in #6422messages-format datasets in DistillationTrainer — use prompt column instead — by @qgallouedec in #6474lmbda on/off-policy mixing in DistillationTrainer — off-policy path fully removed by @qgallouedec in #6458 and #6460clip_ratio/low_min / high_max to be per-completion instead of per-rank-batch-mean (DP-layout-independent). Logging-only, no loss change. By @qgallouedec in #6380use_liger_kernel under DeepSpeed ZeRO-3 by @qgallouedec in #6372use_reentrant=True for PEFT + ZeRO-3 + gradient checkpointing across all trainers by @qgallouedec in #6356precompute_ref_log_probs=True with DeepSpeed in DPO and KTO by @qgallouedec in #6403cast_lm_head_to_fp32 under FSDP2 / DeepSpeed ZeRO-3 by @qgallouedec in #6330</think>-only check. By @qgallouedec in #6363pad_token_id == eos_token_id in GKD and GOLD by @roycho96 in #6357apply_chat_template dropping leading tokens from chosen when the prompt common prefix shrinks against rejected by @sohumt123 in #6433functools.partial or callable instance by @qgallouedec in #6313prepare_peft_model by @sergiopaniego in #6343FileNotFoundError in DPO/KTO precompute_ref_log_probs by @Functionhx in #6348GeometricMixtureWrapper.generate() crash after transformers' MoE decode optimization change by @albertvillanova in #6413vocab_size check for multimodal model configs by @sergiopaniego in #6340IterableDataset in GRPO and RLOO by @albertvillanova in #6324eval_dataset and precompute_ref_log_probs by @albertvillanova in #6370eval_dataset dict passed to evaluate() by @albertvillanova in #6434hasattr guards on lm_head.bias in the Liger path by @qgallouedec in #6425logits/chosen and logits/rejected metrics by @qgallouedec in #6424eos_token_id in non-vLLM on-policy GOLD test stubs by @kashif in #6394BEMACallback bias_power docstring by @DaoyuanLi2816 in #6390max_steps requirement for iterable datasets in GRPO and RLOO by @albertvillanova in #6432AsyncGRPOConfig.max_completion_length description with GRPOConfig by @qgallouedec in #6374AsyncGRPOTrainer.processing_class docstring with GRPOTrainer by @qgallouedec in #6375AsyncGRPOTrainer clip-ratio comment with GRPOTrainer by @qgallouedec in #6376loss_type by @qgallouedec in #6373BEMACallback bias_power=0.0 test by @albertvillanova in #6464eval_dataset input types in core trainer init / evaluate() tests by @albertvillanova in #6336 and #6371eval_dataset init tests for experimental GRPO / Online-DPO / SFT & distillation / preference / BCO+CPO+ORPO+PRM trainers by @albertvillanova in #6395, #6396, #6397, #6398 and #6399environment_factory warning in GRPO tests by @albertvillanova in #6445ensure_weight_tying warning on newer PEFT in GRPO tests / DPO+KTO liger tests by @albertvillanova in #6446 and #6463bitsandbytes _check_is_size FutureWarning in CI by @albertvillanova in #6448IndexPutBackward0 warning from Liger SAPO loss by @albertvillanova in #6436torch.jit.script_method warning on Python 3.14 by @albertvillanova in #6465AnnAssign DeprecationWarning on Python 3.14 by @albertvillanova in #6467create_reference_model deprecation warning by @albertvillanova in #6418dpo invariant reference after #6321 by @qgallouedec in #6421torch.tensor re-wrapping of batch labels in KTO by @albertvillanova in #6431log_multimodal param to GRPOConfig and RLOOConfig to control image logging by @apardyl in #5408Note truncated.
Fix deprecation in CI Sync TRL skills by @albertvillanova in #6200
After many cycles of KTOTrainer ↔ DPOTrainer alignment work, KTO graduates from trl.experimental.kto to the top-level trl package. Same API as DPO/GRPO/SFT — imports move from experimental, tests move to the main test tree, docs no longer flag it as experimental. The experimental path still works and emits a FutureWarning (removal in v2.0.0).
# Before
from trl.experimental.kto import KTOConfig, KTOTrainer
# Now
from trl import KTOConfig, KTOTrainerPer our telemetry, KTO is the 4th most used trainer in TRL — this graduation was overdue.
by @albertvillanova in #6175, #6287 and #6345
Three interrelated changes make agentic RL training with environments substantially more ergonomic.
Environment-owned reward. If your environment_factory env defines a reserved get_reward() method (no args → float), it's called once per completed rollout and treated as a reward source. reward_funcs becomes optional — no more leaking env state back out to trainer-owned reward funcs.
class WordleEnv:
def reset(self, **kwargs):
self._target = sample(words); self._solved = False
def get_reward(self) -> float: # optional, reserved (not a tool)
return 1.0 if self._solved else 0.0
def guess(self, word: str) -> str: # exposed as a tool
self._solved = word == self._target; ...
trainer = GRPOTrainer(
model=model,
train_dataset=dataset,
environment_factory=WordleEnv, # no reward_funcs needed
)Multi-environment support. environment_factory now accepts dict[str, factory] in addition to a single callable. Each dataset row selects its environment via an environment column, and only that env's tools are exposed in that row's prompt — so a coding task and a game can train together in one run without leaking each other's tool schemas. Single-callable usage is unchanged.
Same wiring lands in GRPO, AsyncGRPO, DPPO, and GRPO-with-replay-buffer.
Env-owned reward by @qgallouedec in #6238; multi-env in #6001 and #6002
GRPOConfig now supports both static and adaptive entropy regularization (Skywork-OR1). The bonus encourages exploration and helps prevent premature policy collapse.
GRPOConfig(
entropy_coef=0.01, # static
# or:
use_adaptive_entropy=True, # adjust to target entropy
entropy_target=1.5,
entropy_coef_delta=0.01,
entropy_coef_min=0.0,
entropy_coef_max=0.1,
)Adaptive mode adjusts the coefficient once per optimizer step from window-aggregated entropy (gradient accumulation-aware) and persists entropy_coef in the checkpoint for resume. Not compatible with the Liger kernel.
by @albertvillanova in #6140
quantization_config trainer argument (streamlined QLoRA)QLoRA no longer requires reaching into model_init_kwargs or pre-loading the model manually.
SFTTrainer(
model="meta-llama/Llama-2-7b-hf",
quantization_config=BitsAndBytesConfig(load_in_4bit=True),
peft_config=LoraConfig(),
train_dataset=dataset,
)Added to SFTTrainer, DPOTrainer, GRPOTrainer, RLOOTrainer, RewardTrainer, and KTOTrainer. Sits next to peft_config (the other non-serializable QLoRA ingredient), flows into from_pretrained, and raises if also set in args.model_init_kwargs. Also drops the redundant get_kbit_device_map() line — QLoRA trains identically without it across all tested configurations.
by @qgallouedec in #6157 and #6276
v1.7 added the router load-balancing auxiliary loss to GRPOTrainer / RLOOTrainer / AsyncGRPOTrainer. It's now on DPOTrainer and KTOTrainer too, so post-training MoE models with preference data keeps experts balanced.
by @qgallouedec in #6208 and #6275
chunked_nll via static-shape token packingThe chunked NLL path had data-dependent indexing that broke XLA/Neuron compilation. This PR reworks it to pack valid tokens to the front via a stable argsort on the validity mask, iterate over ceil(n_valid / chunk_size) * chunk_size whole chunks (a tensor, no Python-int sync), and use ignore_index=-100 inside each chunk. GPU behavior is unchanged; Neuron now works too — same tiny memory footprint (~1.33 GiB vs the naive 5.94 GiB alternative) with no per-mask recompilation.
by @michaelbenayoun in #6314
AsyncGRPO micro-batching becomes packing-aware and token-bounded on top of the padding-free path from v1.7:
token_budget tokens with a variable sample count, decoupling peak memory from per_device_train_batch_size. Useful in memory-bound regimes (no gradient checkpointing, long context, very large models).Both ride HF Trainer's existing gradient accumulation — no training-loop surgery, FSDP/EP collectives stay in lockstep.
by @AmineDiro in #6092
GOLDTrainerGOLDTrainer now supports vision-language models end-to-end.
by @Strongich in #5969
by @qgallouedec in #6259
by @albertvillanova in #6159 and #6277
fraction in dataset mixturesYou can now weight a DatasetMixtureConfig by fraction instead of only by explicit counts.
by @qgallouedec in #6199
Continuing the label-refactor from v1.7 (#6037), truncation moves out of the collator and into dataset preparation. SFTTrainer and DPOTrainer now truncate consistently at the same phase.
by @qgallouedec in #6155
A big refactor pass merging DPO / SFT / Reward / KTO tokenization into one shared implementation:
is_vlm parameter — #6298_tokenize a module-level function — #6301_tokenize into a single shared function — #6302apply_chat_template kwargs — #6305chat_template as apply_chat_template_kwargs — #6315All by @albertvillanova. Related: align collators across DPO / SFT / Reward / KTO by @qgallouedec in #6178.
DatasetDict and IterableDatasetDict as eval_dataset in trainers by @albertvillanova in #6322processing_class in CPO/ORPO trainers when omitted by @DaoyuanLi2816 in #6087trust_remote_code to GRPOConfig by @muupan in #4186[async GRPO] Per-generation reset() observation by @qgallouedec in #6072token_budget to vLLM max_model_len by @AmineDiro in #6218entropy and num_tokens metrics in KTO by @qgallouedec in #6257 and #6256TrainerCallback import to the public top-level path by @qgallouedec in #6249get_kbit_device_map() by @qgallouedec in #6158get_dataset_column_names from core and experimental trainers by @albertvillanova in #6272 and #6273GFPOTrainer by @qgallouedec in #6309 — no known usage, and its behavior can be reproduced by adjusting reward_funcs.PAPOTrainer by @qgallouedec in #6235sft_video_llm.py script by @qgallouedec in #6193chunked_nll patch hiding VLM kwargs from generate by @Strongich in #6156OnlineDPOTrainer by @qgallouedec in #6228precompute_ref_log_probs guard for iterable datasets nested in a dict by @albertvillanova in #6307quantization_config + already-instantiated model in DPOTrainer / KTOTrainer by @qgallouedec in #6312KTOTrainer._prepare_dataset type hint by @qgallouedec in #6207ensure_weight_tying warning in liger + PEFT GRPO tests by @albertvillanova in #6188add_response_schema tests for the new parse_response prefix requirement by @qgallouedec in #6236num_key_value_heads = num_attention_heads by @albertvillanova in #6268ValueError instead of NotImplementedError for Liger DPO with PEFT by @albertvillanova in #6225max_steps is required for iterable train datasets by @albertvillanova in #6333epsilon help/docstring wording by @qgallouedec in #6014 (v1.7 window)max_length docstring (truncation direction) by @qgallouedec in #6254compute_metrics type hint (EvalLoopOutput → EvalPrediction) by @qgallouedec in #6253huggingface/doc-builder by @albertvillanova in #6304doc-builder pin by @mishig25 in #6317test_chatml_collator_truncates_keeping_completion_end by @albertvillanova in #6318compute_metrics and Liger by @albertvillanova in #6231 and #6232TestKTOTrainerSlow with test_train_vlm_gemma_3n by @albertvillanova in #6162DataCollatorForChatML raises when completion fills window by @albertvillanova in #6319_get_kl_completion_ids into _get_kl_dataset by @qgallouedec in #6181processing_class in CPO/ORPO trainers when omitted by @DaoyuanLi2816 in #6087Note truncated.
Fix GRPO + vLLM colocate + PEFT hang on non-NVLink hardware by @albertvillanova in #6139
add_response_schema tests for the new parse_response prefix requirement by @qgallouedec in #6236Full Changelog: v1.7.0...v1.7.1
use_transformers_paged was deprecated in v1.4; it's now replaced with proper transformers continuous batching . The old branch silently bypassed impor…
loss_type is now "chunked_nll"The flip announced in v1.6 has landed. Setting loss_type is optional, and the default now resolves to "chunked_nll" — giving every SFTTrainer run ~30% less peak VRAM on average (up to ~50% on large-vocab models) with wall-clock time neutral or slightly faster. No action needed.
The auto-resolve falls back to "nll" when use_liger_kernel=True (the two paths are incompatible). If you want the old behavior — e.g. for custom heads — pin it explicitly:
SFTConfig(loss_type="nll")by @qgallouedec in #5846
Post-training MoE models now correctly include the router load-balancing auxiliary loss, matching the model's own reference forward and SFTTrainer. Enable via model_init_kwargs:
GRPOConfig(
...,
model_init_kwargs={"output_router_logits": True, "router_aux_loss_coef": 0.001},
)Plumbed through _get_per_token_logps_and_entropies (now returns a 3-tuple including aux_loss), folded into the policy loss with grad-accum scaling matched per trainer, and logged as aux_loss. AsyncGRPO recomputes it via load_balancing_loss_func in the chunked LM-head path (same as SFT's chunked path).
by @AmineDiro in #6083, plus router_aux_loss_coef config wiring by @qgallouedec in #6085
Geometric-Mean Policy Optimization lands as an experimental trainer. Replaces GRPO's per-token arithmetic mean of importance ratios with a sequence-level geometric mean (mean of clipped log-ratios, then exp); clipping is one-sided by advantage sign and applied in log space. Default epsilon=0.4 per the paper.
from trl.experimental.gmpo import GMPOConfig, GMPOTrainer
trainer = GMPOTrainer(
model="Qwen/Qwen3-4B",
args=GMPOConfig(epsilon=0.4),
reward_funcs=accuracy_reward,
train_dataset=dataset,
)by @raghulchandramouli in #6078
use_transformers_paged was deprecated in v1.4; it's now replaced with proper transformers continuous batching. The old branch silently bypassed importance-sampling correction (logprobs = None); the new path captures logprobs from output.logprobs and exposes a ContinuousBatchingConfig for KV-cache tuning.
GRPOConfig(
...,
use_transformers_continuous_batching=True,
transformers_continuous_batching_config={
"use_cuda_graph": False,
"max_memory_percent": 0.4, # leave headroom for training
},
)Benchmark (Llama-3.2-1B-Instruct, A100 80GB, GSM8K): 1.25× faster at N=64 generations with -16 GB peak VRAM vs default generate(). Use when N ≥ 32 with variable completion lengths.
use_transformers_paged=True still works and forwards to the new flag with a FutureWarning. Requires transformers>=5.8.0.
by @sergiopaniego in #5765
WeightTransferClient now drives vLLM's native 4-phase RL weight-transfer API instead of the older 2-call flow: pause(mode="keep") → start_weight_update → threaded update_weights + NCCL broadcast → finish_weight_update → resume. Validated end-to-end on H100 across single-node, FSDP2×4 + TP=4, and 2-node FSDP2×4 + DP=2×TP=4 (weight-sync time ≈ 0.18-0.8 s).
by @AmineDiro in #5892
AsyncGRPO now supports the same padding-free path SFT already had. Flattens the batch and uses position_ids-based document boundaries instead of right-padding to the longest sequence — meaningful speedup and memory savings on heterogeneous-length workloads.
by @qgallouedec in #5854
A new trl.experimental.harbor adapter plugs Harbor agentic task suites into GRPOTrainer via environment_factory. Same pattern as the OpenReward integration — one spec wires all three trainer slots:
from trl import GRPOConfig, GRPOTrainer
from trl.experimental.harbor import HarborSpec
spec = HarborSpec("AdithyaSK/data_agent_rl_environment_train", agent="bash", num_tasks=64)
trainer = GRPOTrainer(
model="Qwen/Qwen3-4B",
args=GRPOConfig(num_generations=8, max_steps=50, max_tool_calling_iterations=25),
train_dataset=spec.train_dataset,
environment_factory=spec.environment_factory,
reward_funcs=spec.reward_funcs,
)Built-in bash harness, plus jupyter and terminal_notes example harnesses. Gated by the new trl[harbor] extra.
by @adithya-s-k in #6018
trust_remote_code in trainer configsA single trust_remote_code: bool = False field on the trainer configs now covers the whole load surface — model, processor / tokenizer, reference model, reward model, reward tokenizer, teacher — instead of forcing users to thread it through several independent kwarg dicts.
SFTConfig(trust_remote_code=True)ModelConfig.trust_remote_code is removed to avoid duplicate --trust_remote_code when combining dataclasses; CLI behavior is unchanged.
by @qgallouedec in #5802
The last alignment cycle before graduation: KTO now has parity with DPO on pad_to_multiple_of, sync_ref_model, evaluate(), method order/signature, metric placement (all moved into _compute_loss), and a real text + VLM (including multi-image) test suite.
PRs all by @albertvillanova: #6029, #6030, #6033, #6034, #6035, #6080, #6093, #6148, #6149, #6150, #6152, #6160, #6163.
GRPO and RLOO now support LFM2-VL multimodal inputs end-to-end.
by @zwischenraum in #6114
get_repetition_penalty_reward by @qgallouedec in #6058get_cosine_scaled_reward by @qgallouedec in #6066Label construction moves out of the collator and into dataset preparation, so "what's trainable" is defined in exactly one place. A single batched map produces a labels column where each token keeps its ID when every applicable mask is 1, else -100. Plain LM stays storage-neutral; pre-tokenized datasets with mask columns now go through the same path. Step 1 toward fixing #3927.
{% generation %}-marker training template for Idefics3, enabling assistant_only_loss=True.
num_items_in_batch for gradient accumulation by @behroozazarkhalili in #6006[AsyncGRPO] Rollout worker: set aiohttp limit to max(100, max_inflight_tasks) by @ggcr in #5861evaluate() accept the same dataset types as the trainer by @qgallouedec in #6116unpair_preference_dataset by @albertvillanova in #6161fix(profiling): log ProfilingContext metrics to Trackio backend by @Anai-Guo in #5979.contiguous() calls by @qgallouedec in #6045 and #6046create_reference_model with num_shared_layers was double-allocating the "shared" frozen layers because the loop never assigned _ref_param back. Now it does, so shared layers are held once. By @behroozazarkhalili in #6053chunked_nll mixed Tensor/DTensor error under FSDP2 + PEFT by @albertvillanova in #6065lm_head.weight all-gathers under FSDP2 + chunked_nll by @albertvillanova in #6077target_parameters by @discobot in #6043use_liger_kernel is combined with a PEFT adapter on lm_head by @akshansh47 in #5977[fix] GLM-4-MoE template: turn-terminating token to the turn itself by @qgallouedec in #6044unpair_preference_dataset dropping extra columns by @albertvillanova in #6059fix(gold): preserve vllm prompt special tokens by @he-yufeng in #6063fix(sft): reject transformed datasets during preparation by @he-yufeng in #6054_unpair_row by @albertvillanova in #6062__init__ signatures by @DaoyuanLi2816 in #6011paper_index (experimental page was split) by @DaoyuanLi2816 in #6070docs(rloo): add clarifying comment on KL penalty formula divergence from GRPO by @abderahmane-ai in #6096vllm>=0.22.0 by @sergiopaniego in #6101epsilon help/docstring wording by @qgallouedec in #6014_push_param_to_vllm helper in VLLMGeneration by @albertvillanova in #6004async_grpo to match the rest of trl/experimental by @qgallouedec in #6012num_completions_to_print, epsilon_high fallback, logging_steps docstring, loss variable names, clip-ratio metrics by @qgallouedec in #6020, #6019, #6016, #6013 and #6021chunked_nll loss by @albertvillanova in #6074test: bound memory in test_gkd_trainer_with_liger to avoid OOM on shared runners by @behroozazarkhalili in #6103sft_fa2 invariant test by @qgallouedec in #6069actions/checkout / astral-sh/setup-uv / actions/setup-python / pre-commit/action by @albertvillanova in #6097, #6098, #6099 and #6100parse_version with Version by @albertvillanova in #6164deepspeed < 0.19.2 by @albertvillanova in #6090chore: update tests_transformers_branch.yml by @hf-security-analysis[bot] in #6051chore: update clear_cache.yml by @hf-security-analysis[bot] in #6047chore: update docker-build.yml by @hf-security-analysis[bot] in #6048pr_style_bot workflow by @albertvillanova in #6082_get_train_sampler comment to mention num_iterations > 1 by @anidoesdev in #6125Note truncated.
The old vllm_importance_sampling_cap is deprecated and maps to clip_max .
AsyncRolloutWorker is no longer a thread — it's a spawned child process with its own GIL. The trainer's autograd engine no longer competes with recursive_parse / accuracy_reward for the GIL, which was causing 1-5s stalls in real Qwen3-30B-A3B @ 16k runs and ultimately NCCL watchdog timeouts on other ranks.
Architectural changes:
AsyncRolloutWorker (parent) owns the child process + shared mp.Queue / mp.Value / mp.Event._AsyncRolloutLoop (child-only) handles tokenization, dataset iteration, reward funcs, and asyncio loops.WeightTransferClient owns the NCCL group with vLLM (/pause, /resume, /init_weight_transfer_engine, /update_weights); the rollout child only talks to /v1/completions.Two correctness fixes shipped alongside (they would have conflicted otherwise): broader aiohttp retry (now catches ClientPayloadError) with bounded exponential backoff, and all-NaN reward columns are now preserved — np.nansum was silently returning 0, giving unscorable completions a real advantage signal and pushing the policy away from correct answers (~30% of DeepMath / OpenR1-Math rows).
Note
reward_funcs / tools / environment_factory must now be picklable, and the child runs CPU-only (CUDA_VISIBLE_DEVICES="").
by @AmineDiro in #5749
A new A2POTrainer implements A*-PO from "Accelerating RL for LLM Reasoning with Optimal Advantage Regression". Two stages: an offline V* estimation pass from reference policy samples (with optional filter_all_incorrect to drop prompts where every reference completion fails), then on-policy training with one generation per prompt and a plain least-squares loss on β₂·log(π/π_ref) vs r − V*. No group, no critic, no clipping, no reward normalization.
from trl.experimental.a2po import A2POConfig, A2POTrainer
trainer = A2POTrainer(
model="Qwen/Qwen3-4B",
args=A2POConfig(num_value_samples=8, filter_all_incorrect=True),
train_dataset=dataset,
reward_funcs=accuracy_reward,
)
trainer.train()Designed for binary verifiable rewards (math/code), not open-ended problems.
by @raghulchandramouli in #5940
The biggest KTO ↔ DPO alignment cycle yet — KTOTrainer now supports vision-language models, plus a deep restructuring of compute_loss, KL dataset generation, ref-logp precomputation, activation offloading, sampler strategy, metrics, and more. KTO graduation is very close.
from trl.experimental.kto import KTOConfig, KTOTrainer
trainer = KTOTrainer(
model="Qwen/Qwen2.5-VL-3B-Instruct",
args=KTOConfig(...),
train_dataset=vision_kto_dataset,
)VLM support: by @albertvillanova in #5939. Plus ~20 alignment PRs all by @albertvillanova: #5820, #5849, #5852, #5850, #5866, #5864, #5856, #5872, #5875, #5900, #5901, #5899, #5906, #5909, #5914, #5982, #5936, #5996, #5998, #5999.
The GOLD distillation trainer used to align student/teacher tokens by extending two decoded strings and flushing on equality. It silently broke on any byte-level disagreement — including the common case of one tokenizer prepending BOS while the other doesn't (Llama-3 ↔ Qwen-3). The X-Token paper called this out by name.
Each side now carries (start_byte, end_byte) spans derived once from the fast tokenizer's char offsets, and the walker syncs on cumulative byte boundaries. On the on-policy path, spans come from piece_byte_len over the sampled token ids (not from re-encoding the decoded completion — BPE makes that round-trip non-injective).
Two related fixes shipped: long rows no longer lose the completion (now keeping the last max_length tokens), and the vLLM on-policy original_prompt_text is now decoded from the truncated ids the student actually consumed.
When teacher_model_kind="live" and vllm_mode="server", the vLLM generation server already holds the current student weights (synced every step for rollouts). The new use_teacher_server=True flag scores the teacher's log-probs on that same server instead of running a separate local teacher forward — removing the teacher from the training step entirely.
Supported modes: sampled_token (reverse KL on the realized token) and topk_logits. When buffered batches reuse steps (num_iterations > 1), weights are re-synced before scoring so the teacher never scores stale.
vLLM importance sampling in GRPO now uses a two-sided band [C_min, C_max] instead of a single upper cap, aligning TIS/MIS with IcePop's bidirectional handling of train–inference ratio outliers.
from trl import GRPOConfig
config = GRPOConfig(
vllm_importance_sampling_clip_min=0.5,
vllm_importance_sampling_clip_max=2.0,
vllm_importance_sampling_correction="mask", # or "truncate"
)The old vllm_importance_sampling_cap is deprecated and maps to clip_max.
Day-zero training support for NVIDIA's new model families.
Three more model families with {% generation %} markers (assistant-only loss out of the box):
A new trl/distributed.py introduces a single DistributedBackend class that detects ZeRO stage and FSDP version once, then exposes two context managers (gather_params, summon_full_params) used everywhere. Replaces the scattered getattr(state, "fsdp_plugin", None) / gather_if_zero3 / summon_full_params if ... else nullcontext() boilerplate spread across vllm_generation.py, models/utils.py, and the main trainers. Future deprecations land in one place.
by @albertvillanova in #6000
A two-PR refactor that disentangles SDPO, SDFT, and other self-distillation trainers from their shared base, making each one self-contained and consistent with the rest of the codebase.
by @LeonEricsson in #5862 and #5883
loss_type will change in 1.7Setting SFTConfig.loss_type is now optional, and leaving it unset emits a FutureWarning: in TRL 1.7 the default will switch from "nll" to "chunked_nll". No action needed — you'll just get the new default automatically on upgrade — unless you want to pin the current behavior (e.g. for custom models) with loss_type="nll".
by @qgallouedec in #5997
'None' as CLI value for Optional[T] fields by @qgallouedec in #5843lm_head output projections in chunked SFT loss (GPTNeoX) by @qgallouedec in #5857SFTTrainer: merge entropy and accuracy computation to eliminate redundant logits copy by @flutist in #5897.contiguous() calls in DPOTrainer to reduce peak memory by @flutist in #5926.contiguous() before entropy_from_logits by @qgallouedec in #5930None reward completions from GRPO/RLOO advantage baseline by @AmineDiro in #5902precompute_ref_log_probs with vision datasets in DPO by @albertvillanova in #5867max_length by @lxk8998 in #5927loss_type="chunked_nll" under DeepSpeed ZeRO-3 by @qgallouedec in #5873use_liger_kernel under DeepSpeed ZeRO-3 by @kashif in #5891async_grpo: don't return on queue.Empty by @AmineDiro in #5751generate_batch: inference tensors block inplace ops in background thread by @albertvillanova in #5818 (cross-listed from v1.5 changelog window)encoding="utf-8" when reading .jinja chat templates on Windows by @ColebyPearson in #5869ValueError by pinning kernels < 0.15.1 by @albertvillanova in #5880kernels optional dependency via transformers by @albertvillanova in #5884kernels extra for transformers < 5.1.0 by @albertvillanova in #5928use_liger_kernel guard to SDPO teacher-server validation by @DaoyuanLi2816 in #5994huggingface.co/docs/openenv by @sergiopaniego in #5929bnb_4bit_quant_storage and normalize docstring param headers by @DaoyuanLi2816 in #5993online_dpo_vlm example description by @DaoyuanLi2816 in #5978SFTConfig max_length) by @DaoyuanLi2816 in #5970GRPOConfig epsilon_high help by @DaoyuanLi2816 in #5972trl skills install description by @DaoyuanLi2816 in #6008max_prompt_length argument from GRPO example by @DaoyuanLi2816 in #5964sft.json / dpo.json snapshots after transformers num_items_in_batch fix by @qgallouedec in #5845huggingface/skills by @albertvillanova in #5950.agents by @albertvillanova in #5987docker-build.yml with version parsing by @hf-security-analysis[bot] in #5920Note truncated.
🔒 Gate trainer telemetry on an explicit class-name allowlist by @qgallouedec in https://github.com/huggingface/trl/pull/5851
Full Changelog: https://github.com/huggingface/trl/compare/v1.5.0...v1.5.1
Replace deprecated torch_dtype with dtype across examples, docs, notebooks, tests, and experimental distillation / gold trainers by @qgallouedec in ht…
Three more model families gain training-compatible templates with {% generation %} markers (so assistant_only_loss=True just works):
The chunked LM-head path used by AsyncGRPOTrainer now supports models that use final_logit_softcapping (notably Gemma 2). _ChunkedLogProbFunction applies logit_scale, optional tanh-based softcapping, and temperature consistently in both forward and backward — softcapped models are no longer rejected.
by @mlarnouhet in https://github.com/huggingface/trl/pull/5691
Two more cycles closer to KTO graduation:
compute_loss flow by @albertvillanova in https://github.com/huggingface/trl/pull/5810_compute_loss_liger flow by @albertvillanova in https://github.com/huggingface/trl/pull/5816_BaseTrainer.__init__ now emits a single anonymous huggingface_hub.send_telemetry ping per trainer instantiation, so we can finally see which trainers / model families / distributed backends are actually being used in practice and prioritize accordingly.
The payload is intentionally minimal — TRL version, trainer class name, model architecture, PEFT yes/no, distributed backend (deepspeed/fsdp/ddp/none), bucketed world size, device type, GPU model when available. No user data, no dataset names, no model paths, no hyperparameter values, never sent in CI / offline / HF_HUB_DISABLE_TELEMETRY mode.
See usage_stats.md for what's collected and how to opt out.
by @qgallouedec in https://github.com/huggingface/trl/pull/5758
OpenRewardSpec: fix omitting task-scoped tools during rollout binding (fixes #5727) by @rycerzes in https://github.com/huggingface/trl/pull/5729GRPOTrainer was hanging indefinitely on truncated <tool_call> blocks (a degenerate case that happens naturally when generation hits max_completion_length mid-tool-call). Rewrote the regex to be non-backtracking — worst case goes from O(2ⁿ) to O(n). By @xodn348 in https://github.com/huggingface/trl/pull/5798OffloadActivations — follow-up to v1.4's activation-offloading leak fix. By @butterwecksolutions in https://github.com/huggingface/trl/pull/5730add_hooks by @roycho96 in https://github.com/huggingface/trl/pull/4693vocab_size for DistillationTrainer and GOLDTrainer by @Beichen-Ma in https://github.com/huggingface/trl/pull/5592empty_cache() by @jamie-peterson-ml in https://github.com/huggingface/trl/pull/5799metric_for_best_model for trainer-specific eval metrics by @qgallouedec in https://github.com/huggingface/trl/pull/5811generate_batch: inference tensors blocking inplace ops in background thread by @albertvillanova in https://github.com/huggingface/trl/pull/5818torch_dtype with dtype across examples, docs, notebooks, tests, and experimental distillation / gold trainers by @qgallouedec in https://github.com/huggingface/trl/pull/5717Glm4MoeForCausalLM / Cohere / Cohere2 / Qwen2.5-VL configs with their reference models by @qgallouedec in https://github.com/huggingface/trl/pull/5638, https://github.com/huggingface/trl/pull/5706, https://github.com/huggingface/trl/pull/5707 and https://github.com/huggingface/trl/pull/5739deepstack_visual_indexes and drop the test skip by @qgallouedec in https://github.com/huggingface/trl/pull/5779fullatt_block_indexes out of range for depth=2 by @albertvillanova in https://github.com/huggingface/trl/pull/5805num_heads key in Qwen VL tiny model scripts by @matdou in https://github.com/huggingface/trl/pull/5792model.visual. skip in GRPO/RLOO Qwen2.5-VL tests by @qgallouedec in https://github.com/huggingface/trl/pull/5780AttributeError: 'GptOssConfig' object has no attribute 'num_experts' by @albertvillanova in https://github.com/huggingface/trl/pull/5756apply_model_revisions by removing _commit_hash kwarg by @albertvillanova in https://github.com/huggingface/trl/pull/5762model.visual params by @albertvillanova in https://github.com/huggingface/trl/pull/5806torch < 2.12.0 (later reverted) by @albertvillanova in https://github.com/huggingface/trl/pull/5769pytest --only-rerun by @albertvillanova in https://github.com/huggingface/trl/pull/5784tests_latest.yml by @hf-security-analysis[bot] in https://github.com/huggingface/trl/pull/5733torch_dtype -> dtype by @qgallouedec in https://github.com/huggingface/trl/pull/5717model.visual. skip in GRPO / RLOO Qwen2.5-VL tests by @qgallouedec in https://github.com/huggingface/trl/pull/5780deepstack_visual_indexes and drop the test skip by @qgallouedec in https://github.com/huggingface/trl/pull/5779metric_for_best_model for trainer-specific eval metrics by @qgallouedec in https://github.com/huggingface/trl/pull/5811OpenRewardSpec omitting task‑scoped tools during rollout binding (fixes #5727) by @rycerzes in https://github.com/huggingface/trl/pull/5729Full Changelog: https://github.com/huggingface/trl/compare/v1.4.0...v1.5.0
A new loss_type="chunked_nll" option drastically reduces peak activation memory in SFT by avoiding the full [batch × seq × vocab] logits tensor. Ignor
<img width="2704" height="1455" alt="chunked_loss_idea" src="https://github.com/user-attachments/assets/3957f39b-3e71-4465-949a-22b2cf894d03" />
A new loss_type="chunked_nll" option drastically reduces peak activation memory in SFT by avoiding the full [batch × seq × vocab] logits tensor. Ignored-label tokens are dropped before the lm_head matmul, and the cross-entropy is computed over the remaining tokens in checkpointed chunks (default chunk_size=256, the sweet spot consistent across model sizes and sequence lengths).
from trl import SFTConfig, SFTTrainer
trainer = SFTTrainer(
model="Qwen/Qwen3-4B",
args=SFTConfig(loss_type="chunked_nll"),
train_dataset=dataset,
)
trainer.train()
Peak GPU memory, AdamW fp32:
| Model | Hardware | Seq | nll |
chunked_nll |
|---|---|---|---|---|
| Qwen3-1.7B + LoRA | 1×H100 80GB | 2048 | 47.9 GB | 12.3 GB (3.9× less) |
| Qwen3-4B | 1×H100 80GB | 16384 | OOM | 63.8 GB |
| Qwen3-14B | 8×H100 FSDP2 | 16384 | 58.9 GB | 38.9 GB (1.5× less) |
| Qwen3-32B | 8×H100 FSDP2 | 8192 | OOM | 71.2 GB |
End-to-end, chunked NLL is consistently as fast or faster than nll — and it unlocks sequence lengths that don't fit at all under the standard path.
The chunked path also supports VLMs (https://github.com/huggingface/trl/pull/5684).
by @qgallouedec in https://github.com/huggingface/trl/pull/5575, https://github.com/huggingface/trl/pull/5676 and https://github.com/huggingface/trl/pull/5684
A new trl.experimental.openreward adapter plugs any environment speaking the Open Reward Standard (ORS) protocol into any TRL trainer accepting an environment_factory (GRPOTrainer, AsyncGRPOTrainer). One identifier wires all three trainer slots — dataset, factory, reward_func:
from trl import GRPOConfig, GRPOTrainer
from trl.experimental.openreward import OpenRewardEnv
env = OpenRewardEnv("Eigent/SETA") # or "http://localhost:8000"
trainer = GRPOTrainer(
model="Qwen/Qwen3-4B",
args=GRPOConfig(...),
train_dataset=env.dataset,
environment_factory=env.factory,
reward_funcs=env.reward_func,
)
Tools are bound dynamically from JSON Schema at construction (no per-env wrapper code), and env.dataset autoderives task lists from the ORS task endpoints. The same code path works for envs hosted on the OpenReward platform, self-hosted on any container service, or running locally on localhost. A SETA training example is included.
by @adithya-s-k in https://github.com/huggingface/trl/pull/5696
Unit tests don't catch trainer-level numerical drift (gradient-accumulation normalization bugs, attention-impl divergence (eager ↔ FA2 / kernels)) they silently shift the loss trajectory and users only notice when their run no longer reproduces. (Cf. last year's transformers grad-accum bug, or the "We found two bugs in DeepSpeed" paper.)
A new opt-in pytest -m invariant suite asserts the loss / grad_norm trajectory of short end-to-end SFT/DPO runs against committed reference snapshots, with equivalence classes for configs that should produce identical trajectories (e.g. pdb=1, gas=8 ≡ default; eager ≡ FA2 ≡ kernels). Hardware-pinned to H100 80GB, real pretrained model, full_determinism, fixed seed. Initial coverage: 2 trainers × 2 invariance axes (grad-accum, attn-impl) × gradient-checkpointing equivalence.
by @qgallouedec in https://github.com/huggingface/trl/pull/5686, https://github.com/huggingface/trl/pull/5688 and https://github.com/huggingface/trl/pull/5689
Three new pure helpers in trl.trainer.utils for measuring training efficiency:
compute_flops_per_token(config, seq_len) — handles dense and MoE (Mixtral, Qwen3-MoE, DeepSeek-V2)compute_mfu(flops_per_token, tps, world_size, peak_flops) — Model FLOPs Utilization as a percentageadjusted_mfu(mfu, config, seq_len) — non-causal → causal-corrected (Llama / DS Ulysses convention)by @AmineDiro in https://github.com/huggingface/trl/pull/5698
GRPO's Liger-kernel integration is updated for Liger 0.8.0: delta two-sided clipping, use_bias_correction_kl, and SAPO/VESPO parameters are now forwarded into LigerFusedLinearGRPOLoss. The previous delta + use_liger_kernel guard is removed — both can be combined.
by @kashif in https://github.com/huggingface/trl/pull/5690
A new loss_type="sigmoid_norm" option for DPOConfig implements the per-token (length-normalized) DPO loss used by Tülu 3 / OLMo (paper §5.1.2 eq. 6) to mitigate length bias.
from trl import DPOConfig, DPOTrainer
trainer = DPOTrainer(
model="Qwen/Qwen3-4B",
args=DPOConfig(loss_type="sigmoid_norm"),
train_dataset=dataset,
)
by @BrownianNotion in https://github.com/huggingface/trl/pull/5406
Four more model families gain training-compatible chat templates with {% generation %} markers (assistant-only loss masking) and/or response schemas (tool-calling parsing):
{% generation %} markers by @qgallouedec in https://github.com/huggingface/trl/pull/5675get_training_chat_template now also accepts a processor (not just a tokenizer) — useful for VLMs (https://github.com/huggingface/trl/pull/5560).
Another batch of alignment PRs this cycle. KTO and DPO are now structurally aligned across PEFT handling, model initialization, training-arg grouping, ref-logp precomputation, and metric handling — promotion of KTO out of experimental is imminent.
PRs (all by @albertvillanova): #5659, #5660, #5661, #5679, #5701, #5702, #5703, #5704, #5705, #5714.
parallelism_config with cp_size>1 or sp_size>1 in GRPO/RLOO — fail fast at config init with a clear error instead of mid-training crash. By @kashif in https://github.com/huggingface/trl/pull/5699model_accepts_loss_kwargs=False in DPO and Reward by @albertvillanova in https://github.com/huggingface/trl/pull/5710_tokenizer attribute in experimental trainers by @albertvillanova in https://github.com/huggingface/trl/pull/5566peft_config handling in core / experimental trainers by @albertvillanova in https://github.com/huggingface/trl/pull/5673 and https://github.com/huggingface/trl/pull/5674isinstance with is_peft_model / drop redundant is_peft_available by @albertvillanova in https://github.com/huggingface/trl/pull/5682 and https://github.com/huggingface/trl/pull/5683parse_response by @qgallouedec in https://github.com/huggingface/trl/pull/5561OffloadActivations.__exit__ now syncs the compute/offload streams and clears the stash dictionaries, preventing orphaned offload tensors from leaking onto a dead stream (~0.2 GiB/step accumulation observed during QLoRA vision training before the fix). By @butterwecksolutions in https://github.com/huggingface/trl/pull/5694 and https://github.com/huggingface/trl/pull/5700DistillationTrainer by @k1064190 in https://github.com/huggingface/trl/pull/5594GKDTrainer: fix return_outputs in the Liger kernel path by @roycho96 in https://github.com/huggingface/trl/pull/4688GKDTrainer: fix seq-KD wasted teacher forward by @roycho96 in https://github.com/huggingface/trl/pull/5726GKDTrainer: fix Liger fused JSD path computing wrong loss by @roycho96 in https://github.com/huggingface/trl/pull/5731peft_config to core / experimental trainers by @albertvillanova in https://github.com/huggingface/trl/pull/5664 and https://github.com/huggingface/trl/pull/5665peft_config type hint in experimental trainers by @albertvillanova in https://github.com/huggingface/trl/pull/5666DistillationTrainer by @cmpatino in https://github.com/huggingface/trl/pull/5615Qwen3-4B-Instruct-2507 by @qgallouedec in https://github.com/huggingface/trl/pull/5586Qwen/Qwen3-30B-A3B by @qgallouedec in https://github.com/huggingface/trl/pull/5716DistillationTrainer by @cmpatino in https://github.com/huggingface/trl/pull/5615{% generation %} markers for Cohere2 chat template by @qgallouedec in https://github.com/huggingface/trl/pull/5675get_training_chat_template by @qgallouedec in https://github.com/huggingface/trl/pull/5560parse_response by @qgallouedec in https://github.com/huggingface/trl/pull/5561Full Changelog: https://github.com/huggingface/trl/compare/v1.3.0...v1.4.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →