NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2487 most downloaded on PyPI
Train transformer language models with reinforcement learning.
Last release 5 days ago
29 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
103 releases · first in 2020
One column per quarter.
Fix the server-mode crash at the first weight sync on vLLM 0.20 to 0.25 by @albertvillanova in #7410
Full Changelog: v1.14.0...v1.14.1
trl.losses is gone: DPO, KTO and GRPO stream their own log-probs
trl.losses is gone: DPO, KTO and GRPO stream their own log-probsWarning
from trl.losses import FusedLinearDPOLoss (or FusedLinearKTOLoss, FusedLinearGRPOLoss, FusedLinearJSDLoss) no longer works. The module introduced in v1.13 has been removed.
v1.13 vendored Liger's chunked_loss into trl.losses as a holding action, not a destination (#7063). The problem it was holding: those classes reimplement each trainer's loss math, so TRL carried two copies of every formula, and copies drift. That drift is where the bugs were. GRPO's grad_norm differed from the default path, KTO ignored the reference model under PEFT, KTO class weights and DPO label_smoothing were silently dropped. Worst of all, under plain DDP the fused path ran the loss on the unwrapped model, so DistributedDataParallel's reducer was never armed and gradients were never all-reduced: every rank silently kept its own.
The fix is to stop reimplementing. All DPO, KTO and GRPO need from the fused path is per-token log-probs of the selected tokens, and TRL already had _ChunkedLogProbFunction for exactly that: it streams the vocabulary with an online logsumexp and recomputes in the backward pass. Each trainer now runs backbone -> _ChunkedLogProbFunction -> its own existing loss code, unchanged. One implementation of each loss, and the memory win is kept, because full logits are still never materialized.
Thirteen restrictions existed only because the loss had been rewritten in a form that could not express those options. Most are gone: DPO with use_liger_kernel=True now accepts mixed loss types, f-divergences and precomputed reference log-probs, and GRPO gains entropy and off-policy masking. Still refused: use_weighting, compute_metrics, a PEFT adapter on lm_head, prompt-learning PEFT, and the MoE auxiliary loss.
Single H100, Qwen3-0.6B, bs=4, seq 512:
| config | median step | peak |
|---|---|---|
| v1.13 fused, inline ref | 0.2243 s | 7.05 GB |
| v1.14 chunked, inline ref | 0.2457 s (+9.5%) | 7.05 GB |
v1.14 chunked, precompute_ref_log_probs=True |
0.1971 s (-12.1%) | 4.83 GB (-31%) |
The third row is the point: that configuration did not exist before, because the fused path rejected precomputed reference log-probs outright.
use_liger_kernel=True still works and still enables Liger's model kernels through transformers. In DPO, GRPO and KTO it now selects TRL's chunked log-probability path instead of Liger's fused loss.
The loss head is the only hot path TRL owns: transformers already kernelizes the model internals, and nothing on the Hub covers what happens after the decoder. On one H100 (bf16, V=151936, H=4096, 8192 tokens), selective_log_softmax plus entropy_from_logits cost 12.8 ms and 2.32 GiB over [8, 1024, 151936]. A fused Triton kernel does the same in 0.89 ms, roughly 14x, and lands closer to the fp32 reference than the bf16 path it replaces. GRPO runs this two to four times per step.
It now lives in-tree at trl.kernels and is on by default on CUDA, ROCm and XPU, with the PyTorch implementation as fallback. It was briefly loaded from the Hub; that was reversed because this kernel is pure Triton and compiles at runtime, so the Hub bought nothing while costing a version gate unrelated to whether the kernel works, a download that fails offline, trust_remote_code=True for anyone with kernels installed, a silent except Exception: pass wrapped around all of it, and 13 of 13 tests skipped in CI.
GRPOWithReplayBufferTrainer, BCOTrainer, PRMTrainer, XPOTrainer, NashMDTrainer and the GSPO-token trainer are gone from trl.experimental.
What changed is not the guarantee, which was always that trl.experimental has none, but the fact that the decision is now written down. Removals used to be argued one pull request at a time, with every thread relitigating the policy before it got to the feature. #7182 settled what the judgment weighs, and #7254 put it in the docs: usage, external issues and pull requests that come from people actually running the feature, whether a stable trainer already covers it, maintenance cost, downstream consumers, owner, and age. Deliberately inputs to a judgment rather than a test: no threshold, no notice cycle, and the numbers stay in the removal pull request where you can check them. A paper index entry outlives the code it described.
The numbers are what make the case. BCOTrainer was 1,600 lines carrying its own reference-model handling, DeepSpeed preparation, tokenization and eval loop, for 11 genuine trainings in two months; it was also the only reason TRL depended on scikit-learn and joblib at all, and both are now gone from pyproject.toml. GRPOWithReplayBufferTrainer had hand-copied _generate_and_score_completions from GRPO and drifted to 385 of 616 lines different, for 8 trainings.
Also breaking, and covered above: trl.losses was removed along the way.
compute_loss redirected through the distributed wrapper only for ZeRO-3 and FSDP, so on plain DDP the loss ran on the unwrapped model, DistributedDataParallel.forward() never ran, its reducer was never armed, and every rank silently kept its own gradients. By @qgallouedec in #7247_pack_wrapped sized the output from the sliced view but read from the raw child buffer, so it re-read from the start of the buffer instead of from the slice, and datasets.map(batched=True) passes sliced tables. By @qgallouedec in #6670log_metric on rank 0 alone hung on NCCL, and log_extra crashed inside gather_object with a 1EB allocation. By @behroozazarkhalili in #7382Trainer actually reads, so it no longer realigns them at train time and logs a change the user did not make (part of #7093): eos in the trainers (#7127), written to the text config for composite models (#7315), the reward model pad token (#7318), and the requested pad token applied to vision datasets (#7316, #7352)trl.experimental by @albertvillanova in #7254, the rule the six removals above were made underTiny models, a sweep bringing every tiny test model in line with its reference, all by @albertvillanova:
tie_word_embeddings across the board (#7130)use_mrope (#7332), use_mambapy (#7333), rope_scaling (#7334), cache_implementation and hidden_act (#7335), activation_dropout and prefix (#7336), attention_bias (#7337, #7338)head_dim reduced for Gemma (#7216) and Gemma2 (#7227), ffn_dim passed for OPT (#7232), the Phi3 leftover replaced by the 3 and 3.5 variants (#7239), three with no test consumer retired (#7186)Tests run only when they can be affected, all by @albertvillanova:
main (#7172) and then reverted (#7201), because GitHub buckets cancelled with failure and main looked broken.Release automation. Publishing now happens on a tag push, and refuses unless VERSION in the tagged commit matches the tag (#7165). Release notes come out sectioned, from a label-to-section map (#7166) with documentation and maintenance labels applied automatically from the changed paths (#7205, #7210). Plus the dev version bump (#7160).
Test reliability, all by @albertvillanova: the rerun filter matches the out-of-memory message rather than the exception type (#7292) and needs pytest-rerunfailures>=16.7 to see wrapped exceptions (#7325); ACCELERATE_MIXED_PRECISION is restored between tests (#7321); the Llava assistant-mask tests are un-xfailed now that transformers fixed them (#7373); non-identity adapters in the use_adapter tests (#7275); MODEL_REVISIONS injected into the PEFT loaders (#7288); the reward model pad token covered (#7350).
Workflow hardening. Token permissions narrowed in the test workflows (#7287), the workflows hardened (#7300), and the upstream issue behind the trufflehog digest pin linked (#7214).
Other. Dependabot bumps (#7192, #7299), the latest-release tests moved from T4 to L40S (#7184) with the Slack action checkout workaround dropped (#7183), the kernel publish job run under a personal namespace (#7343), the token sync comments reworded (#7167) and all three generation-config eos shapes covered (#7168), LM-head gradient buffers and GEMMs skipped when its parameters are frozen (#7076), and the end-of-turn loss mask warning made actionable (#7230).
The supported vLLM window moves from >=0.19.1,<=0.28.0 to >=0.20.0,<=0.30.0.
Note truncated.
A new long-context guide plus a runnable example that trains a book-length sequence per step on a single 8×H100 node.
A new long-context guide plus a runnable example that trains a book-length sequence per step on a single 8×H100 node.
Measured: Qwen3-8B at 1,048,576 tokens, 380 s/step, 56.2 GB per GPU (bf16, per_device_train_batch_size=1, loss_type="chunked_nll"). Runs end to end and saves. The guide documents the levers in the order you hit them (chunked_nll, gradient-checkpointing offload, YaRN RoPE) and the constraints (full attention only, no packing). Needs transformers >= 5.16 for checkpointing offload.
lm_head projection now runs on tensor coresThe inner loop of the default loss_type="chunked_nll" was doing h.float() @ w.float().t(). Both operands are already bf16, so the upcast bought no information: it moved the GEMM off the tensor cores onto the fp32 SIMT path and materialized an fp32 copy of the whole lm_head weight (2.03 GB for a 248k vocabulary), rebuilt for every chunk and again on every gradient-checkpoint recompute. In an 8×H100 profile of trl sft on Qwen3.6-35B-A3B those two fp32 SIMT GEMMs were 21.6% of all GPU kernel time.
One chunk, 256 tokens × vocab 248,320 × hidden 2048, 1×H100, bf16, fwd+bwd:
| fwd+bwd | peak mem | |
|---|---|---|
| before | 23.37 ms | 5.99 GB |
| after | 3.86 ms (6.0×) | 3.03 GB |
End to end, 16,384 tokens per step, tokens/s/GPU:
| model | mode | before | after | |
|---|---|---|---|---|
| gemma-3-270m (vocab 262k) | full FT, 1×H100 | 24036 | 31609 | 1.32× |
| Qwen3-0.6B (vocab 152k) | full FT, 1×H100 | 21276 | 26641 | 1.25× |
| Qwen3-8B | full FT, 2×H100 FSDP2 | 3554 | 6009 | 1.69× |
| Qwen3-8B | LoRA r16, 2×H100 FSDP2 | 4531 | 7125 | 1.57× |
| Qwen3-30B-A3B (MoE) | LoRA attn, 2×H100 FSDP2 | 4251 | 5101 | 1.20× |
Same fix applied to the distillation trainers (which paid it twice per chunk, student and teacher). Under accelerate mixed precision the numerics are bit-identical (autocast was already casting the operands back down); without autocast over the projection the GEMM moves to bf16 tensor cores, which is exactly what loss_type="nll" already does.
by @qgallouedec in #6863
trl.lossesTRL's only import from liger-kernel was liger_kernel.chunked_loss (the fused linear DPO / KTO / GRPO / JSD losses). That module is pure PyTorch, and most of the GRPO and DPO variants in it were written for TRL. Upstream reviews had stalled, so the code comes home as trl.losses: FusedLinearDPOLoss, FusedLinearKTOLoss, FusedLinearGRPOLoss, FusedLinearJSDLoss (copied from Liger-Kernel v0.8.2, BSD-2 notice kept, bitwise identical to the installed Liger on random inputs).
No new config flag: use_liger_kernel=True keeps selecting the fused loss exactly as before, and still needs liger-kernel installed because transformers' Trainer patches the model kernels with it.
trl.losses by @qgallouedec in #7059DistillationTrainer by @kashif in #7064.agents/skillsAgent skills for contributing to TRL now ship in the repo at .agents/skills.
by @qgallouedec in #6901
quantization_config argumentThe QLoRA test suites move onto the quantization_config trainer argument added in v1.8, and gain coverage on the VLM paths.
quantization_config by @albertvillanova in #6910quantization_config trainer argument by @albertvillanova in #6927peft>=0.13.0 and drop the 0.12 version guards by @qgallouedec in #7102deepspeed>=0.18.6 and drop the 0.16.4 guard by @albertvillanova in #7118target_parameters use for peft<0.17.0 by @albertvillanova in #7119num_tokens under tensor parallelism by @qgallouedec in #7101output_hidden_states when only last_hidden_state is used by @ciaoyizhen in #4755processing_class by @kashif in #6978AsyncDistillationTrainer up with recent AsyncGRPO changes by @kashif in #6823 and #6987dataset_formatting module by @albertvillanova in #7117FLASH_ATTENTION_VARIANTS constant from the DPO trainer by @albertvillanova in #7013PPOTrainer, PPOConfig, and modeling_value_head.py (PreTrainedModelWrapper, AutoModelForCausalLMWithValueHead, AutoModelForSeq2SeqLMWithValueHead) are gone, along with their tests, doc page, and examples.
Bit of a moment for TRL: PPOTrainer predates every PR in this repo. It landed in dfb6a580 on 2020-03-28, the commit that first added the library, back when the package was called lm_ppo. It was the oldest thing in TRL and the last piece of the original codebase.
Why now: unmaintained (no feature work in over a year, every PPO commit since #4482 a drive-by fix), the only trainer never aligned on the input format (it still took tokenized input_ids), near-zero recorded usage, and a magnet for automated bug hunters filing real reports against code nobody runs. Nobody loses anything: from trl import PPOTrainer already stopped working in v1.10, so anyone actually running PPO is pinned to an older TRL and those installs keep resolving exactly as they do today.
create_reference_model stays (BCO, A2PO and Online DPO use it).
by @qgallouedec in #7020
_ChunkedLogProbFunction backward — backward accepted grad_entropy but never used it, so any loss backpropagating through the entropy output silently got zero gradient from that term, with no error or warning. Only AsyncGRPOTrainer read entropy (under no_grad, for logging), so no user-facing regression today, but this is exported non-experimental code and an entropy-bonus loss term would have hit it silently. By @verma8076 in #6625paper_index paths and add GOLD, IW-OPD, and AsyncDistillation telemetry by @YeonwooSung in #7051A docstring and docs consistency pass across the repo, all by @qgallouedec:
DistillationTrainer signature annotations (#7152), docs formatting pass (#7154), tidy pyproject / Makefile / .gitignore (#7157)Docker images — released and dev builds are now separate workflows, the image version comes from the release tag, and the installed TRL version is pinned:
docker-build.yml by @hf-security-analysis[bot] in #6997Tiny-model config alignment — a sweep bringing every tiny test model's config in line with its reference model, all by @albertvillanova:
layer_types so tiny Gemma3 / Olmo3 cover both attention types (#6962) and tiny Cohere2 (#6963)Other CI, mostly by @albertvillanova:
select instead of extend-select to keep the rule set explicit (#6964), stop dependabot from opening ruff pre-commit bumps (#6965), reformat lambda functions for code clarity (#7092)select instead of extend-select to keep the rule set explicit by @albertvillanova in #6964Note truncated.
⚠️ v1.12.0 is an accidental duplicate of v1.11.0
During the v1.11.0 release, VERSION was bumped to 1.12.0 instead of 1.12.0.dev0. Our publish workflow fires on any push to main that touches VERSION and only skips the upload when the version contains "dev", so trl 1.12.0 was auto-published to PyPI at 19:46 on 2026-08-26, 90 seconds after 1.11.0, with identical code.
trl==1.12.0 on PyPI is bit-identical to trl==1.11.0. There are no features, fixes, or behavior changes in this version.
The version number is burned (PyPI does not allow yanking + re-uploading the same version), so we skip 1.12 entirely for real development. The next feature release is v1.13.0.
Full Changelog: v1.11.0...v1.12.0 (empty)
Ignore benign NumPy __array_wrap__ deprecation and torch inductor TF32 advisory warnings in CI by @albertvillanova in #6850 and #6851
trl vllm-serve used to run a custom FastAPI app around vLLM's LLM, with its own data-parallel fan-out and a worker extension for weight sync. vLLM's own server covers all of that. The custom server is gone: trl/scripts/vllm_serve.py drops from 1218 to ~130 lines and now translates its flags into vllm serve, prints the equivalent command, forwards any extra argument, and sets the three settings TRL needs (--weight-transfer-config, --logprobs-mode processed_logprobs, --max-logprobs -1). fastapi, uvicorn and pydantic leave the vllm extra.
Weight sync now uses vLLM's NCCL weight-transfer engine — tensors are announced once and streamed in a single packed broadcast, instead of one HTTP request + broadcast per tensor. Distillation teacher logprobs come from prompt_logprobs.
Benchmark (GRPO server-mode on H100, Qwen2.5-1.5B / Qwen2.5-VL-3B, same seed, same data):
| Setup | before | after | speedup |
|---|---|---|---|
| Text, TP=1 | 1.08, 0.91 s/step | 0.63, 0.62 s/step | 1.59× |
| Text, TP=2 | 0.96, 0.96 s/step | 0.62, 0.60 s/step | 1.57× |
| VLM, TP=1 | 2.14, 2.13 s/step | 1.48, 1.48 s/step | 1.44× |
Bit-identical completions vs the old server (16/16 over 4 prompts × n=4, logprobs identical to 0.0 over 499 tokens). trl vllm-serve still accepts the same flags — it just forwards to vllm serve now.
by @qgallouedec in #6765
AsyncDistillationTrainer (+ multi-teacher MOPD)Async on-policy distillation, architected like AsyncGRPOTrainer: a background rollout worker generates the student's own completions and scores them against a teacher served over HTTP, so generation and training overlap instead of alternating. The teacher is never loaded locally — just a vLLM server URL.
Also supports MOPD (multi-teacher on-policy distillation): pass more than one entry in teacher_server_urls and route each sample to a teacher via a teacher_id column (e.g. a math teacher and a code teacher, each served independently).
from trl.experimental.async_distillation import AsyncDistillationConfig, AsyncDistillationTrainer
trainer = AsyncDistillationTrainer(
model="Qwen/Qwen3-4B",
args=AsyncDistillationConfig(
teacher_server_urls=["http://teacher-math:8000", "http://teacher-code:8000"],
teacher_top_k=16,
),
train_dataset=dataset, # rows carry a `teacher_id` column for MOPD
)
trainer.train()examples/ was split by format (scripts/ vs notebooks/), which scattered related files (e.g. sft_nemotron_3 existed as both a script and a notebook, in different folders; OpenEnv notebooks lived apart from OpenEnv scripts). Everything is now one folder per story — the name says the method and the task (grpo_wordle, sft_gpt_oss, ppo_tldr), and the folder holds every file the example needs (scripts, notebooks, prompts, chat templates, eval code). Thin single-trainer demo scripts are removed (they live in trl/scripts/ and each trainer doc page has a runnable snippet). Old examples/scripts/… and examples/notebooks/… GitHub / Colab links will 404 — worth flagging in downstream posts.
by @qgallouedec in #6820
AsyncGRPOConfig by @AmineDiro in #6774AsyncGRPOTrainer — drop redundant metric guards, merge train-begin callbacks by @qgallouedec in #6378trl skills installThe TRL skill is now written around the Python API — trainer selection, dataset format, load-bearing config fields, GRPO reward-function contract — with the CLI as a footnote. trl skills install is dropped and the skill source moves out of the package.
trl skills install and move the skill source out of the package by @qgallouedec in #6793DistillationTrainer supports tool callingContinuing the graduation cycle from v1.10.
by @qgallouedec in #6723
DistillationTrainer when the GKD config is fully coveredGKDTrainer now emits a hint when the config is fully covered by the stable DistillationTrainer — no behavior change, just a nudge for users still on the older API.
by @qgallouedec in #6847
position_ids from the SFT collator when sequence parallelism is enabled by @qgallouedec in #6812chunked_nll loss by @qgallouedec in #6842dataset_source to telemetry by @qgallouedec in #6599liger-kernel 0.8.2 and drop the SAPO warning filter by @albertvillanova in #6768None token logprobs from vLLM in GRPO importance sampling by @behroozazarkhalili in #6693[GRPO] Match the Liger dapo / cispo / vespo normalizer to the non-Liger path — completes the fix from v1.10 (#6024) on the Liger path. By @kashif in #5890[SFT] Treat past_key_values and attentions as optional on the model output by @Hakureirm in #6686AsyncGRPOTrainer crashing on environment-owned datasets by @albertvillanova in #6824scale_rewards CLI parsing and warn on ignored feedback by @sergiopaniego in #6341opencode.py proxy trace / log paths after OpenEnv moved them to functions by @sergiopaniego in #6672DistillationTrainer when GKD config is covered (see above)Big CI-hygiene cycle by @albertvillanova. Highlights:
head_dim when the model config does not declare it by @albertvillanova in #6855__array_wrap__ deprecation and torch inductor TF32 advisory warnings in CI by @albertvillanova in #6850 and #6851lob / postgres detectors by @albertvillanova in #6920, #6921 (+ #6926)GITHUB_ACTIONS by @albertvillanova in #6726, #6753, #6749VERSION env from the Docker build step / fix the Slack title of the TRL Docker image build job by @albertvillanova in #6931, #6932RUN_SLOW env variable from the slow tests workflow by @albertvillanova in #6748docstyle-ignore tag from docstrings by @albertvillanova in #6785Note truncated.
The full refactor arc: switched signature columns to prompt (deprecated messages -format), pinned the generation stack to GRPO's, wired the chunked JS…
DistillationTrainer is now a stable trainerAfter a ~30-PR refactor that reshaped its data contract, loss path, generation stack, config surface, and tests, DistillationTrainer and DistillationConfig graduate from trl.experimental.distillation to the top-level trl package. Same import surface as SFT / DPO / GRPO / KTO. The old experimental path still works and emits a FutureWarning (removal in v2.0.0).
# Before
from trl.experimental.distillation import DistillationConfig, DistillationTrainer
# Now
from trl import DistillationConfig, DistillationTrainerAlso lands a trl distillation CLI and moves the tests to tests/. The full refactor arc: switched signature columns to prompt (deprecated messages-format), pinned the generation stack to GRPO's, wired the chunked JSD loss and deleted the full-logit path, cleaned the Liger path to share extraction with the chunked path, adopted GRPO's log() / training_step timing / _save_checkpoint, rebuilt the test suite to GRPO shape, and much more.
by @qgallouedec across ~30 PRs (#6479, #6480, #6481, #6482, #6484, #6487, #6497, #6508, #6509, #6510, #6511, #6512, #6513, #6521, #6522, #6523, #6524, #6525, #6526, #6530, #6537, #6604, #6605, #6606, #6607, #6609, #6610, #6611, #6612, #6613, #6614, #6629, #6630, #6632, #6633, #6634, #6639, #6640, #6641, #6642, #6643, #6644, #6645, #6647, #6653).
DistillationTrainer supports Vision Language ModelsAlongside the graduation, VLMs work end-to-end in DistillationTrainer.
by @qgallouedec in #6650
New experimental loop-owning (black-box) path for AsyncGRPOTrainer — for training external agents like opencode that run their own tool loop, rather than TRL driving each turn as in the environment_factory white-box path.
The agent runs in an OpenEnv session in transparent_proxy mode; an in-sandbox proxy captures each turn's token ids and logprobs. On completion, TRL reads the trace, rebuilds per-turn training rows, and scores the workspace with the session's verify(). Ships with three caller-supplied policy hooks (rollout_reward_fn, train_turn_fn, agent_turn_fn) so you can drop framework aux calls (title generator, context summarizer) and reinforce only the turns you want.
Includes a self-contained examples/scripts/openenv/opencode.py (local subprocess sandbox + DeepCoder held-out stdin/stdout verifier), validated end-to-end on Qwen3-4B.
by @AmineDiro in #6420, plus HF-sandbox variant by @sergiopaniego in #6565
top_p / top_k / min_p / repetition_penalty in AsyncGRPOConfig by @AmineDiro in #6608A new SFT example for google/diffusiongemma-26B-A4B-it implementing the reference recipe: one response block per step, uniform random token corruption with t ~ U(eps, 1), two-pass self-conditioning, flat cross-entropy over the whole canvas plus an autoregressive co-loss on the encoder. LoRA targets attention + dense MLP linears. Requires transformers >= 5.12.0.
Ships with diffusion_gemma.jinja / diffusion_gemma_training.jinja chat templates (with {% generation %} markers) so assistant_only_loss=True works out of the box.
Two default flips this cycle — pinning them explicitly is recommended if you want the old behavior.
SFTConfig.max_completion_length / GRPOConfig.max_completion_length: default bumped from 256 to 512. By @dhruvnigam93 in #6264GRPOConfig.use_bias_correction_kl now defaults to True. By @gowtham-sai-yadav in #6503VLM configurations no longer reject packing / padding-free when the batch happens to be text-only.
by @cris96spa in #6547
target_parameters on peft>=0.20.0 by @albertvillanova in #6591train_dataset types in core trainers by @albertvillanova in #6493 and #6494f_divergence_type / loss type combinations in DPO by @qgallouedec in #6559TQDM_DISABLE in DPO/KTO/BCO reference log-prob loops by @qgallouedec in #6507use_cpu in create_model_from_path device_map default by @Strongich in #6295dtype instead of torch_dtype (deprecated) by @qgallouedec in #6560test_train_moe_peft_model for KTO by @albertvillanova in #6589steps_per_generation != gradient_accumulation_steps — gradients were silently mis-scaled by gradient_accumulation_steps / steps_per_generation (e.g. 0.5× or 4× the intended token-mean in common configs). Default-equivalent configs are unaffected. By @0xadvait in #6024vllm_enable_sleep_mode=True by @muupan in #5313cuda in FSDP2 vLLM weight sync by @verma8076 in #6592prepare_deepspeed crash with CPU offload optimizer by @roycho96 in #5916fix(ppo): exclude padding tokens from entropy calculation by @mukund1985 in #6121fix: avoid CopySlices when scaling policy logits by @DaoyuanLi2816 in #6554[GRPO] Apply the completion mask elementwise in the LuSPO loss aggregation by @YaseenBashaT in #6654[GRPO] Fix entropy bonus normalization inconsistency across loss types by @YaseenBashaT in #6648huggingface/OpenEnv by @sergiopaniego in #6529GRPOConfig loss_type help by @gowtham-sai-yadav in #6477opencode_hf_sandbox.py: loop-owning training in remote HF sandboxes by @sergiopaniego in #6565test: don't hard-code bf16=True on devices that lack bf16 support by @behroozazarkhalili in #6036use_cpu in tests when no CUDA device is available by @albertvillanova in #6504LOAD REPORT tables in test output by @albertvillanova in #6502jmespath only on transformers below 5.13.0 by @albertvillanova in #6597train_dataset=None raises for core trainers by @albertvillanova in #6492loss_type in KTO precompute ref log-probs test by @albertvillanova in #6501test_peft_with_quantization tests for bitsandbytes 0.50.0 by @albertvillanova in #6546torch.get_autocast_gpu_dtype() warning from mamba-ssm kernel by @albertvillanova in #6556add_response_schema when the model already ships a response template by @qgallouedec in #6574openenv_harness.py by @albertvillanova in #6531liger-kernel < 0.8.1 by @albertvillanova in #6517bitsandbytes < 0.50.0 by @albertvillanova in #6544ci: exclude the buggy TruffleHog LOB detector by @behroozazarkhalili in #6674["prompt", "image", "images"] by @qgallouedec in #6481completion_mask alongside labels by @qgallouedec in #6482completion_mask by @qgallouedec in #6484Note truncated.
Revert xfail for NemotronH GRPO/RLOO tests now that transformers#47569 fixed the kernels bug by @albertvillanova in #6558
Full Changelog: v1.9.1...v1.9.2
Fix vLLM server-mode communicator initialization to use the current A… by @sywangyi in #6417
Full Changelog: v1.9.0...v1.9.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →