NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #5538 most downloaded on PyPI
NeMo - a toolkit for Conversational AI
Last release 1 months ago
07 Aug 2026
Release timing varies
gaps range from 2 weeks to 4 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
109 releases · first in 2019
Speech Data Explorer gained S3 reading, comparison mode, security fixes, and tutorial updates (#15500, #15137).
NeMo Speech 3.0 is the first major release after the repo split and rename to NVIDIA-NeMo/Speech. The repo now focuses on ASR, TTS, audio processing, speaker tasks, and SpeechLM. Non-speech Framework, LLM, VLM, diffusion, export/deploy, evaluator components can now be found under separate repos under NVIDIA-NeMo organization.
While this release brings new major features, the central focus is addressing technical debt: removed 800k deprecated LOC, migration to uv package manager for cleaner installs, stronger model test coverage, reduced the number of dependencies, revamped documentation, lighter containers, and AGENTS.md + agentic skills.
examples/ and docs under docs/source/.uv sync reproduces the supported NeMo Speech stack and may replace Python/PyTorch/CUDA inside .venv. To keep your own stack, install PyTorch first, then use uv pip or pip.SpeechLM2 now supports training SpeechLM models with NeMo Automodel through SALMAutomodel, targeting both dense and MoE LLM backbones such as Nemotron 3.
SALMAutomodel adds an Automodel-backed SALM path with native LoRA, deferred configure_model() initialization, shard-aware distributed loading, and Automodel-owned mesh creation (#15447).AutomodelParallelStrategy supports FSDP2, HSDP, TP, CP, and EP, letting SpeechLM training combine sharded data parallelism, tensor/model partitioning, long-context sequence sharding, and MoE expert routing (#15447, #15648, #15679, #15773).NeMo Speech 3.0 adds the train/eval modules behind Nemotron VoiceChat.
NemotronVoiceChat STT + TTS class for validation, offline speech-to-speech inference, and export workflows (#15456).Note: NemotronVoiceChat class is inference/eval only. Train DuplexSTTModel and DuplexEARTTS/speech-decoder modules separately.
SpectrogramToAudio (#14524).NVIDIA-NeMo/Speech (#15783, #15788).uv sync --extra all --extra cu13 (#15769).uv pip or pip, addressing common install feedback (#15769).uv.lock and simplify release builds (#15659, #15668, #15685, #15697).r3.0.0; #15783 bumps the next release to 3.0.uv.lock.torch >= 2.6.0; #15353 documents TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD; #15506/#15018 fix CUDA Python/numba-cuda handling.diskcache; #15502 removes protobuf; #15183 removes CUDA bindings pins; #15438 relaxes kaldialign.trust_remote_code.SECURITY.md.SALMAutomodel with long-context support, encoder chunking, activation-checkpointing controls, and stability fixes.chat_template support.SpectrogramToAudio.LhotseSpeechToTextBpeDataset.input_cfg.yaml directly from AIS buckets.LazyNeMoIterator..nemo torchaudio-preprocessor checkpoint migration.use_bucketing and validates batch size..nemo tar extraction/config processing.PurePosixPath.assert with ValueError in EncDecMultiTaskModel.UTMOSv2Calculator.verbose.One column per quarter.
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
<details><summary>Changelog</summary>
fix: Remove diskcache from requirements (15630) into r2.7.0 by @chtruong814 :: PR: #15635</details>
cp: Fix numba-cuda and cuda-python installation and usage (#15506) by @chtruong814 :: PR: #15540
<details><summary>Changelog</summary>
numba-cuda and cuda-python installation and usage (#15506) by @chtruong814 :: PR: #15540</details>
<details><summary>Changelog</summary>
</details>
cp: Fix cuda-python usage for CUDA graphs (#15416) by @ko3n1g :: PR: #15471
<details><summary>Changelog</summary>
Fix cuda-python usage for CUDA graphs (#15416) by @ko3n1g :: PR: #15471</details>
<details><summary>Changelog</summary>
</details>
Add deprecation notice to modules by @chtruong814 :: PR: #15050
Starting with the next release, NeMo 2.8.0, the following collections will be removed: avlm, diffusion, llm, multimodal, multimodal-autoregressive, nlp, speechlm, vision, vlm, and this repo will focus solely on speech tasks: ASR, TTS, speaker diarization, and speech enhancement.
<details><summary>Changelog</summary>
timestamps=True by @artbataev :: PR: #15298</details>
<details><summary>Changelog</summary>
NemoSTTService by @SangwonSUH :: PR: #15233</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
resume_from_path by @maanug-nv :: PR: #14966</details>
<details><summary>Changelog</summary>
2.7.0rc0.dev0 by @github-actions[bot] :: PR: #14956v2.5.1 by @github-actions[bot] :: PR: #14967r2.5.0 by @github-actions[bot] :: PR: #14990v2.5.3 by @github-actions[bot] :: PR: #15055r2.6.0 by @github-actions[bot] :: PR: #15282r2.7.0 by @github-actions[bot] :: PR: #15351Clarify when to use TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD (15353) into r2.7.0 by @chtruong814 :: PR: #15359Fix macro accuracy when changing labels (15379) into r2.7.0 by @chtruong814 :: PR: #15386[voice agent] fix dependency for nemo26.02 (15380) into r2.7.0 by @chtruong814 :: PR: #15383Remove deprecated LLM, VLM, and diffusion tutorials (15357) into r2.7.0 by @chtruong814 :: PR: #15392fixes nemo tutorial for loading non registered classes (15398) into r2.7.0 by @chtruong814 :: PR: #15399default weights to false (15397) into r2.7.0 by @chtruong814 :: PR: #15401Fixing could not find ctc_segmentation. in CTC tutorial (15403) into r2.7.0 by @chtruong814 :: PR: #15404Adapt to use env variable for adapter mixin model loading (15406) into r2.7.0 by @chtruong814 :: PR: #15407</details>
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
</details>
<details><summary>Changelog</summary>
Updated tutorial on SDE, due to changes in colab and libraries (15137) into r2.6.0 by @chtruong814 :: PR: #15289unset weights_only=False (15312) into r2.6.0 by @chtruong814 :: PR: #15328Update Imports in Audio Notebook (15345) into r2.6.0 by @chtruong814 :: PR: #15346Clarify when to use TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD (15353) into r2.6.0 by @chtruong814 :: PR: #15358</details>
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Non-Speech NeMo 2.0 collections are deprecated and will be removed in a later release. Their functionality is available in the Megatron Bridge repo at…
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
remove ExportDeploy into r2.6.0 by @pablo-garay :: PR: #15053</details>
<details><summary>Changelog</summary>
remove ExportDeploy into r2.6.0 by @pablo-garay :: PR: #15053</details>
<details><summary>Changelog</summary>
2.6.0rc0.dev0 by @github-actions[bot] :: PR: #14512safe_import("modelopt") in llm collection by @kevalmorabia97 :: PR: #14656r2.3.0 by @github-actions[bot] :: PR: #14812v2.4.1 by @github-actions[bot] :: PR: #14828r2.6.0 by @github-actions[bot] :: PR: #14957Bump MCore, TE, Pytorch, and modelopt for 25.11 (14946) into r2.6.0 by @chtruong814 :: PR: #14976Update ctc-segmentation (14991) into r2.6.0 by @chtruong814 :: PR: #14998Pass timeout when running speech functional tests (15012) into r2.6.0 by @chtruong814 :: PR: #15013check asr models (14989) into r2.6.0 by @chtruong814 :: PR: #15002Enable EP in PTQ (15015) into r2.6.0 by @chtruong814 :: PR: #15026Update numba to numba-cuda and update cuda python bindings usage (15018) into r2.6.0 by @chtruong814 :: PR: #15024Add import guards for mcore lightning module (14970) into r2.6.0 by @chtruong814 :: PR: #14981fix loading of hyb ctc rnnt bpe models when using from pretrained (15042) into r2.6.0 by @chtruong814 :: PR: #15045fix: fix update-buildcache workflow after ED remove (15051) into r2.6.0 by @chtruong814 :: PR: #15052chore: update Lightning requirements version (15004) into r2.6.0 by @chtruong814 :: PR: #15049update notebook (15093) into r2.6.0 by @chtruong814 :: PR: #15094Fix: Obsolete Attribute [SDE] (15105) into r2.6.0 by @chtruong814 :: PR: #15106Upgrade NeMo ASR tutorials from Mozilla/CommonVoice to Google/FLEURS (15103) into r2.6.0 by @chtruong814 :: PR: #15107chore: Remove Automodel module (15044) into r2.6.0 by @chtruong814 :: PR: #15084Add deprecation notice to modules (15050) into r2.6.0 by @chtruong814 :: PR: #15110</details>
Nothing published for this version
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Update ctc-segmentation (14991) into r2.5.0 by @chtruong814 :: PR: #15020</details>
cp: Add import guards for mcore lightning module (#14970) into r2.5.0 by @chtruong814 :: PR: #14982
</details>
<details><summary>Changelog</summary>
Add import guards for mcore lightning module (#14970) into r2.5.0 by @chtruong814 :: PR: #14982</details>
<details><summary>Changelog</summary>
</details>
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Feat: Disk space management: for nemo install test (14822) into r2.5.0 by @chtruong814 :: PR: #14937Fix the load checkpointing issue -- onelogger callback gets called multiple time in some case. (14945) into r2.5.0 by @chtruong814 :: PR: #14948</details>
Deprecate Confidence Ensemble models
Collections:
Automodel and Export-Deploy functionality are available in their individual repositories respectively and deprecated in NeMo2
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
2.5.0rc0.dev0 by @github-actions[bot] :: PR: #13944r2.3.0 by @github-actions[bot] :: PR: #14160Dockerfile.speech by @artbataev :: PR: #14185r2.4.0 by @github-actions[bot] :: PR: #14334r2.5.0 by @github-actions[bot] :: PR: #14513Use hugginface_hub for downloading the FLUX checkpoint (14638) into r2.5.0 by @chtruong814 :: PR: #14640Fix function calling notebook (14643) into r2.5.0 by @chtruong814 :: PR: #14650remove service launch scripts (14647) into r2.5.0 by @chtruong814 :: PR: #14648Delete tutorials/llm/llama/biomedical-qa directory (14653) into r2.5.0 by @chtruong814 :: PR: #14654Remove PEFT scheme condition from recipe (14661) into r2.5.0 by @chtruong814 :: PR: #14662fixing kernel restarting when transcribing (14665) into r2.5.0 by @chtruong814 :: PR: #14672Fixing Sortformer training tutorial notebook (14680) into r2.5.0 by @chtruong814 :: PR: #14681Update get_tensor_shapes function whose signature was refactored (14594) into r2.5.0 by @chtruong814 :: PR: #14678Skip trt-llm and vllm install in install test (14663) into r2.5.0 by @chtruong814 :: PR: #14697Fix for \EncDecRNNTBPEModel transcribe() failed with TypeError\ (14698) into r2.5.0 by @chtruong814 :: PR: #14709Fix broken link in Reasoning-SFT.ipynb (14716) into r2.5.0 by @chtruong814 :: PR: #14717Fix deepseek export dtype (14307) into r2.5.0 by @chtruong814 :: PR: #14682remove env var (14739) into r2.5.0 by @chtruong814 :: PR: #14746safe_import("modelopt") in llm collection (#14656)' into 'r2.5.0' by @chtruong814 :: PR: #14771Update prune-distill notebooks to Qwen3 + simplify + mmlu eval (14785) into r2.5.0 by @chtruong814 :: PR: #14789Remove export-deploy, automodel, and eval tutorials (14790) into r2.5.0 by @chtruong814 :: PR: #14792ci: Automodel deprecation warning (14787) into r2.5.0 by @chtruong814 :: PR: #14791</details>
Prerelease: NVIDIA Neural Modules 2.5.0rc0 (2025-08-03)
Prerelease: NVIDIA Neural Modules 2.5.0rc0 (2025-08-03)
Update package_info.py by @ko3n1g :: PR: #14400
<details><summary>Changelog</summary>
Fix callbacks in DSV3 script (14350) into r2.4.0 by @chtruong814 :: PR: #14370Change Llama Embedding Tutorial to use SFT by default (14231) into r2.4.0 by @chtruong814 :: PR: #14303calculate_per_token_loss requirement for context parallel (#14065) (#14282) into r2.4.0 by @chtruong814 :: PR: #14448</details>
Cherry pick Moving export security fixes over here (14254) into r2.4.0 by @chtruong814 :: PR: #14261
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Adding more doc-strings to megatron_parallel.py #12767 by @ko3n1g :: PR: #13824</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
no-concurrency-group-on-main by @ko3n1g :: PR: #13375Run CICD label attached by @ko3n1g :: PR: #13430r2.3.0 by @github-actions[bot] :: PR: #13501__init__ to selective-triggering by @ko3n1g :: PR: #13577init-file-checker by @ko3n1g :: PR: #13684r2.3.1 by @github-actions[bot] :: PR: #13719PYTORCH_CUDA_ALLOC_CONF env var by @malay-nagda :: PR: #13837is_multistorageclient_url by @shunjiad :: PR: #13885r2.4.0 by @github-actions[bot] :: PR: #13945Use jiwer less than 4.0.0 (13997) into r2.4.0 by @ko3n1g :: PR: #13998Remove container license reference (14010) into r2.4.0 by @ko3n1g :: PR: #14017bf16 grads for bf16 jobs (14016) into r2.4.0 by @ko3n1g :: PR: #14020Remove nemo1 stable diffusion test (14018) into r2.4.0 by @ko3n1g :: PR: #140192.4.0rc1.dev0 by @github-actions[bot] :: PR: #14047Fix Loading Custom Quantization Config (13934) into r2.4.0 by @ko3n1g :: PR: #13950[automodel] fix sft notebook (14002) into r2.4.0 by @ko3n1g :: PR: #14003Use average reduction in FSDP grad reduce-scatter when grad dtype is … (13981) into r2.4.0 by @ko3n1g :: PR: #14004GPU memory logging update (13982) into r2.4.0 by @ko3n1g :: PR: #14021Remove kaldiio (14006) into r2.4.0 by @ko3n1g :: PR: #14032Set L2_NeMo_2_Flux_Import_Test to be optional (14056) into r2.4.0 by @ko3n1g :: PR: #14058Bump protobuf to 5.29.5 (14045) into r2.4.0 by @ko3n1g :: PR: #14060Detect hardware before enabling DeepEP (14022) into r2.4.0 by @ko3n1g :: PR: #140682.4.0rc2.dev0 by @github-actions[bot] :: PR: #14115Fix SFT Dataset Bug (13918) into r2.4.0 by @ko3n1g :: PR: #14074Align adapter shape with base linear output shape (14009) into r2.4.0 by @ko3n1g :: PR: #14083[MoE] Update the fp8 precision interface for llama4 and qwen3 (14094) into r2.4.0 by @ko3n1g :: PR: #14104[Llama4] Tokenizer naming update (14114) into r2.4.0 by @ko3n1g :: PR: #14123Bump to pytorch 25.05 container along with TE update (13899) into r2.4.0 by @ko3n1g :: PR: #14145Perf scripts updates (14005) into r2.4.0 by @ko3n1g :: PR: #14129Remove unstructured (14070) into r2.4.0 by @ko3n1g :: PR: #141472.4.0rc3.dev0 by @github-actions[bot] :: PR: #14165Add checkpoint info for NIM Embedding Expor Tutorial (14177) into r2.4.0 by @ko3n1g :: PR: #14178Fix dsv3 script (14007) into r2.4.0 by @ko3n1g :: PR: #14182405b perf script updates (14176) into r2.4.0 by @chtruong814 :: PR: #14195Fix nemotronh flops calculator (14161) into r2.4.0 by @chtruong814 :: PR: #14202Add option to disable gloo process groups (#14156) into r2.4.0 by @chtruong814 :: PR: #14220Remove g2p_en (14204) into r2.4.0 by @chtruong814 :: PR: #14212diffusion mock data null args (14173) into r2.4.0 by @chtruong814 :: PR: #14217perf-scripts: Change b200 config to EP8 (14207) into r2.4.0 by @chtruong814 :: PR: #14223Change RerankerSpecter Dataset question key (14200) into r2.4.0 by @chtruong814 :: PR: #14224Fix the forward when final_loss_mask is not present (14201) into r2.4.0 by @chtruong814 :: PR: #14225Fix Llama Nemotron Nano Importer (14222) into r2.4.0 by @chtruong814 :: PR: #14226[automodel] fix loss_mask pad token (14150) into r2.4.0 by @chtruong814 :: PR: #14227Moving export security fixes over here (14254) into r2.4.0 by @chtruong814 :: PR: #14261Confidence fix for tutorial (14250) into r2.4.0 by @chtruong814 :: PR: #14266added new models to documentation (14264) into r2.4.0 by @chtruong814 :: PR: #14278FIx Flux & Flux_Controlnet initialization issue (#14263) into r2.4.0 by @chtruong814 :: PR: #14273update ffmpeg install (14237) into r2.4.0 by @chtruong814 :: PR: #14279</details>
Prerelease: NVIDIA Neural Modules 2.4.0rc2 (2025-07-09)
Prerelease: NVIDIA Neural Modules 2.4.0rc2 (2025-07-09)
Prerelease: NVIDIA Neural Modules 2.4.0rc1 (2025-07-02)
Prerelease: NVIDIA Neural Modules 2.4.0rc1 (2025-07-02)
Prerelease: NVIDIA Neural Modules 2.4.0rc0 (2025-06-27)
Prerelease: NVIDIA Neural Modules 2.4.0rc0 (2025-06-27)
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit , for acknowledgement please reach out…
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit https://www.nvidia.com/en-us/security/,…
This release addresses known security issues. For the latest NVIDIA Vulnerability Disclosure Information visit https://www.nvidia.com/en-us/security/, for acknowledgement please reach out to the NVIDIA PSIRT team at PSIRT@nvidia.com
Updated vLLMExporter to use vLLM V1 to address a security vulnerability.
</details>
<details><summary>Changelog</summary>
Update vLLMExporter to use vLLM V1 (#13498) into r2.3.0 by @chtruong814 :: PR: #13631</details>
<details><summary>Changelog</summary>
Use explicitly cached canary-1b-flash in CI tests (13237) into r2.3.0 by @ko3n1g :: PR: #13508[automodel] bump liger-kernel to 0.5.8 + fallback (13260) into r2.3.0 by @ko3n1g :: PR: #13308Add recipe and ci scripts for qwen2vl to r2.3.0 by @romanbrickie :: PR: #13336Fix skipme handling (13244) into r2.3.0 by @ko3n1g :: PR: #13376Allow fp8 param gather when using FSDP (13267) into r2.3.0 by @ko3n1g :: PR: #13383Handle boolean args for performance scripts and log received config (13291) into r2.3.0 by @ko3n1g :: PR: #13416new perf configs (13110) into r2.3.0 by @ko3n1g :: PR: #13431Adding additional unit tests for the deploy module (13411) into r2.3.0 by @ko3n1g :: PR: #13449Adding more export tests (13410) into r2.3.0 by @ko3n1g :: PR: #13450[automodel] add FirstRankPerNode (13373) into r2.3.0 by @ko3n1g :: PR: #13559[automodel] deprecate global_batch_size dataset argument (13137) into r2.3.0 by @ko3n1g :: PR: #13560[automodel] fallback FP8 + LCE -> FP8 + CE (#13349) into r2.3.0 by @chtruong814 :: PR: #13561[automodel] add find_unused_parameters=True for DDP (13366) into r2.3.0 by @ko3n1g :: PR: #13601Add CI test for local checkpointing (#13012) into r2.3.0 by @ananthsub :: PR: #13472[automodel] fix --mbs/gbs dtype and chat-template (13598) into r2.3.0 by @akoumpa :: PR: #13613Update t5.py (#13082) to r2.3.0 and bump mcore to f98b1a0 by @chtruong814 :: PR: #13642</details>
ONNX and TensorRT Export for NIM Embedding Container
<details><summary>Changelog</summary>
FastNGramLM -> NGramGPULanguageModel by @artbataev :: PR: #12755</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
FastNGramLM -> NGramGPULanguageModel by @artbataev :: PR: #12755</details>
<details><summary>Changelog</summary>
FastNGramLM -> NGramGPULanguageModel by @artbataev :: PR: #12755</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
r2.2.0 by @github-actions[bot] :: PR: #12585--branch by @ko3n1g :: PR: #12809r2.2.1 by @github-actions[bot] :: PR: #12818Dockerfile.speech by @artbataev :: PR: #12859r2.3.0 by @github-actions[bot] :: PR: #129192.3.0rc3.dev0 by @github-actions[bot] :: PR: #12921[automodel] Add linear ce loss support (12825) into r2.3.0 by @ko3n1g :: PR: #12922DeepSeek V3 Multi Token Prediction (12550) into r2.3.0 by @ko3n1g :: PR: #12928Set L2_NeMo_2_EVAL test to be optional (12949) into r2.3.0 by @ko3n1g :: PR: #12951GB200 LLM performance scripts tuning (12791) into r2.3.0 by @ko3n1g :: PR: #12923Allow configuration of PP communication backend to UCC in nemo2 (11755) into r2.3.0 by @ko3n1g :: PR: #12946guard bitsandbytes based on cuda availability (12937) into r2.3.0 by @ko3n1g :: PR: #12958Hugging Face model deployment support (12628) into r2.3.0 by @ko3n1g :: PR: #12962fix macro-acc for pair-audio eval (12908) into r2.3.0 by @ko3n1g :: PR: #12963Add energon dataset support for Qwen2VL (12831) into r2.3.0 by @ko3n1g :: PR: #12966Make TETransformerLayerAutocast Support Cuda Graph (12075) into r2.3.0 by @ko3n1g :: PR: #12967Use nvidia-lm-eval for evaluation (12902) into r2.3.0 by @ko3n1g :: PR: #12971[NeMo 2.0] Interface for using MXFP8 and FP8 current scaling recipes (12503) into r2.3.0 by @ko3n1g :: PR: #12974Fix trtllm and lightning conflict (12943) into r2.3.0 by @ko3n1g :: PR: #12981Update v3 finetuning recipe (12950) and Specify PP first/last in strategy (12992) into r2.3.0 by @ko3n1g :: PR: #12984Resolve an issue in custom megatron FSDP config setting (12948) into r2.3.0 by @ko3n1g :: PR: #12987Remove getattr_proxy to avoid problematic edge cases (12176) into r2.3.0 by @ko3n1g :: PR: #12990Enable async requests for in-fw deployment with OAI compatible server (12980) into r2.3.0 by @ko3n1g :: PR: #12994initialize model with metadata (12496) into r2.3.0 by @ko3n1g :: PR: #12997Bugfix for logits support for hf deployment (12965) into r2.3.0 by @ko3n1g :: PR: #13001Update nvidia-resiliency-ext to be >= 0.3.0 (12925) into r2.3.0 by @ko3n1g :: PR: #130002.3.0rc4.dev0 by @github-actions[bot] :: PR: #13041Alit/nemotron h (12942) into r2.3.0 by @ko3n1g :: PR: #13007[Automodel] Add TP/SP support with default llama-like sharding plan (12796) into r2.3.0 by @ko3n1g :: PR: #13017Add initial docs broken link check (12977) into r2.3.0 by @ko3n1g :: PR: #13045Fix MoE Init to not use Bias in test_strategy_lib.py (13009) into r2.3.0 by @ko3n1g :: PR: #13014cleaner tflops log name (13005) into r2.3.0 by @ko3n1g :: PR: #13024Improve t5 test coverage (12803) into r2.3.0 by @ko3n1g :: PR: #13025 put the warning on the right place (12909) into r2.3.0 by @ko3n1g :: PR: #13035Temporary disable CUDA graphs in DDP mode for transducer decoding (12907) into r2.3.0 by @ko3n1g :: PR: #13036[automodel] peft fix vlm (13010) into r2.3.0 by @ko3n1g :: PR: #13037Only run the docs link check on the container (13068) into r2.3.0 by @ko3n1g :: PR: #13070Add fp8 recipe option to perf script (13032) into r2.3.0 by @ko3n1g :: PR: #13055Unified ptq export (12786) into r2.3.0 by @ko3n1g :: PR: #13062Fix VP list index out of range from Custom FSDP (13021) into r2.3.0 by @ko3n1g :: PR: #13077Add logging to cancel out PTL's warning about dataloader not being resumable (13072) into r2.3.0 by @ko3n1g :: PR: #13100Fix long sequence generation after new arg introduced in mcore engine (13049) into r2.3.0 by @ko3n1g :: PR: #13104Support Mamba models quantization (12631) into r2.3.0 by @ko3n1g :: PR: #13105Add track_io to user buffer configs (13071) into r2.3.0 by @ko3n1g :: PR: #13111Add fine-tuning dataset function for FineWeb-Edu and update automodel… (13027) into r2.3.0 by @ko3n1g :: PR: #13118Re-add sox to asr requirements (13092) into r2.3.0 by @ko3n1g :: PR: #13120Update Mllama cross attn signature to match update MCore (13048) into r2.3.0 by @ko3n1g :: PR: #13122Fix Exporter for baichuan and chatglm (13095) into r2.3.0 by @ko3n1g :: PR: #131262.3.0rc5.dev0 by @github-actions[bot] :: PR: #13146Guard decord and triton import (12861) into r2.3.0 by @ko3n1g :: PR: #13132Bump TE version and apply patch (13087) into r2.3.0 by @ko3n1g :: PR: #13139Update Llama-Minitron pruning-distillation notebooks from NeMo1 to NeMo2 + NeMoRun (12968) into r2.3.0 by @ko3n1g :: PR: #13141Export and Deploy Tests (13076) into r2.3.0 by @ko3n1g :: PR: #13150ub fp8 h100 fixes (13131) into r2.3.0 by @ko3n1g :: PR: #13153Fix Transducer Decoding with CUDA Graphs in DDP with Mixed Precision (12938) into r2.3.0 by @ko3n1g :: PR: #13154build: Pin modelopt (13029) into r2.3.0 by @chtruong814 :: PR: #13170add fixes for nemotron-h (13073) into r2.3.0 by @JRD971000 :: PR: #13165Add Llama Nemotron Super/Ultra models (13044) into r2.3.0 by @ko3n1g :: PR: #13212Add Blockwise FP8 to PTQ & EP to modelopt resume (12670) into r2.3.0 by @ko3n1g :: PR: #13239[OAI Serving] Validate greedy generation args (redo) (13216) into r2.3.0 by @ko3n1g :: PR: #13242drop sample_alpha in speechlm (13208) into r2.3.0 by @ko3n1g :: PR: #13246[Eval bugfix] Move global eval-related imports inside the evaluate function (13166) into r2.3.0 by @ko3n1g :: PR: #13249[Eval bugfix] Change default val of parallel_requests in eval script (13247) into r2.3.0 by @ko3n1g :: PR: #13253Add tutorial for evaluation with Evals Factory (13259) into r2.3.0 by @ko3n1g :: PR: #13271Fix default token durations (13168) into r2.3.0 by @ko3n1g :: PR: #13261[Evaluation] Add support for nvidia-lm-eval==25.04 (13230) into r2.3.0 by @ko3n1g :: PR: #13274[bug fix] set inference max seq len in inference context (13245) into r2.3.0 by @ko3n1g :: PR: #13276More export and deploy unit tests (13178) into r2.3.0 by @ko3n1g :: PR: #13283Reopen 13040 (13199) into r2.3.0 by @ko3n1g :: PR: #13303Fix nemo1's neva notebook (13218) into r2.3.0 by @ko3n1g :: PR: #13312build: various bumps (13285) into r2.3.0 by @ko3n1g :: PR: #13313ci: Increase cache pool into r2.3.0 by @chtruong814 :: PR: #13317update num nodes in deepseek v3 finetune recipe (13314) into r2.3.0 by @ko3n1g :: PR: #13316Fix neva notebook (13334) into r2.3.0 by @ko3n1g :: PR: #13335Add Llama4 Scout and Maverick Support (#12898) by @ko3n1g :: PR: #13331Fix handling Llama Embedding dimensions param and prompt type in the ONNX export tutorial (13262) into r2.3.0 by @ko3n1g :: PR: #13326Fix transformer offline for CI/CD llama4 tests (#13339) to r2.3.0 by @chtruong814 :: PR: #13340vLLM==0.8.5 update (13350) into r2.3.0 by @ko3n1g :: PR: #13354Add llama4 training recipe (12952) into r2.3.0 by @ko3n1g :: PR: #13386</details>
Prerelease: NVIDIA Neural Modules 2.3.0rc4 (2025-04-21)
Prerelease: NVIDIA Neural Modules 2.3.0rc4 (2025-04-21)
Prerelease: NVIDIA Neural Modules 2.3.0rc3 (2025-04-15)
Prerelease: NVIDIA Neural Modules 2.3.0rc3 (2025-04-15)
Prerelease: NVIDIA Neural Modules 2.3.0rc2 (2025-04-07)
Prerelease: NVIDIA Neural Modules 2.3.0rc2 (2025-04-07)
Fix MoE based models training instability.
</details>
<details><summary>Changelog</summary>
Fix exporter for llama models with shared embed and output layers (12545) into r2.2.0 by @ko3n1g :: PR: #12608Fix TP for LoRA adapter on linear_fc1 (12519) into r2.2.0 by @ko3n1g :: PR: #12607</details>
Remove deprecated tests/infer_data_path.py by @janekl :: PR: #11997
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
sox to SDE by @ko3n1g :: PR: #11882</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
r2.1.0 by @github-actions[bot] :: PR: #11745MCORE_TAG=4dc8977... (2025-01-07) by @ko3n1g :: PR: #11768triton by @ko3n1g :: PR: #119382.2.0rc1 by @github-actions[bot] :: PR: #120232.2.0rc2.dev0 by @github-actions[bot] :: PR: #12040build: Force re-install VCS dependencies (12155) into r2.2.0 by @ko3n1g :: PR: #12191Add function calling SFT NeMo2.0 tutorial (11868) into r2.2.0 by @ko3n1g :: PR: #12180Update TTS code to remove calls to deprecated functions (12153) into r2.2.0 by @ko3n1g :: PR: #12201Fix multi-GPU in-framework deployment (12090) into r2.2.0 by @ko3n1g :: PR: #12172disable moe logging to avoid deepseek hang (12168) into r2.2.0 by @ko3n1g :: PR: #12192build: Pin down transformers (12229) into r2.2.0 by @ko3n1g :: PR: #12230Fix loading extra states from torch tensor (12185) into r2.2.0 by @ko3n1g :: PR: #12226nemo-automodel checkpoint-io refactor (12070) into r2.2.0 by @ko3n1g :: PR: #12234Set L2_Speech_Batch_Size_OOMptimizer_Canary to be optional (12299) into r2.2.0 by @ko3n1g :: PR: #12300build: Exclude tensorstore 0.1.72 (12317) into r2.2.0 by @ko3n1g :: PR: #12318Fix the local path in Sortformer diarizer training tutorial (12135) into r2.2.0 by @ko3n1g :: PR: #12316Add eval requirement to setup.py (12152) into r2.2.0 by @ko3n1g :: PR: #12277Add modelopt to requirements_nlp.txt (12261) into r2.2.0 by @ko3n1g :: PR: #12278Energon ckpt multimodal (12245) into r2.2.0 by @ko3n1g :: PR: #12307[nemo1] Fix Mamba/Bert loading from checkpoint after TE extra states were introduced (12275) into r2.2.0 by @ko3n1g :: PR: #12314fix masked loss calculation (12255) into r2.2.0 by @ko3n1g :: PR: #12286build: Bump mcore (12320) into r2.2.0 by @ko3n1g :: PR: #12328[automodel] re-enable FSDP2 tests (12325) into r2.2.0 by @ko3n1g :: PR: #12331[automodel] fix loss reporting (12303) into r2.2.0 by @ko3n1g :: PR: #12334[automodel] remove fix_progress_bar from fsdp2 strategy (12339) into r2.2.0 by @ko3n1g :: PR: #12347Fix NeMo1 Bert Embedding Dataset args (12341) into r2.2.0 by @ko3n1g :: PR: #12349Fix NeMo1 sequence_len_offset in Bert fwd (12350) into r2.2.0 by @ko3n1g :: PR: #12359Add nemo-run recipe for evaluation (12301) into r2.2.0 by @ko3n1g :: PR: #12352Add DeepSeek-R1 Distillation NeMo 2.0 tutorial (12187) into r2.2.0 by @ko3n1g :: PR: #123552.2.0rc4.dev0 by @github-actions[bot] :: PR: #12363[automodel] add lr scheduler (12351) into r2.2.0 by @ko3n1g :: PR: #12361[automodel] add distributed data sampler (12326) into r2.2.0 by @ko3n1g :: PR: #12373[NeVA] Fix for CP+THD (12366) into r2.2.0 by @ko3n1g :: PR: #12375Ignore attribute error when serializing mcore specs (12353) into r2.2.0 by @ko3n1g :: PR: #12383Avoid init_ddp for inference (12011) into r2.2.0 by @ko3n1g :: PR: #12385[docs] fix notebook render (12374) into r2.2.0 by @ko3n1g :: PR: #12394Neva finetune scripts and PP fix (12387) into r2.2.0 by @ko3n1g :: PR: #12397[automodel] update runner tags for notebooks (12428) into r2.2.0 by @ko3n1g :: PR: #12431[automodel] update examples (12411) into r2.2.0 by @ko3n1g :: PR: #12432Evaluation docs (12348) into r2.2.0 by @ko3n1g :: PR: #12460Update prompt format (12452) into r2.2.0 by @ko3n1g :: PR: #12455Fixing a wrong Sortformer Tutorial Notebook path. (12479) into r2.2.0 by @ko3n1g :: PR: #12480added a needed checks and changes for bugfix (12400) into r2.2.0 by @Ssofja :: PR: #12447[automodel] fix loss/tps reporting across ranks (12389) into r2.2.0 by @ko3n1g :: PR: #12413enable fsdp flag for FSDP2Strategy (12392) into r2.2.0 by @ko3n1g :: PR: #12429Fix lita notebook issue (12474) into r2.2.0 by @ko3n1g :: PR: #12476 Changed the argument types passed to metrics calculation functions (12500) into r2.2.0 by @ko3n1g :: PR: #12502added needed fixes (12495) into r2.2.0 by @ko3n1g :: PR: #12509update transformers version requirements (12475) into r2.2.0 by @ko3n1g :: PR: #12507[checkpoint] Log timings for checkpoint IO save and load (11972) into r2.2.0 by @ko3n1g :: PR: #12520few checkings needed because of the change of asr models output (12499) into r2.2.0 by @ko3n1g :: PR: #12513Remove _attn_implementationinLlamaBidirectionalModel constructor (12364) into r2.2.0 by @ko3n1g :: PR: #12525Configure FSDP to keep module params (12074) into r2.2.0 by @ko3n1g :: PR: #12524[automodel] docs (11942) into r2.2.0 by @ko3n1g :: PR: #12530[automodel] update examples' comments (12518) and [automodel] Move PEFT to configure_model (#12491) into r2.2.0 by @ko3n1g :: PR: #12529update readme to include latest pytorch version (12539) into r2.2.0 by @ko3n1g :: PR: #12577</details>
Prerelease: NVIDIA Neural Modules 2.2.0rc3 (2025-02-25)
Prerelease: NVIDIA Neural Modules 2.2.0rc3 (2025-02-25)
Prerelease: NVIDIA Neural Modules 2.2.0rc2 (2025-02-17)
Prerelease: NVIDIA Neural Modules 2.2.0rc2 (2025-02-17)
Prerelease: NVIDIA Neural Modules 2.2.0rc1 (2025-02-04)
Prerelease: NVIDIA Neural Modules 2.2.0rc1 (2025-02-04)
Prerelease: NVIDIA Neural Modules 2.2.0rc0 (2025-02-02)
Prerelease: NVIDIA Neural Modules 2.2.0rc0 (2025-02-02)
Added deprecation notice by @Ssofja :: PR: #11133
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
BaseMegatronSampler for compatibility with PTL's _BatchProgress by @ashors1 :: PR: #11016ckpt_to_weights_subdir from MegatronCheckpointIO by @ashors1 :: PR: #10897attention_bias argument in transformer block and transformer layer modules, addressing change in MCore by @yaoyu-33 :: PR: #11289MCORE_TAG=67a50f2... (2024-11-28) by @ko3n1g :: PR: #11427L2_Megatron_LM_To_NeMo_Conversion by @ko3n1g :: PR: #11484</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Dockerfile.ci (2024-10-14) by @ko3n1g :: PR: #10871Dockerfile.ci (2024-10-17) by @ko3n1g :: PR: #10919Dockerfile.ci (2024-10-21) by @ko3n1g :: PR: #10965Dockerfile.ci (2024-10-22) by @ko3n1g :: PR: #10979Dockerfile.ci (2024-10-23) by @ko3n1g :: PR: #11001Dockerfile.ci (2024-10-27) by @ko3n1g :: PR: #11051Dockerfile.ci (2024-10-28) by @ko3n1g :: PR: #11054Dockerfile.ci (2024-10-30) by @ko3n1g :: PR: #11092Dockerfile.ci (2024-11-05) by @ko3n1g :: PR: #11159Dockerfile.ci (2024-11-06) by @ko3n1g :: PR: #11174Dockerfile.ci (2024-11-07) by @ko3n1g :: PR: #11196Dockerfile.ci (2024-11-08) by @ko3n1g :: PR: #11222Dockerfile.ci (2024-11-11) by @ko3n1g :: PR: #11247Dockerfile.ci (2024-11-12) by @ko3n1g :: PR: #11254bump mcore to templates by @ko3n1g :: PR: #11229MCORE_TAG=aded519... (2024-11-12) by @ko3n1g :: PR: #11260pull_request_target by @ko3n1g :: PR: #11263workflow_event by @ko3n1g :: PR: #11322MCORE_TAG=bd677bf... (2024-12-06) by @ko3n1g :: PR: #11492r2.1.0 by @github-actions[bot] :: PR: #11556Add fix docstring for speech commands (11638) into r2.1.0 by @ko3n1g :: PR: #11639Add fix docstring for VAD (11659) into r2.1.0 by @ko3n1g :: PR: #11660Downgrading the 'datasets' package from 3.0.0 to 2.21.0 for Multilang_ASR.ipynb and ASR_CTC_Language_Finetuning.ipynb (11675) into r2.1.0 by @ko3n1g :: PR: #11677Rename multimodal data module - EnergonMultiModalDataModule (11654) into r2.1.0 by @ko3n1g :: PR: #11685r2.1.0rc2 by @ko3n1g :: PR: #11693</details>
Prerelease: NVIDIA Neural Modules 2.1.0rc2 (2024-12-21)
Prerelease: NVIDIA Neural Modules 2.1.0rc2 (2024-12-21)
Prerelease: NVIDIA Neural Modules 2.1.0rc1 (2024-12-20)
Prerelease: NVIDIA Neural Modules 2.1.0rc1 (2024-12-20)
Nothing published for this version
Added deprecation notice by @Ssofja :: PR: #11133
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Bump Dockerfile.ci (2024-09-09) by @ko3n1g :: PR: #10423
MCORE interface for TP-only FP8 AMAX reduction by @erhoo82 :: PR: #10437
Remove Apex dependency if not using MixedFusedLayerNorm by @cuichenx :: PR: #10468
Add missing import guards for causal_conv1d and mamba_ssm dependencies by @janekl :: PR: #10429
Update doc for fp8 trt-llm export by @Laplasjan107 :: PR: #10444
set mock in GPTDatasetConfig by @akoumpa :: PR: #10435
Remove running validating after finetuning by @huvunvidia :: PR: #10560
Extending modelopt spec for TEDotProductAttention by @janekl :: PR: #10523
Fix mb_calculator import in lora tutorial by @BoxiangW :: PR: #10624
.nemo conversion bug fix by @dimapihtar :: PR: #10598
Updating modelopt spec for Mixtral by @janekl :: PR: #10660
Require setuptools>=70 and update deprecated api by @thomasdhc :: PR: #10659
Akoumparouli/fix get tokenizer list by @akoumpa :: PR: #10596
[McoreDistOptim] fix the naming to match apex.dist by @gdengk :: PR: #10707
[fix] Ensures disabling exp_manager with exp_manager=null does not error by @terrykong :: PR: #10651
[feat] Update get_model_parallel_src_rank to support tp-pp-dp ordering by @terrykong :: PR: #10652
feat: Migrate GPTSession refit path in Nemo export to ModelRunner for Aligner by @terrykong :: PR: #10654
[MCoreDistOptim] Add assertions for McoreDistOptim and fix fp8 arg specs by @gdengk :: PR: #10748
Fix for crashes with tensorboard_logger=false and VP + LoRA by @vysarge :: PR: #10792
Adding init_model_parallel to FabricMegatronStrategy by @marcromeyn :: PR: #10733
Moving steps to MegatronParallel to improve UX for Fabric by @marcromeyn :: PR: #10732
Adding setup_megatron_optimizer to FabricMegatronStrategy by @marcromeyn :: PR: #10833
Make FabricMegatronMixedPrecision match MegatronMixedPrecision by @marcromeyn :: PR: #10835
Fix VPP bug in MegatronStep by @marcromeyn :: PR: #10847
Expose drop_last in MegatronDataSampler by @farhadrgh :: PR: #10837
Move collectiob.nlp imports inline for t5 by @marcromeyn :: PR: #10877
Use a context-manager when opening files by @akoumpa :: PR: #10895
Packed sequence bug fixes by @cuichenx :: PR: #10898
ckpt convert bug fixes by @dimapihtar :: PR: #10878
remove deprecated ci tests by @dimapihtar :: PR: #10922
Adithyare/oai chat completion by @arendu :: PR: #10785
Update T5 tokenizer (adding additional tokens to tokenizer config) by @huvunvidia :: PR: #10972
Add support and recipes for HF models via AutoModelForCausalLM by @akoumpa :: PR: #10962
BaseMegatronSampler for compatibility with PTL'''s _BatchProgress by @ashors1 :: PR: #11016ckpt_to_weights_subdir from MegatronCheckpointIO by @ashors1 :: PR: #10897</details>
deprecate NeMo NLP tutorial by @dimapihtar :: PR: #9864
<details><summary>Changelog</summary>
audio for transcription by @titu1994 :: PR: #9201<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
scheduler method by @sararb :: PR: #9609PYTORCH_VERSION variable by @artbataev :: PR: #9736val_loss by @ashors1 :: PR: #9814save_dir by @ashors1 :: PR: #9954</details>
Cleanup deprecated files and temporary changes by @cuichenx :: PR: #9088
Megatron Core RETRO
Pretraining, conversion, evaluation, SFT, and PEFT for:
Embedding Models Fine Tuning
BERT models
Video capabilities with NeVa
Distributed Checkpointing
Multimodal LLM (LLAVA/NeVA)
<details><summary>Changelog</summary>
audio for transcription (#9201) by @titu1994 :: PR: #9235</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
shard_id parsing in LazyNemoTarredIterator, enables AIS dataloading by @pzelasko :: PR: #9077</details>
[NLP] Remove replace_sampler_ddp (deprecated in Trainer) by @janekl :: PR: #7981
Announcement - https://nvidia.github.io/NeMo/blogs/2024/2024-02-canary/
Previously, the RNNT metric was stateful while the CTC one was not (r1.22.0, r1.23.0)
Therefore this calculation in the RNNT joint for fused operation worked properly. However with the unification of metrics in r1.23.0, a bug was introduced where only the last sub-batch of metrics calculates the scores and does not accumulate. This is patched via https://github.com/NVIDIA/NeMo/pull/8587 and will be fixed in the next release.
Workaround: Explicitly disable fused batch size during inference using the following command
from omegaconf import open_dict
model = ...
decoding_cfg = model.cfg.decoding
with open_dict(decoding_cfg):
decoding_cfg.fused_batch_size = -1
model.change_decoding_strategy(decoding_cfg)
Note: This bug does not affect scores calculated via model.transcribe() (since it does not calculate metrics during inference, just text), or using the transcribe_speech.py or speech_to_text_eval.py in examples/asr.
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:24.01.speech
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Announcement - https://nvidia.github.io/NeMo/blogs/2024/2024-01-parakeet/
Announcement - https://nvidia.github.io/NeMo/blogs/2024/2024-01-parakeet/
Announcement - https://nvidia.github.io/NeMo/blogs/2024/2024-01-parakeet-tdt/
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.10
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
deprecation warning by @arendu :: PR: #7193
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.08
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
[TTS] corrected misleading deprecation warnings. by @XuesongYang :: PR: #6702
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.06
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Remove deprecated arg compute_on_step. See #6979.
This release is a small patch to fix torchmetrics.
compute_on_step. See #6979.Sharded Manifests for Tarred Datasets #6395
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.04
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
For the complete release note, please see NeMo 1.18.0 Release Notes
For the complete release note, please see NeMo 1.18.0 Release Notes
This patch release fixes a major bug in ASR Bucketing datasets that was introduced in r1.17.0 in PR https://github.com/NVIDIA/NeMo/pull/6191. Due to this bug, while each bucket is randomly shuffled before selection on each rank, only a single bucket would loop infinitely - without continuing onto subsequent buckets.
Effect: Significantly worse WER would be obtained since not all buckets would be used.
This has been patched and should work correctly in 1.18.1 onwards.
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.03
GPT-2B-001, trained on 1.1T tokens with 4K sequence length.
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
[TTS] deprecate AudioToCharWithPriorAndPitchDataset. by @XuesongYang :: PR: #5959
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.02
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Multi-channel dereverberation algorithm
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:23.01
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Greedy timestamp decoding with inference script
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.12
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Hybrid CTC + Transducer loss ASR #5364
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.11
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
[TTS] deprecate TextToWaveform base class. by @XuesongYang :: PR: #5205
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.09
<details><summary>Issues</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.08
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
Fix logger reference by @SeanNaren :: PR: #4786
Fix error with class method reference in msdd by @SeanNaren :: PR: #4865
Add sync for logging calls to ensure aggregation across devices by @SeanNaren :: PR: #4876
Fix saving the last checkpoint when using val check interval by @SeanNaren :: PR: #4905
Add support for skipping validation on resume + extend saving last ckpt test by @SeanNaren :: PR: #4922
Move trainer calls for ssl models to training and validation steps only by @sam1373 :: PR: #4685
Change Num Partitions size expansion fix by @aklife97 :: PR: #4719
upgrade to PTL 1.7 by @nithinraok :: PR: #4672
Fixing outputs of infer() and use of NeMo length regulator helper by @borisfom :: PR: #4724
bug fix: enable async grad reduction when DP > 1 by @erhoo82 :: PR: #4740
Add LayerNorm1P, weight decay for LN and unscaled initialization by @mikolajblaz :: PR: #4743
Data Simulator by @chooper1 :: PR: #4686
jenkins data simulator fix by @nithinraok :: PR: #4751
Mutiscale Diarization Decoder (MSDD) model and module files by @tango4j :: PR: #4650
Fix logging in gradient clipping with PTL 1.7.2 by @MaximumEntropy :: PR: #4769
Fix checkpoint restoring by @nithinraok :: PR: #4777
avoid data clipping after convolution with rir samples by @nithinraok :: PR: #4806
Fixed in_features dim if bidirectional is True by @farisalasmary :: PR: #4588
Fix float/integer type error in WER.update() by @fujimotos :: PR: #4816
[Speech Data Explorer] An option to explicitly specify the base dir by @anteju :: PR: #4678
adding instancenorm as an option for conv normalization by @bmwshop :: PR: #4827
Fix small spelling mistakes by @SeanNaren :: PR: #4839
[Tutorials] Fix matplotlib version and directory name in Multispeaker_Simulator by @anteju :: PR: #4804
Update diarization folder structure by @tango4j :: PR: #4823
Missing types in clustering by @SeanNaren :: PR: #4858
add new models by @Jorjeous :: PR: #4852
Fix decoding for T5 models with RPE by @MaximumEntropy :: PR: #4847
Update Speaker Diarization notebooks with unknown oracle_num_speakers by @fayejf :: PR: #4861
Fix mha bug by @yzhang123 :: PR: #4859
Updates to adapter training by @arendu :: PR: #4842
Changes to MSDD code after review, fix test log call by @SeanNaren :: PR: #4881
Fixed output of BERT to be [batch x seq x hidden] by @michalivne :: PR: #4887
Add AMI dataset script by @SeanNaren :: PR: #4864
Update label_models.py by @stevehuang52 :: PR: #4891
Update tutorials.rst for question answering by @Zhilin123 :: PR: #4895
removed unused imports for all domains. by @XuesongYang :: PR: #4901
Fix ptl_load_state not providing cls by @MaximumEntropy :: PR: #4914
Remove unused cv collection by @okuchaiev :: PR: #4907
Add mixed-representation config to PhonemizerTokenizer by @rlangman :: PR: #4904
Fix implicit bug in _AudioLabelDataset by @stevehuang52 :: PR: #4923
Fix and refactor label models by @fayejf :: PR: #4913
Sparrowhawk deployment fix by @ekmb :: PR: #4928
Upgrade to NGC PyTorch 22.08 Container by @ericharper :: PR: #4929
Fixes for Cherry Picked PRs by @titu1994 :: PR: #4962
Fix cherry pick workflow by @ericharper :: PR: #4964
check for active conda environment by @nithinraok :: PR: #4970
fix label models restoring issue from weighted cross entropy by @nithinraok :: PR: #4968
Add simple pre-commit file (#4983) by @SeanNaren :: PR: #4995
Fix bug in Squeezeformer Conv block by @titu1994 :: PR: #5011
Fix bugs by @Zhilin123 :: PR: #5036
Add black to pre-commit (#5027) by @SeanNaren :: PR: #5045
Fix bug in question answering tutorial by @Zhilin123 :: PR: #5049
Missing fixes from r1.11.0 to T5 finetuning eval by @MaximumEntropy :: PR: #5054
P&C docs by @jubick1337 :: PR: #5068
probabilites -> probabilities by @nithinraok :: PR: #5078
Notebook bug fixes by @vadam5 :: PR: #5084
update strategy in notebook from ddp_fork to dp by @Zhilin123 :: PR: #5088
Fix Unhashable type list for Numba Cuda spec augment kernel by @titu1994 :: PR: #5093
Remove numba import by @titu1994 :: PR: #5095
T5 prompt learning fixes missing from r.11.0 merge by @MaximumEntropy :: PR: #5075
T5 Decoding with PP > 2 fix by @MaximumEntropy :: PR: #5091
Multiprocessing fix by @jubick1337 :: PR: #5106
[Bug fix] PC lexical + audio by @ekmb :: PR: #5109
bugfix: pybtex.database.InvalidNameString: Too many commas in author … by @XuesongYang :: PR: #5112
</details>
Deprecated old scripts for ljspeech. by @XuesongYang :: PR: #4780
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.07
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.05
<details><summary>Issues</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Deprecation by @blisc :: PR: #4082
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.04
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
Megatron BERT export does not currently work in the NVIDIA NGC PyTorch 22.03 container. The issue will be fixed in the NGC PyTorch 22.04 container.
Megatron BERT export does not currently work in the NVIDIA NGC PyTorch 22.03 container. The issue will be fixed in the NGC PyTorch 22.04 container.
Bump TTS deprecation version to 1.9 by @blisc :: PR: #3955
<details><summary>Issues</summary>
</details>
For additional information regarding NeMo containers, please visit: https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo
docker pull nvcr.io/nvidia/nemo:22.03
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
<details><summary>Changelog</summary>
</details>
GPT dataloader improvements and fixes by @crcrpar :: PRs #3826 , #3665
Your coding agent can read these notes before it upgrades. Set up the MCP server →