NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1264 most downloaded on PyPI
PyTorch Lightning is the lightweight PyTorch wrapper for ML researchers. Scale your models. Write less boilerplate.
Last release 8 days ago
10 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
223 releases · first in 2019
Fixed install/upgrade - removing single quote
self.lightningignore (#16080)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.5...1.8.5.post0
Added Lightning{Flow,Work}.lightningignores attributes to programmatically ignore files before uploading to the cloud
Lightning{Flow,Work}.lightningignores attributes to programmatically ignore files before uploading to the cloud (#15818).lightningignore that ignores venv (#16056)DDPStrategy import in app framework (#16029)AutoScaler raising an exception when non-default cloud compute is specified (#15991)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.4.post0...1.8.5
One column per quarter.
Fixed MultiNode Component to use separate cloud computes
L.app.structures (#15964)XLAProfiler not recording anything due to mismatching of action names (#15885)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.4...1.8.4.post0
Fixed torch.jit.script-ing a LightningModule causing an unintended error message about deprecated use_amp property
code_dir argument to tracer run (#15771)lightning run model to launch a LightningLite accelerated script (#15506)lightning delete app to delete a lightning app on the cloud (#15783)AutoScaler component (#15769)ready of the LightningFlow to inform when the Open App should be visible (#15921)_start_method to customize how to start the works (#15923)configure_layout method to the LightningWork which can be used to control how the work is handled in the layout of a parent flow (#15926)lightning run app organization/name (#15941)MultiNode components now warn the user when running with num_nodes > 1 locally (#15806)BuildConfig(requirements=[...]) is passed but a requirements.txt file is already present in the Work (#15799)BuildConfig(dockerfile="...") is passed but a Dockerfile file is already present in the Work (#15799)SingleProcessRuntime (#15933)enable_spawn method of the WorkRunExecutor (#15812)L.app.structures would cause multiple apps to be opened and fail with an error in the cloud (#15911)ImportError on Multinode if package not present (#15963)shuffle=False having no effect when using DDP/DistributedSampler (#15931)fit_loop.restarting to be False for lr finder (#15620)torch.jit.script-ing a LightningModule causing an unintended error message about deprecated use_amp property (#15947)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.3...1.8.4
Nothing published for this version
Fixed the PyTorch Inference locally on GPU
Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.3...1.8.3
Nothing published for this version
Deduplicate top-level lighting CLI command groups
lightning add ssh-key CLI command has been transitioned to lightning create ssh-keylightning remove ssh-key CLI command has been transitioned to lightning delete ssh-keyLightningTrainerScript start-up time (#15751)StreamlitFrontend to support upload in localhost (#15684)LightningFlow (#15750)tensorboard to tensorboardx in TensorBoardLogger (#15728)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.2...1.8.3
Added title and description to ServeGradio
.lightning file (#15654)LightningLite(strategy="ddp_spawn", ...) to LightningLite(strategy="ddp", ...) when on an LSF cluster (#15103)Trainer(strategy="ddp_spawn", ...) to Trainer(strategy="ddp", ...) when on an LSF cluster (#15103](https://github.com/PyTorchLightning/pytorch-lightning/issues/15103))Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.1...1.8.2
Added the start method to the work
start method to the work (#15523)MultiNode Component to run with distributed computation with any frameworks (#15524)RunWorkExecutor to the work and provides default ones for the MultiNode Component (#15561)start_with_flow flag to the LightningWork which can be disabled to prevent the work from starting at the same time as the flow (#15591)bi-directional delta updates between the flow and the works (#15582)--setup flag to lightning run app CLI command allowing for dependency installation via app comments (#15577)flow.flows to be recursive wont to align the behavior with the flow.works (#15466)params argument in TracerPythonScript.run no longer prepends -- automatically to parameters (#15518)lightning CLI taking a long time to error out when the cloud is not reachable (#15412)srun detection causing permission errors (#15485)lightning_lite causing a warning 'Redirects are currently not supported in Windows or MacOs' (#15610)TensorBoardLogger not validating the input array type when logging the model graph (#15323)ColossalAIStrategy at import time when torch.distributed is not available (#15535)fs.listdir with file URI instead of path in CheckpointConnector (#15413)BaseFinetuning callback not setting the track_running_stats attribute for batch normaliztion layers (#15063)WandbLogger(log_model=True|'all) raising an error and not being able to serialize tensors in the metadata (#15544)Trainer(precision=16) and fused optimizers such as Adam(..., fused=True) (#15544)AttributeError with Bagua Strategy (#12534)pytorch_lightning causing a warning 'Redirects are currently not supported in Windows or MacOs' (#15610)Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.0...1.8.1
Implement freeze batchnorm with freezing track running stats by @PososikTeam in https://github.com/Lightning-AI/lightning/pull/15063
Full Changelog: https://github.com/Lightning-AI/lightning/compare/1.8.0...1.8.0.post1
We removed some Callback hooks that were ambiguous to use Removed deprecated callback hooks (#14834):
The core team is excited to announce the release of Lightning 1.8 :zap:
Lightning v1.8 is the culmination of work from 52 contributors who have worked on features, bug-fixes, and documentation for a total of over 550+ commits since v1.7.
<a name="highlights"></a>
Colossal-AI focuses on improving efficiency when training large-scale AI models with billions of parameters. With the new Colossal-AI strategy in Lightning 1.8, you can train existing models like GPT-3 with up to half as many GPUs as usually needed. You can also train models up to twice as big with the same number of GPUs, saving you significant cost. Here is how you use it:
# Select the strategy with good defaults
trainer = Trainer(strategy="colossalai")
# or tune parameters to your liking
from lightning.pytorch.strategies import ColossalAIStrategy
trainer = Trainer(strategy=ColossalAIStrategy(placement_policy="cpu", ...))
You can find Colossal-AI's benchmarks with Lightning on GPT-2 here.
Under the hood, Colossal-AI implements different parallelism algorithms that are especially interesting for the development of SOTA transformer models:
Learn how to install and use Colossal-AI effectively with Lightning here.
NOTE: This strategy is marked as experimental. Stay tuned for more updates in the future.
Introducing encrypted secrets (#14612), a feature requested by Lightning App users :tada:!
Encrypted secrets allow you to securely pass private data to your apps, like API keys, access tokens, database passwords, or other credentials, without exposing them in your code.
Add a secret to your Lightning account in lightning.ai (read more here)
Add an environment variable to your app to read the secret:
# somewhere in your Flow or Work:
GitHubComponent(api_token=os.environ["API_TOKEN"])
Pass the secret to your app run with the following command:
lightning run app app.py --cloud --secret API_TOKEN=github_api_token
These secrets are encrypted and stored in the Lightning database. Nothing except your app can access the value.
NOTE: This is an experimental feature.
Introducing CLI commands for apps (#13602)! As a Lightning App builder, if you want to easily create a CLI interface for users to interract with your app, then this is for you.
Here is an example where users can dynamically create notebooks from the CLI.
All you need to do is implement the configure_commands hook on the LightningFlow:
import lightning as L
from commands.notebook.run import RunNotebook
class Flow(L.LightningFlow):
...
def configure_commands(self):
# Return a list of dictionaries with commands:
return [{"run notebook": RunNotebook(method=self.run_notebook)}]
app = L.LightningApp(Flow())
Once the app is running with lightning run app app.py, you can connect to the app with the following command:
lightning connect {app name} -y
and run the command that was configured:
lightning run notebook --name=my_notebook_name
<a style="visibility:hidden">For a full tutorial and running example, visit our docs. TODO: add to docs</a> NOTE: This is an experimental feature.
In Lightning v1.7, we introduced an integration for PyTorch FSDP in the form of our FSDP strategy, which allows you to train huge models with billions of parameters sharded across hundreds of GPUs and machines.
# Native FSDP implementation
trainer = Trainer(strategy="fsdp_native")
We are continuing to improve the support for this feature by adding automatic wrapping of layers for use cases where the model fits into CPU memory, but not into GPU memory (#14383).
Here are some examples:
Case 1: Model is so large that it does not fit into CPU memory.
Construct your layers in the configure_sharded_model hook and wrap the large ones you want to shard across GPUs:
class MassiveModel(LightningModule):
...
# Create model here and wrap the large layers for sharding
def configure_sharded_model(self):
for i, layer in enumerate(self.block):
self.block[i] = wrap(layer)
...
Case 2: Model fits into CPU memory, but not into GPU memory. In Lightning v1.8, you no longer need to do anything special here, as we can automatically wrap the layers for you using FSDP's policy:
model = MassiveModel()
trainer = Trainer(
accelerator="gpu",
devices=8,
strategy="fsdp_native", # or strategy="fsdp" for fairscale
precision=16
)
# Automatically wraps the layers here:
trainer.fit(model)
Case 3: Model fits into GPU memory. No action required, use any strategy you want.
Note: if you want to manually wrap layers for more control, you can still do that!
Read more about FSDP and how layer wrapping works in our docs.
In this release, we focused on Tuner improvements and introduced two new callbacks that can help you customize the batch size finder and learning rate finder as per your use case.
You can customize the BatchSizeFinder callback to run at different epochs. This feature is useful while fine-tuning models since you can't always use the same batch size after unfreezing the backbone.
from lightning.pytorch.callbacks import BatchSizeFinder
class FineTuneBatchSizeFinder(BatchSizeFinder):
def __init__(self, milestones, *args, **kwargs):
super().__init__(*args, **kwargs)
self.milestones = milestones
def on_fit_start(self, *args, **kwargs):
return
def on_train_epoch_start(self, trainer, pl_module):
if trainer.current_epoch in self.milestones or trainer.current_epoch == 0:
self.scale_batch_size(trainer, pl_module)
trainer = Trainer(callbacks=[FineTuneBatchSizeFinder(milestones=(5, 10))])
trainer.fit(...)
Run batch size finder for validate/test/predict.
from lightning.pytorch.callbacks import BatchSizeFinder
class EvalBatchSizeFinder(BatchSizeFinder):
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
def on_fit_start(self, *args, **kwargs):
return
def on_test_start(self, trainer, pl_module):
self.scale_batch_size(trainer, pl_module)
trainer = Trainer(callbacks=[EvalBatchSizeFinder()])
trainer.test(...)
You can now use the LearningRateFinder callback to run at different intervals. This feature is useful when fine-tuning models, for example.
from lightning.pytorch.callbacks import LearningRateFinder
class FineTuneLearningRateFinder(LearningRateFinder):
def __init__(self, milestones, *args, **kwargs):
super().__init__(*args, **kwargs)
self.milestones = milestones
def on_fit_start(self, *args, **kwargs):
return
def on_train_epoch_start(self, trainer, pl_module):
if trainer.current_epoch in self.milestones or trainer.current_epoch == 0:
self.lr_find(trainer, pl_module)
trainer = Trainer(callbacks=[FineTuneLearningRateFinder(milestones=(5, 10))])
trainer.fit(...)
Even though the LightningCLI class is designed to help in the implementation of command line tools, there are instances when it might be more desirable to run directly from Python. In Lightning 1.8, you can now do this (#14596):
from lightning.pytorch.cli import LightningCLI
def cli_main(args):
cli = LightningCLI(MyModel, ..., args=args)
...
Anywhere in your program, you can now call the CLI directly:
cli_main(["--trainer.max_epochs=100", "--model.encoder_layers=24"])
Learn about all features of the LightningCLI!
Multi-node training on a SLURM cluster has been supported since the inception of Lightning Trainer, and has seen several improvements over time thanks to many community contributions. And we just keep going! In this release, we've added two quality of life improvements:
The preemption/termination signal is now configurable (#14626):
# the default signal is SIGUSR1
trainer = Trainer(plugins=[SLURMEnvironment(requeue_signal=signal.SIGUSR1)])
# customize it for your cluster
trainer = Trainer(plugins=[SLURMEnvironment(requeue_signal=signal.SIGHUP)])
Automatic requeuing of jobs now also works for array jobs (#15040)! Array jobs are a convenient way to group/launch several scripts at once. When the SLURM scheduler interrupts your jobs, Lightning will save a checkpoint, resubmit a new job, and, once the scheduler allocates resources, the Trainer will resume from where it left off.
Read more about our SLURM integration here.
<a name="bc-changes"></a>
This section outlines notable changes that are not backward compatible with previous versions. The full list of changes and removals can be found in the CHANGELOG below.
The signature and behavior of the on_load_checkpoint and on_save_checkpoint callback hooks have changed (#14835):
Before:
def on_save_checkpoint(self, trainer, pl_module, checkpoint):
...
# previously, we were able to return state here
return state
def on_load_checkpoint(self, trainer, pl_module, callback_state):
# previously, only the state for this callback was passed in as argument
...
Now:
def on_save_checkpoint(self, trainer, pl_module, checkpoint):
...
# returning a value here is no longer supported
# you can modify the checkpoint dict directly
return None
def state_dict(self):
...
# Now, return state from this new method
return state
def on_load_checkpoint(self, trainer, pl_module, checkpoint):
# previously, only the state for this callback was passed in as argument
...
def load_state_dict(self, state):
# Now, the state for this callback gets passed to this new method
...
The on_save_checkpoint and on_load_checkpoint hooks on the LightningDataModule have been removed in favor of the state_dict and load_state_dict methods:
-def on_save_checkpoint(self, checkpoint):
- checkpoint["banana"] = self.banana
+def state_dict(self):
+ return dict(banana=self.banana)
-def on_load_checkpoint(self, checkpoint):
- self.banana = checkpoint["banana"]
+def load_state_dict(self, state):
+ self.banana = state["banana"]
We removed some Callback hooks that were ambiguous to use Removed deprecated callback hooks (#14834):
| Old name | New name |
|---|---|
on_batch_start |
on_train_batch_start |
on_batch_end |
on_train_batch_end |
on_epoch_start |
on_train_epoch_start |
on_epoch_start |
on_validation_epoch_start |
on_epoch_start |
on_test_epoch_start |
on_pretrain_routine_start |
on_fit_start |
We cleaned up the properties related to device indices (#14829).
The attributes Trainer.{devices,gpus,num_gpus,ipus,tpu_cores,num_processes,root_gpu,data_parallel_device_ids} have been removed in favor of accelerator-agnostic attributes:
trainer = Trainer(...)
# access the number of devices the trainer uses on this machine ...
print(trainer.num_devices)
# ... or the device IDs
print(trainer.device_ids)
In previous versions of Lightning, switching between the "gloo" and "nccl" backends for multi-GPU, multi-node training was possible through setting an environment variable like so:
PL_TORCH_DISTRIBUTED_BACKEND="gloo" python train.py
But not all strategies support changing the backend in this way. From now on, the backend has to be set in the code (#14693):
trainer = Trainer(strategy=DDPStrategy(process_group_backend="gloo"))
The default remains "nccl", and you should choose "gloo" only for debugging purposes.
Logging with multiple loggers can be super useful (and super easy with Lightning). For example, you could be using one logger to record sensitive image logs to a hosted MLFlow server within your organization, and at the same time log loss curves online to WandB.
trainer = Trainer(
loggers=[WandbLogger(...), MLFlowLogger(...)]
)
Here are two major changes that apply when using multiple loggers in 1.8:
Checkpoints and profiler reports no longer go to a strange folder with a long, hard to remember name (#14325). From now on, these arifacts will land in the version folder of the first logger in the list.
The loggers used to be wrapped by a LoggerCollection object, so that when you accessed trainer.logger you could log to all of them simultaneously. However, this "magic" caused confusion and errors among users and we decided to simplify this (#14283):
# now returns the first logger in the list
print(trainer.logger)
# access all loggers in a list with plural
loggers = trainer.loggers
for logger in loggers:
logger.do_something()
<a name="deprecations"></a>
Why is Lightning deprecating APIs in every release?
Many users have this question, and it is a fair one! Deprecations are a normal part of API evolution in all software. We continually improve Lightning, which means we make APIs like class names, methods, hooks and arguments clear, easy to remember, and general enough to adopt more functionality in the future. Sometimes we have to let old things go to build new and better products.
Learn more about our deprecation window here.
So far, we have followed the pattern of removing deprecated functionality and APIs after two minor versions of deprecation. From Lightning 1.8 onward, we will additionaly convert warnings to error messages after the deprecation phase ends. This way, we can greatly improve the upgrade experience with helpful messages for users who skip more than two minor Lightning versions. The exception to this rule are experimental features, which are marked as such in our documentation.
Here is a summary of major deprecations introduced in 1.8:
| API | Removal version | Alternative |
|---|---|---|
Argument Trainer(amp_level=...) |
1.10 | Trainer(plugins=[ApexMixedPrecisionPlugin(amp_level=...)]) |
Function unwrap_lightning_module |
1.10 | Strategy.lightning_module |
Function unwrap_lightning_module_sharded |
1.10 | Strategy.lightning_module |
Import pl.core.mixins.DeviceDtypeModuleMixin |
1.10 | No longer supported |
Argument LightningCLI(save_config_filename=...) |
1.10 | LightningCLI(save_config_kwargs=dict(config_filename=...)) |
Argument LightningCLI(save_config_overwrite=...) |
1.10 | LightningCLI(save_config_kwargs=dict(overwrite=...)) |
Argument LightningCLI(save_config_multifile=...) |
1.10 | LightningCLI(save_config_kwargs=dict(multifile=...)) |
Enum TrainerFn.TUNING |
1.10 | No longer supported |
Enum RunningStage.TUNING |
1.10 | No longer supported |
Attribute Trainer.tuning |
1.10 | No longer supported |
<a name="changelog"></a>
<details><summary>Added</summary>
load_state_dict and state_dict hooks for LightningFlow components (#14100)--secret option to CLI to allow binding secrets to app environment variables when running in the cloud (#14612)DESCRIPTION attribute (#15193LightningWork to the LightningApp (#15215JustPyFrontend to ease UI creation with https://github.com/justpy-org/justpy (#15002)configure_api and configure_commands to be executed in the Rest API process (#15098</details>
<details><summary>Changed</summary>
--instance-types option when creating clusters (#15314)</details>
<details><summary>Fixed</summary>
</details>
<details><summary>Added</summary>
ddp_fork (and associated alias strategies) with CUDA GPUs (#14983)BatchSizeFinder callback (#11089)LearningRateFinder callback (#13802)method argument which will determine when to run the BatchSizeFinder: one of fit, validate, test or predict (#11089)seed_everything with rank info (#14031)DDPFullyShardedNativeStrategy (#14252)LightningDataModule.from_datasets (#14185)DDPShardedStrategy (#14208)DDPFullyShardedStrategy (#14383)lightning_utilities package (
#14475,
#14537,
#14556,
#14558,
#14575,
#14620)args parameter to LightningCLI to ease running from within Python (#14596)WandbLogger.download_artifact and WandbLogger.use_artifact for managing artifacts with Weights and Biases (#14551)LightningLite.setup() does not have all parameters on the same device (#14822)CometLogger now flags the Comet Experiments as being created from Lightning for analytics purposes (#14906)ckpt_path="hpc" keyword for checkpoint loading (#14911)SaveConfigCallback (#14998)inference_mode flag to Trainer to let users enable/disable inference mode during evaluation (#15034)LightningLite.no_backward_sync for control over efficient gradient accumulation with distributed strategies (#14966)srun command in SLURM and that environment variables are not conflicting (#15011)python -i and an interactive-incompatible strategy (#15293)</details>
<details><summary>Changed</summary>
Trainer.{fit,validate,test,predict,tune} methods now raise a useful error message if the input is not a LightningModule (#13892)MisconfigurationException if batch transfer hooks are overriden with IPUAccelerator (#13961)LightningModule (#13738)on_before_batch_transfer for DPStrategy and IPUAccelerator (#14023)Trainer will now raise an error (#14341)torch.cuda rng state to the aggregate _collect_rng_states() and _set_rng_states() (#14384)trainer.should_stop to not stop in between an epoch and run until min_steps/min_epochs only (#13890)pyDeprecate dependency is no longer installed (#14472)LightningEnvironment when number of SLURM tasks does not correspond to number of processes in Trainer (#14300)lightning_lite.precision.Precision base class (#14798)
PrecisionPlugin.backward signature changed: The closure_loss argument was renamed to tensorPrecisionPlugin.{pre_,post_}backward signature changed: The closure_loss argument was renamed to tensor and moved as the first argumentPrecisionPlugin.optimizer_step signature changed: The model, optimizer_idx and closure arguments need to be passed as keyword arguments nowPL_DISABLE_FORK environment variable introduced in v1.7.4 (#14631)MLFlowLogger.finalize() now sets the status to FAILED when an exception occurred in Trainer, and sets the status to FINISHED on successful completion (#12292)model.double() when using precision=64 in Lightning Lite (#14827)ckpt_path has been set (#14911)Callback.on_load_checkpoint now gets the full checkpoint dictionary and the callback_state argument was renamed checkpoint (#14835)save_hyperparameters() to before the deepcopy (#15132)torch.cuda.device_count and from PyTorch 1.14 and higher, Lightning will configure PyTorch to use a NVML-based check for torch.cuda.is_available. (#15110, #15133)NeptuneLogger now uses neptune.init_run instead of the deprecated neptune.init to initialize a run (#15393)</details>
<details><summary>Deprecated</summary>
LightningDeepSpeedModule (#14000)amp_level from Trainer in favour of passing it explictly via precision plugin (#13898)pytorch_lightning.utiltiies.meta functions in favor of built-in https://github.com/pytorch/torchdistx support (#13868)unwrap_lightning_module and unwrap_lightning_module_sharded utility functions in favor of accessing the unwrapped LightningModule on the strategy directly (#13738)pl_module argument in LightningParallelModule, LightningDistributedModule, LightningShardedDataParallel, LightningBaguaModule and LightningDeepSpeedModule wrapper classes (#13738)on_colab_kaggle function (#14247)pl.core.mixins.DeviceDtypeModuleMixin class (#14511, #14548)pytorch_lightning.utilities.xla_device (#14514, #14550)
inner_f functionpl_multi_process functionXLADeviceUtils.xla_available staticmethodXLADeviceUtils.tpu_device_exists staticmethod in favor of pytorch_lightning.accelerators.TPUAccelerator.is_available()pytorch_lightning.utilities.distributed.tpu_distributed in favor of lightning_lite.accelerators.tpu.tpu_distributed (#14550)pytorch_lightning.utilities.cloud_io in favor of lightning_lite.utilities.cloud_io (#14515)pytorch_lightning.utilities.apply_func in favor of lightning_utilities.core.apply_func (#14516, #14537)pytorch_lightning.utilities.device_parser (#14492, #14753)
pytorch_lightning.utilities.device_parser.determine_root_gpu_device in favor of lightning_lite.utilities.device_parser.determine_root_gpu_devicepytorch_lightning.utilities.device_parser.parse_gpu_ids in favor of lightning_lite.utilities.device_parser.parse_gpu_idspytorch_lightning.utilities.device_parser.is_cuda_available in favor of lightning_lite.accelerators.cuda.is_cuda_availablepytorch_lightning.utilities.device_parser.num_cuda_devices in favor of lightning_lite.accelerators.cuda.num_cuda_devicespytorch_lightning.utilities.device_parser.parse_cpu_cores in favor of lightning_lite.accelerators.cpu.parse_cpu_corespytorch_lightning.utilities.device_parser.parse_tpu_cores in favor of lightning_lite.accelerators.tpu.parse_tpu_corespytorch_lightning.utilities.device_parser.parse_hpus in favor of pytorch_lightning.accelerators.hpu.parse_hpusSaveConfigCallback parameters in LightningCLI.__init__: save_config_kwargs, save_config_overwrite and save_config_multifile. New save_config_kwargs parameter should be used instead (#14998)TrainerFn.TUNING, RunningStage.TUNING and trainer.tuning property (#15100)pl.utilities.distributed.AllGatherGrad implementation in favor of PyTorch's (#15364)</details>
<details><summary>Removed</summary>
Trainer.training_type_plugin property in favor of Trainer.strategy (#14011)DDP2Strategy (#14026)DistributedType and DeviceType enum classes (#14045)rank_zero_warn warning category positionally (#14470)Trainer.get_deprecated_arg_names() (#14415)on_train_batch_end(outputs) format when multiple optimizers are used and TBPTT is enabled (#14373)training_epoch_end(outputs) format when multiple optimizers are used and TBPTT is enabled (#14373)pytorch_lightning.utiltiies.meta functions in favor of built-in https://github.com/pytorch/torchdistx support (#13868)LoggerCollection; Trainer.logger and LightningModule.logger now returns the first logger when more than one gets passed to the Trainer (#14283)trainer.lr_schedulers (#14408)LightningModule.{on_hpc_load,on_hpc_save} hooks in favor of the general purpose hooks LightningModule.{on_load_checkpoint,on_save_checkpoint} (#14315)neptune-client API in the NeptuneLogger (#14727)weights_save_path Trainer argumnent and Trainer.weights_save_path property (#14424)pytorch_lightning.utilities.distributed.rank_zero_only in favor of pytorch_lightning.utilities.rank_zero.rank_zero_onlypytorch_lightning.utilities.distributed.rank_zero_debug in favor of pytorch_lightning.utilities.rank_zero.rank_zero_debugpytorch_lightning.utilities.distributed.rank_zero_info in favor of pytorch_lightning.utilities.rank_zero.rank_zero_infopytorch_lightning.utilities.warnings.rank_zero_warn in favor of pytorch_lightning.utilities.rank_zero.rank_zero_warnpytorch_lightning.utilities.warnings.rank_zero_deprecation in favor of pytorch_lightning.utilities.rank_zero.rank_zero_deprecationpytorch_lightning.utilities.warnings.LightningDeprecationWarning in favor of pytorch_lightning.utilities.rank_zero.LightningDeprecationWarningTrainer.num_processes attribute in favour of Trainer.num_devices (#14423)Trainer.data_parallel_device_ids hook in favour of Trainer.device_ids (#14422)TrainerCallbackHookMixin (#14401)BaseProfiler and AbstractProfiler classes (#14404)PL_TORCH_DISTRIBUTED_BACKEND, in favor of setting the process_group_backend in the strategy constructor (#14693)Callback.on_configure_sharded_model in favor of Callback.setupCallback.on_before_accelerator_backend_setup in favor of Callback.setupCallback.on_batch_start in favor of Callback.on_train_batch_startCallback.on_batch_end in favor of Callback.on_train_batch_endCallback.on_epoch_start in favor of Callback.on_{train,validation,test}_epoch_startCallback.on_epoch_end in favor of Callback.on_{train,validation,test}_epoch_endCallback.on_pretrain_routine_{start,end} in favor of Callback.on_fit_startTrainer.{devices,gpus,num_gpus,ipus,tpu_cores} in favor of the accelerator-agnostic Trainer.num_devices (#14829)LightningIPUModule (#14830)Logger.agg_and_log_metrics hook in favour of Logger.log_metrics and the agg_key_funcs and agg_default_func arguments. (#14840)PrecisionPlugin.on_load_checkpoint and PrecisionPlugin.on_save_checkpoint (#14833)Trainer.root_gpu attribute in favor of Trainer.strategy.root_device (#14829)Trainer.use_amp and LightningModule.use_amp attributes (#14832)Callback.on_init_start and Callback.on_init_end (#14867)Trainer.run_stage in favor of Trainer.{fit,validate,test,predict} (#14870)SimpleProfiler.profile_iterable and AdvancedProfiler.profile_iterable attributes (#14864)Trainer.verbose_evaluate (#14884)Trainer.should_rank_save_checkpoint (#14885)TrainerOptimizersMixin (#14887)Trainer.lightning_optimizers (#14889)TrainerDataLoadingMixin (#14888)Trainer.call_hook in favor of Trainer._call_callback_hooks, Trainer._call_lightning_module_hook, Trainer._call_ttp_hook, and Trainer._call_accelerator_hook (#14869)Trainer.{validated,tested,predicted}_ckpt_path (#14897)device_stats_monitor_prefix_metric_keys (#14890)LightningDataModule.on_save/load_checkpoint hooks (#14909)Callback.on_save_checkpoint in favor of implementing Callback.state_dict (#14835)</details>
<details><summary>Fixed</summary>
LightningLite.setup() not setting the .device attribute correctly on the returned wrapper (#14822)StochasticWeightAveraging callback (#14836)NeptuneLogger() (#14919)save_dir is overridden by None dir when using CLI (#14878)LightningDataModule.load_state_dict hook while restoring checkpoint using LightningDataModule.load_from_checkpoint (#14883)SaveConfigCallback instances should only save the config once to allow having the overwrite=False safeguard when using LightningCLI(..., run=False) (#14927)StopIteration exception is raised while using an IterableDataset (#14940)Trainer support for PyTorch built without distributed support (#14971)StochasticWeightAveraging callback (#14866)LightningCLI parse_env and description in subcommands (#15138)multiprocessing.Pool after importing Lightning (#15292)RichProgressBar together with checkpointing (#15319)RichProgressBar crashing when used with distributed strategies (#15376)RichProgressBar not resetting the internal state for the sanity check progress (#15377)</details>
Full commit list: https://github.com/PyTorchLightning/pytorch-lightning/compare/1.7.0...1.8.0
<a name="contributors"></a>
@akihironitta @ananthsub @AndresAlgaba @ar90n @Atharva-Phatak @awaelchli @BongYang @Borda @carmocca @dependabot @donlapark @ethanwharris @Felonious-Spellfire @hhsecond @jerome-habana @JustinGoheen @justusschock @kaushikb11 @krishnakalyan3 @krshrimali @luca-medeiros @manangoel99 @manskx @mauvilsa @MrShevan @nicolai86 @nmiculinic @otaj @Queuecumber @rlizzo @rohitgr7 @rschireman @SeanNaren @speediedan @tchaton @tshu-w
@Birch-san @clementpoiret @HalestormAI @thongonary @alecmerdler @adam-lightning @yurijmikhalevich @lijm1358 @robert-s-lee @panos-is @kacperlukawski @alro923 @dmitsf @Anner-deJong @cschell @nishantb06 @Callidior @j0rd1smit @MarcSkovMadsen @KralaBenjamin @robertomest @daniel347x @pierocor @datumbox @nohalon @pritamsoni-hsr @nandwalritik @gilfree @ritsuki1227 @christopher-nguyen-re @JulesGM @jgbos @dconathan @jsr-p @NeoKish @Blaizzy @suyash-811 @alexkuzmik @ziyadsheeba @geoffrey-g-delhomme @amrutha1098 @AlessioQuercia @ver217 @Helias @zxvix @1SAA @fabiofumarola @luca3rd @kimpty @PaulLerner @rbracco @wouterzwerink
If we forgot somebody or you have a suggestion, find support here :zap:
Chuck Norris can write functions of infinite recursion ... and have them return.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixed the availability check for the neptune-client package
TensorBoardLogger.finalize creating a new experiment when none was created during the Trainer's execution (#14762)TypeError on import when torch.distributed is not available (#14809)@awaelchli @Borda @carmocca @dependabot @otaj @raoakarsha
If we forgot someone due to not matching commit email with GitHub account, let us know :)
Improved the error messaging when passing Trainer.method(model, x_dataloader=None) with no module-method implementations available
Trainer.method(model, x_dataloader=None) with no module-method implementations available (#14614)mode="power" (#14372)self.log-ing a tensor would create a user warning from PyTorch about cloning tensors (#14599)torch.distributed is not available (#14454)@akihironitta @awaelchli @Borda @carmocca @dependabot @krshrimali @mauvilsa @pierocor @rohitgr7 @wangraying
If we forgot someone due to not matching commit email with GitHub account, let us know :)
Squeezed tensor values when logging with LightningModule.log
LightningModule.log (#14489)WandbLogger save_dir is not set after creation (#14326)Trainer.estimated_stepping_batches when maximum number of epochs is not set (#14317)@carmocca @dependabot @robertomest @rohitgr7 @tshu-w
If we forgot someone due to not matching commit email with GitHub account, let us know :)
Added an environment variable PL_DISABLE_FORK that can be used to disable all forking in the Trainer
PL_DISABLE_FORK that can be used to disable all forking in the Trainer (#14319)LightningDataModule hparams parsing (#12806)lr_find() so that the correct LR schedule is used for the actual training (#14113)@rohitgr7 @tanmoyio @justusschock @cschell @carmocca @Callidior @awaelchli @j0rd1smit @dependabot @Borda @otaj
Fixed an assertion error when using a ReduceOnPlateau scheduler with the Horovod strategy
ReduceOnPlateau scheduler with the Horovod strategy (#14215)AttributeError when accessing LightningModule.logger and the Trainer has multiple loggers (#14234)RichProgressBar (#14296)logging in the configure_gradient_clipping hook after unintended removal in v1.7.2 (#14298)reload_dataloaders_every_n_epochs for validation (#13964)@awaelchli @Borda @carmocca @dependabot @kaushikb11 @otaj @rohitgr7
Avoid metadata.entry_points deprecation warning on Python 3.10
FullyShardedNativeNativeMixedPrecisionPlugin to handle precision for DDPFullyShardedNativeStrategy (#14092)on_before_batch_transfer, transfer_batch_to_device, on_after_batch_transfer, configure_gradient_clipping, clip_gradients (#14069)MisconfigurationException if batch transfer hooks are overriden with IPUAccelerator (13961)WandbLogger is now "lightning_logs" (#14145)WandbLogger.name property no longer returns the name of the experiment, and instead returns the project's name (#14145)AttributeError when multiple DataLoader classes are imported (#14117)LightningModule or LightningDataModule (#14151)LightningModule.cuda() gets called without specifying a device index and the current cuda device was not 0 (#14128)sync_dist when using torchmetrics (#14143)metadata.entry_points deprecation warning on Python 3.10 (#14052)WandbLogger would be set to the project name instead of a randomly generated string (#14145)DataLoader and BatchSampler when instantiated inside *_dataloader hooks (#14212)@adamreeve @akihironitta @awaelchli @Borda @carmocca @dependabot @otaj @rohitgr7
Casted only floating point tensors to fp16 with IPUs
DeepSpeedStrategy (#14000)NeptuneLogger dependency being unrecognized (#13988)max_epochs even when fast_dev_run was set (#13262)precision="mixed" being used with DeepSpeedStrategy and IPUStrategy (#14041)ddp_find_unused_parameters to be set False, whereas the intended default is True (#14095)@adamjstewart @akihironitta @awaelchli @Birch-san @carmocca @clementpoiret @dependabot @rohitgr7
The new version renders the registries and the auto_registry flag, introduced in 1.6.0, unnecessary, so we have deprecated them.
The core team is excited to announce the release of PyTorch Lightning 1.7 :zap:
PyTorch Lightning 1.7 is the culmination of work from 106 contributors who have worked on features, bug-fixes, and documentation for a total of over 492 commits since 1.6.0.
<a name="highlights"></a>
For those using PyTorch 1.12 on M1 or M2 Apple machines, we have created the MPSAccelerator. MPSAccelerator enables accelerated GPU training on Apple’s Metal Performance Shaders (MPS) as a backend process.
NOTE
Support for this accelerator is currently marked as experimental in PyTorch. Because many operators are still missing, you may run into a few rough edges.
# Selects the accelerator
trainer = pl.Trainer(accelerator="mps")
# Equivalent to
from pytorch_lightning.accelerators import MPSAccelerator
trainer = pl.Trainer(accelerator=MPSAccelerator())
# Defaults to "mps" when run on M1 or M2 Apple machines
# to avoid code changes when switching computers
trainer = pl.Trainer(accelerator="gpu")
PyTorch 1.12 also added native support for Fully Sharded Data Parallel (FSDP). Previously, PyTorch Lightning enabled this by using the fairscale project. You can now choose between both options.
NOTE
Support for this strategy is marked as beta in PyTorch.
# Native PyTorch implementation
trainer = pl.Trainer(strategy="fsdp_native")
# Equivalent to
from pytorch_lightning.strategies import DDPFullyShardedNativeStrategy
trainer = pl.Trainer(strategy=DDPFullyShardedNativeStrategy())
# For reference, FairScale's implementation can be used with
trainer = pl.Trainer(strategy="fsdp")
Collaborative Training solves the need for top-tier multi-GPU servers by allowing you to train across unreliable machines such as local ones or even preemptible cloud compute across the Internet.
Under the hood, we use Hivemind. This provides de-centralized training across the Internet.
from pytorch_lightning.strategies import HivemindStrategy
trainer = pl.Trainer(
strategy=HivemindStrategy(target_batch_size=8192),
accelerator="gpu",
devices=1
)
For more information, check out the docs.
So far, the only multi-GPU strategy supported in Jupyter notebooks (including Grid.ai, Google Colab, and Kaggle, for example) has been the Data-Parallel (DP) strategy (strategy="dp"). DP, however, has several limitations that often obstruct users' workflows. It can be slow, it's incompatible with TorchMetrics, it doesn't persist state changes on replicas, and it's difficult to use with non-primitive input- and output structures.
In this release, we've added support for Distributed Data Parallel in Jupyter notebooks using the fork mechanism to address these shortcomings. This is only available for MacOS and Linux (sorry Windows!).
NOTE
This feature is experimental.
This is how you use multi-device in notebooks now:
# Train on 2 GPUs in a Jupyter notebook
trainer = pl.Trainer(accelerator="gpu", devices=2)
# Can be set explicitly
trainer = pl.Trainer(accelerator="gpu", devices=2, strategy="ddp_notebook")
# Can also be used in non-interactive environments
trainer = pl.Trainer(accelerator="gpu", devices=2, strategy="ddp_fork")
By default, the Trainer detects the interactive environment and selects the right strategy for you. Learn more in the full documentation.
If a run is configured to save to the same directory as a previous run and ModelCheckpoint(save_last=True) is enabled, the "last" checkpoint is now versioned with a simple -v1 suffix to avoid overwriting the existing "last" checkpoint. This mimics the behaviour for checkpoints that monitor a metric.
In certain scenarios, like when running in a cloud spot instance with fault-tolerant training enabled, it is useful to load the latest available checkpoint. It is now possible to pass the string ckpt_path="last" in order to load the latest available checkpoint from the set of existing checkpoints.
trainer = Trainer(...)
trainer.fit(..., ckpt_path="last")
In some cases, for example iteration based training, it is useful to run validation after every N number of training batches without being limited by the epoch boundary. Now, you can enable validation based on total training batches.
trainer = Trainer(..., val_check_interval=N, check_val_every_n_epoch=None)
trainer.fit(...)
For example, given 5 epochs of 10 batches, setting N=25 would run validation in the 3rd and 5th epoch.
PyTorch Lightning provides the DeviceStatsMonitor callback to monitor the stats of the hardware currently used. However, users often also want to monitor the stats of other hardware. In this release, we have added an option to additionally monitor CPU stats:
from pytorch_lightning.callbacks import DeviceStatsMonitor
# Log both CPU stats and GPU stats
trainer = pl.Trainer(callbacks=DeviceStatsMonitor(cpu_stats=True), accelerator="gpu")
# Log just the GPU stats
trainer = pl.Trainer(callbacks=DeviceStatsMonitor(cpu_stats=False), accelerator="gpu")
# Equivalent to `DeviceStatsMonitor()`
trainer = pl.Trainer(callbacks=DeviceStatsMonitor(cpu_stats=True), accelerator="cpu")
The CPU stats are gathered using the psutil package.
It is now possible to use custom samplers in a distributed environment without the need to set replace_ddp_sampler=False and wrap your sampler manually with the DistributedSampler.
PyTorch 1.9 introduced torch.inference_mode, which is a faster alternative for torch.no_grad. Lightning will now use inference_mode wherever possible during evaluation.
In Pytorch 1.11, operations that do not have a deterministic implementation can be set to throw a warning instead of an error when ran in deterministic mode. This is now supported by our Trainer:
trainer = pl.Trainer(deterministic="warn")
After the latest updates to jsonargparse, the library supporting the LightningCLI, there's now complete support for shorthand notation. This includes automatic support for shorthand notation to all arguments, not just the ones that are part of the registries, plus support inside configuration files.
+ # pytorch_lightning==1.7.0
trainer:
callbacks:
- - class_path: pytorch_lightning.callbacks.EarlyStopping
+ - class_path: EarlyStopping
init_args:
monitor: "loss"
A header with the version that generated the config is now included.
All subclasses for a given base class can be specified by name, so there's no need to explicitly register them. The only requirement is that the module where the subclass is defined is imported prior to parsing.
from pytorch_lightning.cli import LightningCLI
import my_code.models
import my_code.optimizers
cli = LightningCLI()
# Now use any of the classes:
# python trainer.py fit --model=Model1 --optimizer=CustomOptimizer
The new version renders the registries and the auto_registry flag, introduced in 1.6.0, unnecessary, so we have deprecated them.
Support was also added for list appending; for example, to add a callback to an existing list that might be already configured:
$ python trainer.py fit \
- --trainer.callbacks=EarlyStopping \
+ --trainer.callbacks+=EarlyStopping \
--trainer.callbacks.patience=5 \
- --trainer.callbacks=LearningRateMonitor \
+ --trainer.callbacks+=LearningRateMonitor \
--trainer.callbacks.logging_interval=epoch
Entry Points are an advanced feature in Python's setuptools that allow packages to expose metadata to other packages. In Lightning, we allow an arbitrary package to include callbacks that the Lightning Trainer can automatically use when installed, without you having to manually add them to the Trainer. This is useful in production environments where it is common to provide specialized monitoring and logging callbacks globally for every application.
A setup.py file for a callbacks plugin package could look something like this:
from setuptools import setup
setup(
name="my-package",
version="0.0.1",
entry_points={
# Lightning will look for this key here in the environment:
"pytorch_lightning.callbacks_factory": [
"monitor_callbacks=factories:my_custom_callbacks_factory"
]
},
)
Read more about callback entry points in our docs.
EarlyStopping messagesOur EarlyStopping callback implementation, by default, logs the stopping messages on every rank when it's run in a distributed environment. This was done in case the monitored values were not synchronized. However, some users found this verbose. To avoid this, you can now set a flag:
from pytorch_lightning.callbacks import EarlyStopping
trainer = pl.Trainer(callbacks=EarlyStopping(..., log_rank_zero_only=True))
Checkpoint class for extra customizationIf you want to customize ModelCheckpoint callback, without all the extra functionality this class provides, this release provides an empty class Checkpoint for easier inheritance. In all internal code, the check is made against the Checkpoint class in order to ensure everything works properly for custom classes.
Setting overfit_batches=N, now enables validation and runs N number of validation batches during trainer.fit.
# Uses 1% of each train & val set
trainer = Trainer(overfit_batches=0.01)
# Uses 10 batches for each train & val set
trainer = Trainer(overfit_batches=10)
DeviceStatsMonitor callback can now be used to automatically monitor and log device stats during the training stage with Habana devices.
from pytorch_lightning import Trainer
from pytorch_lightning.callbacks import DeviceStatsMonitor
device_stats = DeviceStatsMonitor()
trainer = Trainer(accelerator="hpu", callbacks=[device_stats])
LightningDataModule.load_from_checkpointNow, hyper-parameters from LightningDataModule save to checkpoints and reload when training is resumed. And just like you use LightningModule.load_from_checkpoint to load a model using a checkpoint filepath, you can now load LightningDataModule using the same hook.
# Lad weights without mapping ...
datamodule = MyLightningDataModule.load_from_checkpoint('path/to/checkpoint.ckpt')
# Or load weights and hyperparameters from separate files.
datamodule = MyLightningDataModule.load_from_checkpoint(
'path/to/checkpoint.ckpt',
hparams_file='/path/to/hparams_file.yaml'
)
# Override some of the params with new values
datamodule = MyLightningDataModule.load_from_checkpoint(
'path/to/checkpoint.ckpt',
batch_size=32,
num_workers=10,
)
When serving models in production, it generally is a good pratice to ensure that the model can be served and optimzed before starting training to avoid wasting money.
To do so, you can import a ServableModule (an nn.Module) and add it as an extra base class to your base model as follows:
from pytorch_lightning import LightningModule
from pytorch_lightning.serve import ServableModule
class ProductionReadyModel(LightningModule, ServableModule):
...
To make your model servable, you would need to implement three hooks:
configure_payload: Describe the format of the payload (data sent to the server).configure_serialization: Describe the functions used to convert the payload to tensors (de-serialization) and tensors to payload (serialization)serve_step: The method used to transform the input tensors to a dictionary of prediction tensors.from pytorch_lightning.serve import ServableModule, ServableModuleValidator
class ProductionReadyModel(LitModule, ServableModule):
def configure_payload(self):
# 1: Access the train dataloader and load a single sample.
image, _ = self.trainer.train_dataloader.loaders.dataset[0]
# 2: Convert the image into a PIL Image to bytes and encode it with base64
pil_image = T.ToPILImage()(image)
buffered = BytesIO()
pil_image.save(buffered, format="JPEG")
img_str = base64.b64encode(buffered.getvalue()).decode("UTF-8")
payload = {"body": {"x": img_str}}
return payload
def configure_serialization(self):
deserializers = {"x": Image(224, 224).deserialize}
serializers = {"output": Top1().serialize}
return deserializers, serializers
def serve_step(self, x: torch.Tensor) -> Dict[str, torch.Tensor]:
return {"output": self.model(x)}
Finally, add the ServableModuleValidator callback to the Trainer to validate the model is servable on_train_start. This uses a FastAPI server.
pl_module = ProductionReadyModel()
trainer = Trainer(..., callbacks=[ServableModuleValidator()])
trainer.fit(pl_module)
Have a look at the full example here.
You can now save checkpoints asynchronously using the AsyncCheckpointIO plugin without blocking your training process. To enable this, you can pass a AsyncCheckpointIO plugin to the Trainer.
from pytorch_lightning.plugins.io import AsyncCheckpointIO
trainer = Trainer(plugins=[AsyncCheckpointIO()])
Have a look at the full example here.
<a name="bc-changes"></a>
This section outlines notable changes that are not backward compatible with previous versions. The full list of changes and removals can be found in the CHANGELOG below.
The DDP2 strategy, previously known as the DDP2 plugin, has been part of Lightning since its inception. Due to both the technical challenges in maintaining the plugin after PyTorch's removal of the multi-device support in DistributedDataParallel, as well as a general lack of interest, we have decided to retire the strategy entirely.
In previous versions, metrics logged inside epoch-end hooks were forcefully synced. This makes the sync_dist flag irrelevant and causes communication overhead that might be undesired. In this release, we've removed this behaviour and instead warn the user that synchronization might be desired.
<a name="deprecations"></a>
| API | Removal version | Alternative |
|---|---|---|
Import pytorch_lightning.loggers.base.LightningLoggerBase |
1.9 | pytorch_lightning.loggers.logger.Logger |
Import pytorch_lightning.callbacks.base.Callback |
1.9 | pytorch_lightning.callbacks.callback.Callback |
Import pytorch_lightning.core.lightning.LightningModule |
1.9 | pytorch_lightning.core.module.LightningModule |
Import pytorch_lightning.loops.base.Loop |
1.9 | pytorch_lightning.loops.loop.Loop |
Import pytorch_lightning.profiler |
1.9 | pytorch_lightning.profilers |
Arguments Trainer(num_processes=..., gpus=..., tpu_cores=..., ipus=...) |
2.0 | Trainer(accelerator=..., devices=...) |
Argument LightningCLI(seed_everything_default=None) |
1.9 | LightningCLI(seed_everything_default=False) |
Method Trainer.reset_train_val_dataloaders() |
1.9 | Trainer.reset_{train,val}_dataloader |
Import pytorch_lightning.utilities.cli module |
1.9 | pytorch_lightning.cli |
Objects pytorch_lightning.utilities.cli.{OPTIMIZER,LR_SCHEDULER,MODEL,DATAMODULE,CALLBACK,LOGGER}_REGISTRY |
1.9 | Not necessary anymore |
Argument LightningCLI(auto_registry=...) |
1.9 | Not necessary anymore |
Argument Trainer(strategy="ddp2") and class pytorch_lightning.strategies.DDP2Strategy |
1.8 | No longer supported |
<a name="changelog"></a>
<details><summary>Added</summary>
ServableModule and its associated callback called ServableModuleValidator to ensure the model can served (#13614)PossibleUserWarning (#13377)log_rank_zero_only to EarlyStopping to disable logging to non-zero rank processes (#13233)ckpt_path="last" (#12816)LightningDataModule.load_from_checkpoint to support loading datamodules directly from checkpoint (#12550)Trainer.save_checkpoint() without a model attached (#12772)DeepSpeedStrategy on unsupported accelerators (#12699)torch.inference_mode for evaluation and prediction (#12715)val_check_interval to a value higher than the amount of training batches when check_val_every_n_epoch=None (#11993)pytorch_lightning version as a header in the CLI config files (#12532)Callback registration through entry points (#12739)Trainer(deterministic="warn") to warn instead of fail when a non-deterministic operation is encountered (#12588)__next__ calls (#12124)predict_dataset argument in LightningDataModule.from_datasets to create predict dataloaders (#12942)DeviceStatsMonitor (#12228)DistributedSamplerWrapper (#12959)LightningDataModule hooks (#12971)Checkpoint class to inherit from (#13024)DeviceStatsMonitor (#11795)teardown() method to Accelerator (#11935)timeout argument to DDPStrategy and DDPSpawnStrategy. (#13244, #13383)XLAEnvironment cluster environment plugin (#11330)FitLoop stopping conditions are met (#9749)DummyLogger (#13224Trainer reference for ensembles of LightningModules (#13638MPSAccelerator (#13123)</details>
<details><summary>Changed</summary>
accelerator="gpu" now automatically selects an available GPU backend (CUDA and MPS currently) (#13642)extract_batch_size (#12573)weights_save_path/name/version/checkpoints to weights_save_path/checkpoints (#12372)weights_save_path/name1_name2/version1_version2/checkpoints to weights_save_path/checkpoints (#12372)swa_lrs argument in StochasticWeightAveraging callback as required (#12556)LightningCLI's shorthand notation changed to use jsonargparse native feature (#12614)LightningCLI changed to use jsonargparse native support for list append (#13129)seed_everything_default argument in the LightningCLI to type Union[bool, int]. If set to True a seed is automatically generated for the parser argument --seed_everything. (#12822, #13110)add_argparse_args function. (#12504)limit_train_batches (#12885)DataLoader instantiated inside a *_dataloader hook will not set the passed arguments as attributes anymore (#12981)WandbLogger will now use the run name in the logs folder if it is provided, and otherwise the project name (#12604)sync_dist=True on epoch end (13364)val_check_interval(int) to consider total train batches processed instead of _batches_that_stepped for validation check during training (#12832auto_device_count, is_available & get_device_name methods based on the latest torch habana package (#13423)BatchSampler when running on multiple IPUs (#13854)</details>
<details><summary>Deprecated</summary>
pytorch_lightning.accelerators.gpu.GPUAccelerator in favor of pytorch_lightning.accelerators.cuda.CUDAAccelerator (#13636)pytorch_lightning.loggers.base.LightningLoggerBase in favor of pytorch_lightning.loggers.logger.Logger, and deprecated pytorch_lightning.loggers.base in favor of pytorch_lightning.loggers.logger (#120148)pytorch_lightning.callbacks.base.Callback in favor of pytorch_lightning.callbacks.callback.Callback (#13031)num_processes, gpus, tpu_cores, and ipus from the Trainer constructor in favor of using the accelerator and devices arguments (#11040)LightningCLI(seed_everything_default=None) in favor of False (#12804).pytorch_lightning.core.lightning.LightningModule in favor of pytorch_lightning.core.module.LightningModule (#12740)pytorch_lightning.loops.base.Loop in favor of pytorch_lightning.loops.loop.Loop (#13043)Trainer.reset_train_val_dataloaders() in favor of Trainer.reset_{train,val}_dataloader (#12184)pytorch_lightning.utilities.cli.LightningCLI in favor of equivalent copies in pytorch_lightning.cli.LightningCLI (#13767)pytorch_lightning.profiler in favor of pytorch_lightning.profilers (#12308)</details>
<details><summary>Removed</summary>
IndexBatchSamplerWrapper.batch_indices (#13565)LightningModule.add_to_queue and LightningModule.get_from_queue method (#13600)pytorch_lightning.core.decorators.parameter_validation from decorators (#13514)Logger.close method (#13149)weights_summary argument from the Trainer constructor (#13070)flush_logs_every_n_steps argument from the Trainer constructor (#13074)process_position argument from the Trainer constructor (13071)checkpoint_callback argument from the Trainer constructor (#13027)on_{train,val,test,predict}_dataloader hooks from the LightningModule and LightningDataModule (#13033)TestTubeLogger (#12859)pytorch_lightning.core.memory.LayerSummary and pytorch_lightning.core.memory.ModelSummary (#12593)summarize method from the LightningModule (#12559)model_size property from the LightningModule class (#12641)stochastic_weight_avg argument from the Trainer constructor (#12535)progress_bar_refresh_rate argument from the Trainer constructor (#12514)prepare_data_per_node argument from the Trainer constructor (#12536)pytorch_lightning.core.memory.{get_gpu_memory_map,get_memory_profile} (#12659)terminate_on_nan argument from the Trainer constructor (#12553)XLAStatsMonitor callback (#12688)pytorch_lightning.callbacks.progress.progress (#12658)dim and size arguments from the LightningDataModule constructor(#12780)train_transforms argument from the LightningDataModule constructor(#12662)log_gpu_memory argument from the Trainer constructor (#12657)GPUStatsMonitor callback (#12554)val_transforms argument from the LightningDataModule constructor (#12763)test_transforms argument from the LightningDataModule constructor (#12773)Trainer(max_steps=None) (#13591)dataloader_idx argument from on_train_batch_start/end hooks Callback and LightningModule (#12769, #12977)get_progress_bar_dict property from LightningModule (#12839)Strategy.post_dispatch() hook (#13461)pytorch_lightning.callbacks.lr_monitor.LearningRateMonitor.lr_sch_names (#13353)Trainer.slurm_job_id in favor of SLURMEnvironment.job_id (#13459)DDP2Strategy (#12705)LightningDistributed (#13549)master_address and master_port in favor of main_address and main_port (#13458)KubeflowEnvironment.is_using_kubelfow(), LSFEnvironment.is_using_lsf() and TorchElasticEnvironment.is_using_torchelastic() in favor of the detect() method (#13458)Callback.on_keyboard_interrupt (#13438)LightningModule.on_post_move_to_device (#13548)TPUSpawnStrategy.{tpu_local_core_rank,tpu_global_core_rank} attributes in favor of TPUSpawnStrategy.{local_rank,global_rank} (#11163)SingleTPUStrategy.{tpu_local_core_rank,tpu_global_core_rank} attributes in favor of SingleTPUStrategy.{local_rank,global_rank}(#11163)</details>
<details><summary>Fixed</summary>
DataLoaders when instantiated in *_dataloader hook (#12981)BatchSamplers when instantiated in *_dataloader hook #13640)LightningLite.setup() now properly supports pass-through when looking up attributes (#12597)LightningCLI signature parameter resolving for some lightning classes (#13283)pytorch_lightning.utilities.distributed.gather_all_tensors to handle tensors of different dimensions (#12630)Trainer.predict(return_predictions=False) to track prediction's batch_indices (#13629)CheckpointIO plugin with strategies (#13785)val_check_interval=int and check_val_every_n_epoch=None (#12832ReduceLROnPlateau scheduler if reduce_on_plateau is set by the user in scheduler config (#13838)global_step while restoring logging step for old checkpoints (#13645)precision=16 on IPU, the cast has been moved off the IPU onto the host, making the copies from host to IPU cheaper (#13880)amp_level for DeepSpeedPrecisionPlugin to O2 (#13897)TQDMProgressBar reset and update to show correct time estimation (2/2) (#13962)</details>
Full commit list: https://github.com/PyTorchLightning/pytorch-lightning/compare/1.6.0...1.7.0
<a name="contributors"></a>
@akashkw @akihironitta @aniketmaurya @awaelchli @Benjamin-Etheredge @Borda @carmocca @catalys1 @daniellepintz @edenlightning @edward-io @EricWiener @fschlatt @ftorres16 @jerome-habana @justusschock @karthikrangasai @kaushikb11 @krishnakalyan3 @krshrimali @mauvilsa @nikvaessen @otaj @pre-commit-ci @puhuk @raoakarsha @rasbt @rohitgr7 @SeanNaren @s-rog @talregev @tchaton @tshu-w @twsl @weiji14 @williamFalcon @WrRan
@alvitawa @aminst @ankitaS11 @ar90n @Atharva-Phatak @bibhabasumohapatra @BongYang @code-review-doctor @CompRhys @Cyprien-Ricque @dependabot @digital-idiot @DN6 @donlapark @ekagra-ranjan @ethanfurman @gautierdag @georgestein @HallerPatrick @HenryLau0220 @hhsecond @himkt @HMellor @igorgad @inwaves @ishtos @JeroenDelcour @JiahaoYao @jiny419 @jinyoung-lim @JustinGoheen @jxmorris12 @Keiku @kingjuno @lsy643 @luca-medeiros @lukasugar @maciek-pioro @mads-oestergaard @manskx @martinosorb @MohammedAlkhrashi @MrShevan @myxik @naisofly @NathanielDamours @nayoungjun @niberger @nitinramvelraj @nninept @pbsds @Pragyanstha @PrajwalBorkar @Prometheos2 @rampartrange @rhjohnstone @rschireman @samz5320 @Schinkikami @semaphore-egg @shantam-8 @shenoynikhil @sisilmehta2000 @s-kumano @stanbiryukov @talregev @tanmoyio @tkonopka @vumichien @wangherr @yhl48 @YongWookHa
If we forgot somebody or you have a suggestion, find support here :zap:
Chuck Norris can unit-test entire applications with a single assert.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixed estimated_stepping_batches requiring distributed comms in configure_optimizers for the DeepSpeedStrategy
estimated_stepping_batches requiring distributed comms in configure_optimizers for the DeepSpeedStrategy (#13350).set_epoch() also on batch samplers if the dataloader has one wrapped in a distributed sampler (#13396)@adamjstewart @akihironitta @awaelchli @Borda @martinosorb @rohitgr7 @SeanNaren
Added all DDP params to be exposed through hpu parallel strategy
torch.backends.cudnn.benchmark=False by default (unlike in v1.6.{0-4}) after speed and memory problems depending on the data used. Please consider tuning Trainer(benchmark) manually. (#13154)torch.backends.cudnn.benchmark when Trainer(benchmark=...) is not set (#13154)Trainer(precision=64) during evaluation which now uses the wrapped precision module (#12983)LightningModule for evaluation during trainer.fit for BaguaStrategy (#12983)LightningModule so it can be deleted (#12897)materialize_module setting a module's child recursively (#12870)Profiler to the Trainer (#13084)DDPStrategy and DDPSpawnStrategy to initialize optimizers only after moving the module to the device (#11952)@akihironitta @ananthsub @ar90n @awaelchli @Borda @carmocca @dependabot @jerome-habana @mads-oestergaard @otaj @rohitgr7
Fixed trainer.logger deprecation message
rich.console.Console throughout codebase (#12886)DeepspeedStrategy (#12887)trainer.logger deprecation message (#12671)ShardedStrategy (#12915)Trainer.validate and Trainer.test (#12857)KFoldLoop (#12441)optimizer_zero_grad from being called after IPU execution (#12913)fuse_modules to be qat-aware for torch>=1.11 (#12891)DDPFullyShardedStrategy when precision=16 (#12965)TQDMProgressBar reset and update to show correct time estimation (#12889)@akihironitta @carmocca @HMellor @jerome-habana @kaushikb11 @krshrimali @mauvilsa @niberger @ORippler @otaj @rohitgr7 @SeanNaren
Fixed ImportError when torch.distributed is not available.
ImportError when torch.distributed is not available. (#12794)ModelCheckpoint monitors with dots (#12783)@akihironitta @alvitawa @awaelchli @Borda @carmocca @code-review-doctor @ethanfurman @HenryLau0220 @krshrimali @otaj
Support strategy argument being case insensitive
strategy argument being case insensitive (#12528)TQDMProgressBar (#12563)average_parameters multiple times per optimizer step (#12452)super().__init__() (#12609)LightningLite.run method accepts user-defined arguments (#12629)rank_zero_only decorator in LSF environments (#12587)nn.Module is not saved under hparams (#12669)MisconfigurationException when the accelerator is available but the user passes invalid ([]/0/"0") values to the devices flag (#12708)auto_select_gpus with the accelerator and devices API (#12608)@akihironitta @awaelchli @Borda @carmocca @kaushikb11 @krshrimali @mauvilsa @otaj @pre-commit-ci @rohitgr7 @semaphore-egg @tkonopka @wayi1
If we forgot someone due to not matching the commit email with the GitHub account, let us know :]
Fixed security vulnerabilities CVE-2020-1747 and CVE-2020-14343 caused by the PyYAML dependency
The core team is excited to announce the PyTorch Lightning 1.6 release ⚡
<a name="highlights"></a>
PyTorch Lightning 1.6 is the work of 99 contributors who have worked on features, bug-fixes, and documentation for a total of over 750 commits since 1.5. This is our most active release yet. Here are some highlights:
Lightning 1.6 now supports the Habana® framework, which includes Gaudi® AI training processors. Their heterogeneous architecture includes a cluster of fully programmable Tensor Processing Cores (TPC) along with its associated development tools and libraries and a configurable Matrix Math engine.
You can leverage the Habana hardware to accelerate your Deep Learning training workloads simply by passing:
trainer = pl.Trainer(accelerator="hpu")
# single Gaudi training
trainer = pl.Trainer(accelerator="hpu", devices=1)
# distributed training with 8 Gaudi
trainer = pl.Trainer(accelerator="hpu", devices=8)
The Bagua Strategy is a deep learning acceleration framework that supports multiple, advanced distributed training algorithms with state-of-the-art system relaxation techniques. Enabling Bagua, which can be considerably faster than vanilla PyTorch DDP, is as simple as:
trainer = pl.Trainer(strategy="bagua")
# or to choose a custom algorithm
trainer = pl.Trainer(strategy=BaguaStrategy(algorithm="gradient_allreduce") # default
The Accelerator, Strategy, and Plugin APIs are a core part of PyTorch Lightning. They're where all the distributed boilerplate lives, and we're constantly working to improve both them and the overall PyTorch Lightning platform experience.
In this release, we've made some large changes to achieve that goal. Not to worry, though! The only users affected by these changes are those who use custom implementations of Accelerator and Strategy (TrainingTypePlugin) as well as certain Plugins. In particular, we want to highlight the following changes:
All TrainingTypePlugins have been renamed to Strategy (#11120). Strategy is a more appropriate name because it encompasses more than simply training communcation. This change is now aligned with the changes we implemented in 1.5, which introduced the new strategy and devices flags to the Trainer.
# Before
from pytorch_lightning.plugins import DDPPlugin
# New
from pytorch_lightning.strategies import DDPStrategy
The Accelerator and PrecisionPlugin have moved into Strategy. All strategies now take an optional parameter accelerator and precision_plugin (#11022, #10570).
Custom Accelerator implementations must now implement two new abstract methods: is_available() (#11797) and auto_device_count() (#10222). The latter determines how many devices get used by default when specifying Trainer(accelerator=..., devices="auto").
We redesigned the process creation for spawn-based strategies such as DDPSpawnStrategy and TPUSpawnStrategy (#10896). All spawn-based strategies now spawn processes immediately upon calling Trainer.{fit,validate,test,predict}, which means the hooks/callbacks prepare_data, setup, configure_sharded_model and teardown all run under an initialized process group. These changes align the spawn-based strategies with their non-spawn counterparts (such as DDPStrategy).
We've also exposed the process group backend for use. For example, you can now easily enable fairring like this:
# Explicitly specify the process group backend if you choose to
ddp = pl.strategies.DDPStrategy(process_group_backend="fairring")
trainer = Trainer(strategy=ddp, accelerator="gpu", devices=8)
In a similar fashion, if installing torch>=1.11, you can enable DDP static graph to apply special runtime optimizations:
trainer = Trainer(devices=4, strategy=DDPStrategy(static_graph=True))
LightningCLI improvementsIn the previous release, we added shorthand notation support for registered components. In this release, we added a flag to automatically register all available components:
from pytorch_lightning.utilities.cli import LightningCLI
LightningCLI(auto_registry=True)
We have also added support for the ReduceLROnPlateau scheduler with shorthand notation:
$ python script.py fit --optimizer=Adam --lr_scheduler=ReduceLROnPlateau --lr_scheduler.monitor=metric_to_track
If you need to customize the learning rate scheduler configuration, you can do so by overriding:
class MyLightningCLI(LightningCLI):
@staticmethod
def configure_optimizers(lightning_module, optimizer, lr_scheduler=None):
return {"optimizer": optimizer, "lr_scheduler": {"scheduler": lr_scheduler, ...}}
Finally, loggers are also now configurable with shorthand:
$ python script.py fit --trainer.logger=WandbLogger --trainer.logger.name="my_lightning_run"
We've added the ability to turn the automatic resubmission on or off when a job gets interrupted by the SLURM controller (via signal handling). Users who prefer to let their code handle the resubmission (for example, when submitit is used) can now pass:
from pytorch_lightning.plugins.environments import SLURMEnvironment
trainer = pl.Trainer(plugins=SLURMEnvironment(auto_requeue=False))
The Fault-tolerance training under manual optimization now tracks optimization progress. We also changed the graceful exit signal from SIGUSR1 to SIGTERM for better support inside cloud instances.
An additional feature we're excited to announce is support for consecutive trainer.fit() calls.
trainer = pl.Trainer(max_epochs=2)
trainer.fit(model)
# now, run 2 more epochs
trainer.fit_loop.max_epochs = 4
trainer.fit(model)
The Loop's state is now included as part of the checkpoints saved by the library. This enables finer restoration of custom loops.
We've also made it easier to replace Lightning's loops with your own. For example:
class MyCustomLoop(pl.loops.TrainingEpochLoop):
...
trainer = pl.Trainer(...)
trainer.fit_loop.replace(epoch_loop=MyCustomLoop)
# Trainer runs the fit loop with your new epoch loop!
trainer.fit(model)
In previous versions, Lightning required that the DataLoader instance set its input arguments as instance attributes. This meant that custom DataLoaders also had this hidden requirement. In this release, we do this automatically for the user, easing the passing of custom loaders:
class MyDataLoader(torch.utils.data.DataLoader):
def __init__(self, a=123, *args, **kwargs):
- # this was required before
- self.a = a
super().__init__(*args, **kwargs)
trainer.fit(model, train_dataloader=MyDataLoader())
As of this release, Lightning no longer pre-fetches 1 extra batch if it doesn't need to. Previously, doing so would conflict with the internal pre-fetching done by optimized data loaders such as FFCV's. You can now define your own pre-fetching value like this:
class MyCustomLoop(pl.loops.FitLoop):
@property
def prefetch_batches(self):
return 7 # lucky number 7
trainer = pl.Trainer(...)
trainer.fit_loop = MyCustomLoop(min_epochs=trainer.min_epochs, max_epochs=trainer.max_epochs)
LightningModule.lr_scheduler_stepLightning now allows the use of custom learning rate schedulers that aren't natively available in PyTorch. A great example of this is Timm Schedulers.
When using custom learning rate schedulers relying on an API other than PyTorch's, you can now define the LightningModule.lr_scheduler_step with your desired logic.
from timm.scheduler import TanhLRScheduler
class MyLightningModule(pl.LightningModule):
def configure_optimizers(self):
optimizer = ...
scheduler = TanhLRScheduler(optimizer, ...)
return {"optimizer": optimizer, "lr_scheduler": {"scheduler": scheduler, "interval": "epoch"}}
def lr_scheduler_step(self, scheduler, optimizer_idx, metric):
scheduler.step(epoch=self.current_epoch) # timm's scheduler need the epoch value
This release introduces new hooks to standardize all stateful components to use state_dict and load_state_dict, mimicking the PyTorch API. The new hooks receive their own component's state and replace most usages of the previous on_save_checkpoint and on_load_checkpoint hooks.
def MyCallback(pl.Callback):
- def on_save_checkpoint(self, trainer, pl_module, checkpoint):
- return {'x': self.x}
- def on_load_checkpoint(self, trainer, pl_module, checkpoint):
- self.x = x
+ def state_dict(self):
+ return {'x': self.x}
+ def load_state_dict(self, checkpoint):
+ self.x = x
Trainer.estimated_stepping_batchesYou can use built-in Trainer.estimated_stepping_batches to compute the total number of stepping batches needed for the complete training.
The property takes gradient accumulation factor and distributed setting into consideration when performing this computation so that you don't have to derive it manually:
class MyLightningModule(pl.LightningModule):
def configure_optimizers(self):
optimizer = ...
scheduler = torch.optim.lr_scheduler.OneCycleLR(
optimizer, max_lr=1e-3, total_steps=self.trainer.estimated_stepping_batches
)
return {"optimizer": optimizer, "lr_scheduler": scheduler}
Trainer.num_devices and Trainer.device_idsIn the past, retrieving the number of devices used, or their IDs, posed a considerable challenge. Additionally, doing so required knowing which property to access based on the current Trainer configuration.
To simplify this process, we've deprecated the per-accelerator properties to have accelerator agnostic properties. For example:
- num_devices = max(1, trainer.num_gpus, trainer.num_processes)
- if trainer.tpu_cores:
- num_devices = max(num_devices, trainer.tpu_cores)
+ num_devices = trainer.num_devices
Fault Tolerance has limitations that require specific information about your data-loading structure.
It is now possible to resolve those limitations by enabling manual fault tolerance where you can write your own logic and specify how exactly to checkpoint your own datasets and samplers. You can do so using this environment flag:
$ PL_FAULT_TOLERANT_TRAINING=MANUAL python script.py
Check out this video for a dive into the internals of this flag.
We introduced a new plugin class for wrapping layers of a model with synchronization logic for multiprocessing.
class MyLayerSync(pl.plugins.LayerSync):
...
layer_sync = MyLayerSync(...)
trainer = Trainer(sync_batchnorm=True, plugins=layer_sync, strategy="ddp")
There has been much progress in the field of ML Accelerators, and the list of accelerators is constantly expanding.
We've made it easier for users to try out new accelerators by enabling support for registering custom Accelerator classes in Lightning.
from pytorch_lightning.accelerators import Accelerator, AcceleratorRegistry
class SOTAAccelerator(Accelerator):
def __init__(self, x):
...
AcceleratorRegistry.register("sota_accelerator", SOTAAccelerator, x=123)
# the following works now:
trainer = Trainer(accelerator="sota_accelerator")
<a name="bc-changes"></a>
Here is a selection of notable changes that are not backward compatible with previous versions. The full list of changes and removals can be found in the CHANGELOG below.
Following our 4 PyTorch release window, this release supports PyTorch 1.8 to 1.11. Support for PyTorch 1.7 has been removed.
Following Python's end-of-life, support for Python 3.6 has been removed.
AcceleratorConnector rewriteTo support new accelerator and stategy features, we completely rewrote our internal AcceleratorConncetor class. No backwards compatibility was maintained so it is likely to have broken your code if it was using this class.
current_epoch boundaryTo resolve fault-tolerance issues, we changed where the current epoch value gets increased.
trainer.current_epoch is now increased by 1 on_train_end. This means that if a model is run for 3 epochs (0, 1, 2), trainer.current_epoch will now return 3 instead of 2 after trainer.fit(). This can also impact custom callbacks that acess this property inside this hook.
This also impacts checkpoints saved during an epoch (e.g. on_train_epoch_end). For example, a Trainer(max_epochs=1, limit_train_batches=1) instance that saves a checkpoint will have the current_epoch=0 value saved instead of current_epoch=1.
global_step boundaryTo resolve fault-tolerance issues, we changed where the global step value gets increased.
Access to trainer.global_step during an intra-training validation hook will now correctly return the number of optimizer steps taken already. In pseudocode:
training_step()
+ global_step += 1
validation_if_necessary()
- global_step += 1
Saved checkpoints that use the global step value as part of the filename are now increased by 1 for the same reason. A checkpoint saved after 1 step will be now be named step=1.ckpt instead of step=0.ckpt.
The trainer.global_step value will now account for TBPTT or multiple optimizers. Users setting Trainer({min,max}_steps=...) under these circumstances will need to adjust their values.
training_step when using DataParallelWhen using Trainer(strategy="dp"), all the tensors returned by training_step were previously reduced to a scalar (https://github.com/PyTorchLightning/pytorch-lightning/pull/11594). This behavior was especially confusing when outputs needed to be collected into the training_epoch_end hook.
From now on, outputs are no longer reduced except for the loss tensor, unless you implement training_step_end, in which case the loss won't get reduced either.
Previous versions were lenient in that the lack of GPU devices defaulted to running on CPU. This meant that users' code could be running much slower without them ever noticing that it was running on CPU.
We suggest passing Trainer(accelerator="auto") when this leniency is desired.
<a name="changelog"></a>
<details><summary>Added</summary>
MLFlowLogger (#12290)backward_passes_per_step (#11911)DETAIL log level to provide useful logs for improving monitoring and debugging of batch jobs (#11008)SLURMEnvironment(auto_requeue=True|False) to control whether Lightning handles the requeuing (#10601)_Stateful protocol to detect if classes are stateful (#10646)_FaultTolerantMode enum used to track different supported fault tolerant modes (#10645)_rotate_worker_indices utility to reload the state according the latest worker (#10647)_terminate_gracefully to all processes and add support for DDP (#10638)DataLoaders returned in the *_dataloader() methods, i.e., automatic replacement of samplers now works with custom types of DataLoader (#10680)DataLoader implementation is not well implemented and we need to reconstruct it (#10719)Loop's state by default in the checkpoint (#10784)Loop.replace to easily switch one loop for another (#10324)--lr_scheduler=ReduceLROnPlateau to the LightningCLI (#10860)LightningCLI.configure_optimizers to override the configure_optimizers return value (#10860)LightningCLI(auto_registry) flag to register all subclasses of the registerable components automatically (#12108)max_epochs in the Trainer is not set (#10700)LightningModule.configure_callbacks without wrapping it into a list (#11060)console_kwargs for RichProgressBar to initialize inner Console (#10875)LightningCLI (#11533)LOGGER_REGISTRY instance to register custom loggers to the LightningCLI (#11533)Trainer arguments limit_*_batches, overfit_batches, or val_check_interval are set to 1 or 1.0 (#11950)PrecisionPlugin.teardown method (#10990)LightningModule.lr_scheduler_step (#10249)DataFetcher (#11606)optimizer.step. This can be useful for LightningLite users, manual optimization users, or users overriding LightningModule.optimizer_step (#11711)MisconfigurationException if user provided opt_idx in scheduler config doesn't match with actual optimizer index of its respective optimizer (#11247)loggers property to Trainer which returns a list of loggers provided by the user (#11683)loggers property to LightningModule which retrieves the loggers property from Trainer (#11683)CombinedLoader for the training data (#11648)DistributedSampler during validation/testing (#11479)Bagua training strategy (#11146)poptorch.DataLoader in a *_dataloader hook (#12116)rank_zero module to centralize utilities (#11747)_Stateful support for LightningDataModule (#11637)_Stateful support for PrecisionPlugin (#11638)Accelerator.is_available to check device availability (#11797)Trainer (#11888)nn.Module with save_hyperparameters() (#12068)estimated_stepping_batches property to Trainer (#11599)on_load_checkpoint/on_save_checkpoint callback and LightningModule hooks (#12149)LayerSync and NativeSyncBatchNorm plugins (#11754)storage_options argument to Trainer.save_checkpoint() to pass to custom CheckpointIO implementations (#11891)device_ids and num_devices property to Trainer (#12151)Callback.state_dict() and Callback.load_state_dict() methods (#12232)AcceleratorRegistry (#12180)apply_to_collections (#11889)</details>
<details><summary>Changed</summary>
benchmark flag optional and set its value based on the deterministic flag (#11944)_print_results method of the EvaluationLoop (#11332)EvaluationLoop (#12427)prog_bar flag to False in LightningModule.log_grad_norm (#11472)init_dist_connection() when torch distributed is not available (#10418)monitor argument in the EarlyStopping callback is no longer optional (#10328)MisconfigurationException when enable_progress_bar=False and a progress bar instance has been passed in the callback list (#10520)trainer.connectors.env_vars_connector._defaults_from_env_vars to utilities.argsparse._defaults_from_env_vars (#10501)LightningCLI required for the new major release of jsonargparse v4.0.0 (#10426)refresh_rate_per_second parameter to refresh_rate for RichProgressBar signature (#10497)PrecisionPlugin into TrainingTypePlugin and updated all references (#10570)signal.SIGTERM to gracefully exit instead of signal.SIGUSR1 (#10605)Loop.restarting=... now sets the value recursively for all subloops (#11442)batch_size cannot be inferred from the current batch if it contained a string or was a custom batch object (#10541)overfit_batches > 0 is set in the Trainer (#9709)Accelerator to TrainingTypePlugin (#10596)Trainer to the Strategy (#11444)batch_to_device method from Accelerator to TrainingTypePlugin (#10649)DDPSpawnPlugin no longer overrides the post_dispatch plugin hook (#10034)LightningModule.{add_to_queue,get_from_queue} hooks no longer get a torch.multiprocessing.SimpleQueue and instead receive a list based queue (#10034)training_step, validation_step, test_step and predict_step method signatures in Accelerator and updated input from caller side (#10908)DDPSpawnPlugin and related plugins save (#10934)LoggerCollection returns only unique logger names and versions (#10976)DDPSpawnPlugin, TPUSpawnPlugin, etc.) (#10896)
Trainer.{fit,validate,test,predict}prepare_data, setup, configure_sharded_model and teardown now run under initialized process group for spawn-based plugins just like their non-spawn counterpartsMisconfigurationExceptions will now be raised as ProcessRaisedException (torch>=1.8) or as Exception (torch<1.8)TrainingTypePlugin.pre_dispatch() method and merged it with TrainingTypePlugin.setup() (#11137)batch_to_device entry in profiling from stage-specific to generic, to match profiling of other hooks (#11031)NeptuneLogger (#11015)__getstate__ and __setstate__ of RichProgressBar (#11100)DDPPlugin and DDPSpawnPlugin and their subclasses now remove the SyncBatchNorm wrappers in teardown() to enable proper support at inference after fitting (#11078)Accelerator instance to the TrainingTypePlugin; all training-type plugins now take an optional parameter accelerator (#11022)TrainingTypePlugin to Strategy (#11120)
ParallelPlugin to ParallelStrategy (#11123)DataParallelPlugin to DataParallelStrategy (#11183)DDPPlugin to DDPStrategy (#11142)DDP2Plugin to DDP2Strategy (#11185)DDPShardedPlugin to DDPShardedStrategy (#11186)DDPFullyShardedPlugin to DDPFullyShardedStrategy (#11143)DDPSpawnPlugin to DDPSpawnStrategy (#11145)DDPSpawnShardedPlugin to DDPSpawnShardedStrategy (#11210)DeepSpeedPlugin to DeepSpeedStrategy (#11194)HorovodPlugin to HorovodStrategy (#11195)TPUSpawnPlugin to TPUSpawnStrategy (#11190)IPUPlugin to IPUStrategy (#11193)SingleDevicePlugin to SingleDeviceStrategy (#11182)SingleTPUPlugin to SingleTPUStrategy (#11182)TrainingTypePluginsRegistry to StrategyRegistry (#11233)ResultCollection, ResultMetric, and ResultMetricCollection classes as protected (#11130)trainer.checkpoint_connector as protected (#11550)FitLoop instead of the TrainingEpochLoop (#11201)Strategy classes to the strategies directory (#11226)training_type_plugin file to strategy (#11239)DeviceStatsMonitor to group metrics based on the logger's group_separator (#11254)UserWarning if evaluation is triggered with best ckpt and trainer is configured with multiple checkpoint callbacks (#11274)Trainer.logged_metrics now always contains scalar tensors, even when a Python scalar was logged (#11270)MisconfigurationException to ModuleNotFoundError when rich isn't available (#11360)trainer.current_epoch value is now increased by 1 during and after on_train_end (#8578)trainer.global_step value now accounts for multiple optimizers and TBPTT splits (#11805)trainer.global_step value is now increased right after the optimizer.step() call which will impact users who access it during an intra-training validation hook (#11805)ModelCheckpoint(filename='{step}') is different compared to previous versions. A checkpoint saved after 1 step will be named step=1.ckpt instead of step=0.ckpt (#11805)ABC for Accelerator: Users need to implement auto_device_count (#11521)parallel_devices property in ParallelStrategy to be lazy initialized (#11572)TQDMProgressBar to run a separate progress bar for each eval dataloader (#11657)SimpleProfiler(extended=False) summary based on mean duration for each hook (#11671)shuffle=False for eval dataloaders (#11575)training_step_end is overridden (#11594)training_epoch_end hook will no longer receive reduced outputs from training_step and instead get the full tensor of results from all GPUs (#11594)lightning_logs for consistency (#11762)accelerator_connector (#11448)find_unused_parameters=True (#12425)limit_batches=0 (#11576)is_global_zero check in training_epoch_loop before logger.save. If you have a custom logger that implements save the Trainer will now call save on all ranks by default. To change this behavior add @rank_zero_only to your save implementation (#12134)trainer.logger_connector as protected (#12195)Strategy.process_dataloader function call from fit/evaluation/predict_loop.py to data_connector.py (#12251)ModelCheckpoint(save_last=True, every_n_epochs=N) now saves a "last" checkpoint every epoch (disregarding every_n_epochs) instead of only once at the end of training (#12418)sync_batchnorm now only apply it when fitting (#11919)supporters.py so that in the accumulator element (for loss) is created directly on the device (#12430)EarlyStopping.on_save_checkpoint and EarlyStopping.on_load_checkpoint in favor of EarlyStopping.state_dict and EarlyStopping.load_state_dict (#11887)BaseFinetuning.on_save_checkpoint and BaseFinetuning.on_load_checkpoint in favor of BaseFinetuning.state_dict and BaseFinetuning.load_state_dict (#11887)BackboneFinetuning.on_save_checkpoint and BackboneFinetuning.on_load_checkpoint in favor of BackboneFinetuning.state_dict and BackboneFinetuning.load_state_dict (#11887)ModelCheckpoint.on_save_checkpoint and ModelCheckpoint.on_load_checkpoint in favor of ModelCheckpoint.state_dict and ModelCheckpoint.load_state_dict (#11887)Timer.on_save_checkpoint and Timer.on_load_checkpoint in favor of Timer.state_dict and Timer.load_state_dict (#11887)</details>
<details><summary>Deprecated</summary>
training_type_plugin property in favor of strategy in Trainer and updated the references (#11141)Trainer.{validated,tested,predicted}_ckpt_path and replaced with read-only property Trainer.ckpt_path set when checkpoints loaded via Trainer.{fit,validate,test,predict} (#11696)ClusterEnvironment.master_{address,port} in favor of ClusterEnvironment.main_{address,port} (#10103)DistributedType in favor of _StrategyType (#10505)precision_plugin constructor argument from Accelerator (#10570)DeviceType in favor of _AcceleratorType (#10503)Trainer.slurm_job_id in favor of the new SLURMEnvironment.job_id() method (#10622)IndexBatchSamplerWrapper.batch_indices in favor of IndexBatchSamplerWrapper.seen_batch_indices (#10870)on_init_start and on_init_end callback hooks (#10940)Trainer.call_hook in favor of Trainer._call_callback_hooks, Trainer._call_lightning_module_hook, Trainer._call_ttp_hook, and Trainer._call_accelerator_hook (#10979)TrainingTypePlugin.post_dispatch in favor of TrainingTypePlugin.teardown (#10939)ModelIO.on_hpc_{save/load} in favor of CheckpointHooks.on_{save/load}_checkpoint (#10911)Trainer.run_stage in favor of Trainer.{fit,validate,test,predict} (#11000)Trainer.lr_schedulers in favor of Trainer.lr_scheduler_configs which returns a list of dataclasses instead of dictionaries (#11443)Trainer.verbose_evaluate in favor of EvaluationLoop(verbose=...) (#10931)Trainer.should_rank_save_checkpoint Trainer property (#11068)Trainer.lightning_optimizers (#11444)TrainerOptimizersMixin and moved functionality to core/optimizer.py(#11155)on_train_batch_end(outputs) format when multiple optimizers are used and TBPTT is enabled (#12182)training_epoch_end(outputs) format when multiple optimizers are used and TBPTT is enabled (#12182)TrainerCallbackHookMixin (#11148)TrainerDataLoadingMixin and moved functionality to Trainer and DataConnector (#11282)pytorch_lightning.callbacks.device_stats_monitor.prefix_metric_keys (#11254)Callback.on_epoch_start hook in favour of Callback.on_{train/val/test}_epoch_start (#11578)Callback.on_epoch_end hook in favour of Callback.on_{train/val/test}_epoch_end (#11578)LightningModule.on_epoch_start hook in favor of LightningModule.on_{train/val/test}_epoch_start (#11578)LightningModule.on_epoch_end hook in favor of LightningModule.on_{train/val/test}_epoch_end (#11578)on_before_accelerator_backend_setup callback hook in favour of setup (#11568)on_batch_start and on_batch_end callback hooks in favor of on_train_batch_start and on_train_batch_end (#11577)on_configure_sharded_model callback hook in favor of setup (#11627)pytorch_lightning.utilities.distributed.rank_zero_only in favor of pytorch_lightning.utilities.rank_zero.rank_zero_only (#11747)pytorch_lightning.utilities.distributed.rank_zero_debug in favor of pytorch_lightning.utilities.rank_zero.rank_zero_debug (#11747)pytorch_lightning.utilities.distributed.rank_zero_info in favor of pytorch_lightning.utilities.rank_zero.rank_zero_info (#11747)pytorch_lightning.utilities.warnings.rank_zero_warn in favor of pytorch_lightning.utilities.rank_zero.rank_zero_warn (#11747)pytorch_lightning.utilities.warnings.rank_zero_deprecation in favor of pytorch_lightning.utilities.rank_zero.rank_zero_deprecation (#11747)pytorch_lightning.utilities.warnings.LightningDeprecationWarning in favor of pytorch_lightning.utilities.rank_zero.LightningDeprecationWarningon_pretrain_routine_start and on_pretrain_routine_end callback hooks in favor of on_fit_start (#11794)LightningModule.on_pretrain_routine_start and LightningModule.on_pretrain_routine_end hooks in favor of on_fit_start (#12122)agg_key_funcs and agg_default_func parameters from LightningLoggerBase (#11871)LightningLoggerBase.update_agg_funcs (#11871)LightningLoggerBase.agg_and_log_metrics in favor of LightningLoggerBase.log_metrics (#11832)weights_save_path to the Trainer constructor in favor of adding the ModelCheckpoint callback with dirpath directly to the list of callbacks (#12084)pytorch_lightning.profiler.AbstractProfiler in favor of pytorch_lightning.profiler.Profiler (#12106)pytorch_lightning.profiler.BaseProfiler in favor of pytorch_lightning.profiler.Profiler (#12150)BaseProfiler.profile_iterable (#12102)LoggerCollection in favor of trainer.loggers (#12147)PrecisionPlugin.on_{save,load}_checkpoint in favor of PrecisionPlugin.{state_dict,load_state_dict} (#11978)LightningDataModule.on_save/load_checkpoint in favor of state_dict/load_state_dict (#11893)Trainer.use_amp in favor of Trainer.amp_backend (#12312)LightingModule.use_amp in favor of Trainer.amp_backend (#12315)PL_TORCH_DISTRIBUTED_BACKEND (#11745)ParallelPlugin.torch_distributed_backend in favor of DDPStrategy.process_group_backend property (#11745)ModelCheckpoint.save_checkpoint in favor of Trainer.save_checkpoint (#12456)Trainer.devices in favor of Trainer.num_devices and Trainer.device_ids (#12151)Trainer.root_gpu in favor of Trainer.strategy.root_device.index when GPU is used (#12262)Trainer.num_gpus in favor of Trainer.num_devices when GPU is used (#12384)Trainer.ipus in favor of Trainer.num_devices when IPU is used (#12386)Trainer.num_processes in favor of Trainer.num_devices (#12388)Trainer.data_parallel_device_ids in favor of Trainer.device_ids (#12072)Callback.on_save_checkpoint in favor of returning state in Callback.state_dict for checkpointing (#11887)Callback.on_load_checkpoint(callback_state) in favor of passing the callback state to Callback.load_state_dict and in 1.8, passing the entire checkpoint dictionary to Callback.on_load_checkpoint(checkpoint) (#11887)Trainer.gpus in favor of Trainer.device_ids or Trainer.num_devices (#12436)Trainer.tpu_cores in favor of Trainer.num_devices (#12437)</details>
<details><summary>Removed</summary>
method in pytorch_lightning.utilities.model_helpers.is_overridden (#10507)ClusterEnvironment.creates_children (#10339)TrainerModelHooksMixin.is_function_implemented and TrainerModelHooksMixin.has_arg (#10322)pytorch_lightning.utilities.device_dtype_mixin.DeviceDtypeModuleMixin in favor of pytorch_lightning.core.mixins.device_dtype_mixin.DeviceDtypeModuleMixin (#10442)LightningModule.loaded_optimizer_states_dict property (#10346)Trainer.fit(train_dataloader=), Trainer.validate(val_dataloaders=), and Trainer.test(test_dataloader=) (#10325)has_prepared_data, has_setup_fit, has_setup_validate, has_setup_test, has_setup_predict, has_teardown_fit, has_teardown_validate, has_teardown_test and has_teardown_predict datamodule lifecycle properties (#10350)every_n_val_epochs parameter of ModelCheckpoint (#10366)import pytorch_lightning.profiler.profilers in favor of import pytorch_lightning.profiler (#10443)configure_slurm_dpp from accelerator connector (#10370)num_nodes and sync_batchnorm from DDPPlugin, DDPSpawnPlugin, DeepSpeedPlugin (#10357)is_slurm_managing_tasks from AcceleratorConnector (#10353)LightningModule.log(tbptt_reduce_fx, tbptt_reduce_token, sync_dist_op) (#10423)Plugin.task_idx (#10441)master_params from PrecisionPlugin (#10372)training_step. For example, return {'loss': ..., 'foo': foo.detach()} will now be necessary if foo has gradients which you do not want to store (#10424)Accelerator base class:
transfer_batch_to_device hook. The new argument dataloader_idx is now required (#10480)utilities.distributed.rank_zero_{warn/deprecation} (#10451)mode argument from ModelSummary class (#10449)Trainer.train_loop property in favor of Trainer.fit_loop (#10482)Trainer.train_loop property in favor of Trainer.fit_loop (#10482)disable_validation property from Trainer (#10450)CheckpointConnector.hpc_load property in favor of CheckpointConnector.restore (#10525)reload_dataloaders_every_epoch from Trainer in favour of reload_dataloaders_every_n_epochs (#10481)precision_plugin attribute from Accelerator in favor of its equivalent attribute precision_plugin in the TrainingTypePlugin (#10570)DeepSpeedPlugin.{precision,amp_type,amp_level} properties (#10657)on_before_batch_transfer, transfer_batch_to_device and on_after_batch_transfer hooks in LightningModule (#10603)return_result from the DDPSpawnPlugin.spawn() method (#10867)TrainingTypePlugin.results and corresponding properties in subclasses (#10034)mp_queue attribute from DDPSpawnPlugin and TPUSpawnPlugin (#10034)_move_optimizer_state method overrides from TPUSpawnPlugin and SingleTPUPlugin (#10849)should_rank_save_checkpoint property from TrainingTypePlugin (#11070)model_sharded_context method from Accelerator (#10886)pre_dispatch from the PrecisionPlugin (#10887)setup_optimizers_in_pre_dispatch from the strategies and achieve the same logic in setup and pre_dispatch methods (#10906)pre_dispatch, dispatch and post_dispatch from the Accelerator (#10885)training_step, test_step, validation_step and predict_step from the Accelerator (#10890)TrainingTypePlugin.start_{training,evaluating,predicting} hooks and the same in all subclasses (#10989, #10896)Accelerator.on_train_start (#10999)Strategy.init_optimizers in favor of Strategy.setup_optimizers (#11236)profile("training_step_and_backward") in Closure class since we already profile calls training_step and backward (#11222)Strategy.optimizer_zero_grad (#11246)Strategy.on_gpu (#11537)Strategy.on_tpu property (#11536)LightningLoggerBase.experiment (#11603)FitLoop.current_epoch getter and setter (#11562)_short_id in NeptuneLogger (#11517)log_text and log_image from the LightningLoggerBase API (#11857)profile("model_forward") in favor of profiling training_step (#12032)get_mp_spawn_kwargs from DDPSpawnStrategy and TPUSpawnStrategy in favor of configuration in the _SpawnLauncher (#11966)_aggregate_metrics, _reduce_agg_metrics, and _finalize_agg_metrics from LightningLoggerBase (#12053)AcceleratorConnector.device_type property (#12081)AcceleratorConnector.num_nodes (#12107)AcceleratorConnector.has_ipu property (#12111)AcceleratorConnector.use_ipu property (#12110)AcceleratorConnector.has_tpu property (#12109)AcceleratorConnector.use_dp property (#12112)configure_sync_batchnorm from ParallelStrategy and all other strategies that inherit from it (#11754)sync_batchnorm from strategies (#11754)AcceleratorConnector.root_gpu property (#12262)AcceleratorConnector.tpu_id property (#12387)AcceleratorConnector.num_gpus property (#12384)AcceleratorConnector.num_ipus property (#12386)AcceleratorConnector.num_processes property (#12388)AcceleratorConnector.parallel_device_ids property (#12072)AcceleratorConnector.devices property (#12435)AcceleratorConnector.parallel_devices property (#12075)AcceleratorConnector.tpu_cores property (#12437)</details>
<details><summary>Fixed</summary>
ModelCheckpoint could delete last checkpoint from the old directory when dirpath has changed during resumed training (#12225)ModelCheckpoint could delete older checkpoints when dirpath has changed during resumed training (#12045)HorovodStrategy.teardown() did not complete gracefully if an exception was thrown during callback setup #11752PyYAML dependency (#11099){test,validation}_epoch_end with multiple dataloaders (#11132)TPUSpawnPlugin handling the XLA_USE_BF16 environment variable incorrectly (#10990)Trainer.lightning_optimizers (#11155)configure_optimizers with the CLI (#11672)SimpleProfiler summary (#11414)DistributedSampler to the poptorch.DataLoader when IPUs are used (#12114)LightningModule.{un,}toggle_model when only 1 optimizer is used (#12088)RichProgressbar to display the metrics logged only on main progress bar (#11690)RichProgressBar progress when refresh rate does not evenly divide the total counter (#11668)RichProgressBar progress validation bar total when using multiple validation runs within a single training epoch (#11668)RichProgressBarTheme styles after detecting light theme on colab (#10993)_ddp_params_and_buffers_to_ignore (#11949)AttributeError when calling save_hyperparameters and no parameters need saving (#11827)logger=None is passed to the Trainer (#12249)ModelCheckpoint was still set even if no checkpoint was saved (#12418)ModelCheckpoint was overriding the epoch and step logged values (#12418)epoch and step values with ModelCheckpoint would fail (#12418)DDPFullyShardedStrategy (#12267)</details>
Full commit list: https://github.com/PyTorchLightning/pytorch-lightning/compare/1.5.0...1.6.0
<a name="contributors"></a>
@akihironitta @ananthsub @awaelchli @Borda @borisdayma @carmocca @daniellepintz @edward-io @ethanwharris @four4fish @jjenniferdai @kaushikb11 @kingyiusuen @kragniz @mauvilsa @ninginthecloud @popfido @rohitgr7 @SeanNaren @speediedan @tchaton @tshu-w @twsl @williamFalcon
@a-gardner1 @abhi-rf @abhinavarora @adamreeve @adamviola @AJSVB @akashkw @amin-nejad @AndresAlgaba @ant0nsc @armanal @bhadreshpsavani @CAIQT @catalys1 @chaddy1004 @chunyang-wen @circlecrystal @Code-Cornelius @Cyber-Machine @dennisbappert @DuYicong515 @edpizzi @franp9am @ftorres16 @ggare-cmu @guyang3532 @Honzys @idiomaticrefactoring @isvogor-foi @jerome-habana @jgibson2 @jlhbaseball15 @jona-0 @JoostvDoorn @josafatburmeister @konstantinjdobler @Kr4is @krishnakalyan3 @krshrimali @lemairecarl @lucmos @manangoel99 @mathemusician @mayeroa @mbortolon97 @NathanGodey @Nesqulck @nithinraok @ORippler @os1ma @peterdudfield @Piyush-97 @puhuk @qqueing @quancs @Raahul-Singh @Raalsky @Rajathbharadwaj @rasbt @rharish101 @rhjohnstone @rjkilpatrick @RobertLaurella @roschly @rsokl @rusty1s @SauravMaheshkar @sethvargo @shabie @shivammehta007 @srb-cv @ThomVett @wangraying @whokilleddb @zredeaux65
If we forgot someone or have any suggestion, let us know in Slack :zap:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixed an issue to avoid validation loop run on restart
on_epoch logged values on train epoch end (#11689)step argument in WandbLogger.log_image work (#11716)restore_optimizers for mapping states (#11757)DPStrategy, the batch is not explicitly moved to the device (#11780)trainer.validate() (#11700)Trainer.weights_save_path for fault-tolerant training (#11776)@ananthsub @borda @circlecrystal @NathanGodey @nithinraok @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Pinned sphinx-autodoc-typehints with <v1.15
SaveConfigCallback (#11532)LSFEnvironment to use LSB_DJOB_RANKFILE environment variable instead of LSB_HOSTS for determining node rank and main address (#10825)IterableDataset (#11507)@ajtritt @akihironitta @carmocca @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed LightningCLI race condition while saving the config
LightningCLI race condition while saving the config (#11199)log(reduce_fx=min|max) (#11310)reload_dataloaders_every_n_epochs and check_val_every_n_epoch (#10948)@adamviola @akihironitta @awaelchli @Borda @carmocca @edpizzi
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Avoid the deprecated onnx.export(example_outputs=...) in torch 1.10
NeptuneLogger when using DDP (#11030)onnx.export(example_outputs=...) in torch 1.10 (#11116)LightningModule after training with Trainer(sync_batchnorm=True) (#11078)AttributeError occuring when using a CombinedLoader (multiple dataloaders) for prediction (#11111)Trainer(track_grad_norm=..., logger=False) would fail (#11114)bf16 precision on CPU (#11161)ModelCheckpoint callback now saves and restores attributes best_k_models, kth_best_model_path, kth_value, and last_model_path (#10995)@awaelchli @borchero @carmocca @guyang3532 @kaushikb11 @ORippler @Raalsky @rohitgr7 @SeanNaren
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed a bug where the DeepSpeedPlugin arguments cpu_checkpointing and contiguous_memory_optimization were not being forwarded to deepspeed correctly
cpu_checkpointing and contiguous_memory_optimization were not being forwarded to deepspeed correctly (#10874)NeptuneLogger causing checkpoints to be uploaded with a duplicated file extension (#11015)LightningModule (#10991)RichProgressBar (#10913)CombinedLoader while checking for warning raised with eval dataloaders (#10994)on_epoch logged values on train epoch end (#11069)trainer.validate (#11069)@carmocca @jona-0 @kaushikb11 @Raalsky @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Disabled batch_size extraction for torchmetric instances because they accumulate the metrics internally
SignalConnector not restoring the default signal handlers on teardown when running on SLURM or with fault-tolerant training enabled (#10611)SignalConnector._has_already_handler check for callable type (#10483)rich version is less than 10.2.2 (#10839)BasePredictionWriter hooks when using a dataloader with num_workers > 0 (#10870)torch_xla.debug for torch-xla<1.8 (#10836)DDPSpawnPlugin and related plugins leaving a temporary checkpoint behind (#10934)TypeError occuring in the SingalConnector.teardown() method (#10961)@awaelchli @carmocca @four4fish @kaushikb11 @lucmos @mauvilsa @Raalsky @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed support for --key.help=class with the LightningCLI
--key.help=class with the LightningCLI (#10767)_compare_version for python packages (#10762)SummaryWriter not close before spawning the processes (#10777)on_step=False, on_epoch=True to on_step=True, on_epoch=False (#10756)@awaelchli @carmocca @kaushikb11 @rohitgr7 @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed ShardedTensor state dict hook registration to check if torch distributed is available
ShardedTensor state dict hook registration to check if torch distributed is available (#10621)self.log not respecting a tensor's dtype when applying computations (#10076)_wrap_init popping unexisting keys from DataLoader signature parameters (#10613)LightningModule.log (#10408)Trainer(move_metrics_to_cpu=True) not moving the evaluation logged results to CPU (#10631){validation,test}_step outputs getting moved to CPU with Trainer(move_metrics_to_cpu=True) (#10631)@ananthsub @awaelchli @carmocca @jiwidi @kaushikb11 @qqueing @rohitgr7 @shabie @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed scripting causing false positive deprecation warnings (#10470, #10555)
CombinedLoader and max_size_cycle didn't receive a DistributedSampler (#10374)utilities.apply_to_collection (#9702)isinstance not working with init_meta_context, materialized model not being moved to the device (#10493)overfit_batches to only replace the sample when SequentialSampler is not used (#10486)DeviceDtypeModuleMixin (#10559)@a-gardner1 @awaelchli @carmocca @justusschock @Raahul-Singh @rohitgr7 @SeanNaren @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed apply_to_collection(defaultdict)
apply_to_collection(defaultdict) (#10316)DataLoader(batch_size=None) is passed (#10345)__init__ arguments for sub-classed DataLoader re-instantiation in Lite (#10334)CSVLogger after a call to CSVLogger.save (#10388)PostLocalSGD when torch.distributed not available (#10359)on_step=True in epoch-level hooks causing unintended side-effects. Logging with on_step=True in epoch-level hooks will now correctly raise an error (#10409)RichProgressBar (#10428)persistent_workers being deleted on every iteration (#10434)@EspenHa @four4fish @peterdudfield @rohitgr7 @tchaton @kaushikb11 @awaelchli @Borda @carmocca
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Changes in Accelerators and Plugins were made without deprecation due to their experimental state. The API is expected to become stable in 1.6.
The PyTorch Lightning team and its community are excited to announce Lightning 1.5, introducing support for LightningLite, Fault-tolerant Training, Loop Customization, Lightning Tutorials, LightningCLI V2, RichProgressBar, CheckpointIO Plugin, Trainer Strategy flag, and more!
<a name="highlights"></a>
Lightning 1.5 marks our biggest release yet. Over 60 contributors have worked on features, bugfixes and documentation improvements for a total of 640 commits since v1.4. Here are some highlights:
Fault-tolerant Training is a new internal mechanism that enables PyTorch Lightning to recover from a hardware or software failure. This is particularly interesting while training in the cloud with preemptive instances which can shutdown at any time. Once a Lightning experiment unexpectedly exits, a temporary checkpoint is saved that contains the exact state of all loops and the model. With this new experimental feature, you will be able to restore your training mid-epoch on the exact batch and continue training as if it never got interrupted.
PL_FAULT_TOLERANT_TRAINING=1 python train.py
LightningLite enables pure PyTorch users to scale their existing code to any kind of hardware while retaining full control over their own loops and optimization logic.
With just a few lines of code and no large refactoring, you get support for multi-device, multi-node, running on different accelerators (CPU, GPU, TPU), native automatic mixed precision (half and bfloat16), and double precision, in just a few seconds. And no special launcher required! Check out our documentation to find out how you can get one step closer to boilerplate-free research!
class Lite(LightningLite):
def run(self):
# Let Lite setup your dataloader(s)
train_loader = self.setup_dataloaders(torch.utils.data.DataLoader(...))
model = Net() # .to() not needed
optimizer = optim.Adam(model.parameters())
# Let Lite setup your model and optimizer
model, optimizer = self.setup(model, optimizer)
for epoch in range(5):
for data, target in train_loader:
optimizer.zero_grad()
output = model(data) # data is already on the device
loss = F.nll_loss(output, target)
self.backward(loss) # instead of loss.backward()
optimizer.step()
Lite(accelerator="gpu", devices="auto").run()
The new Loop API lets advanced users swap out the default gradient descent optimization loop at the core of Lightning with a different optimization paradigm. This is part of our effort to make Lightning the simplest, most flexible framework to take any kind of deep learning research to production.
Read our comprehensive introduction to loops
We integrated with Rich and created a new and improved progress bar for Lightning. Try it out:
pip install rich
from pytorch_lightning import Trainer
from pytorch_lightning.callbacks import RichProgressBar
trainer = Trainer(callbacks=[RichProgressBar()])
<a name="strategy-and-devices"></a>
With the new strategy and devices arguments in the Trainer, it is now easer to switch from one hardware to another.
<div align="center">
| Before | After |
|---|---|
Trainer(accelerator="ddp", gpus=2) |
Trainer(accelerator="gpu", devices=2, strategy="ddp") |
Trainer(accelerator="ddp_cpu", num_processes=2) |
Trainer(accelerator="cpu", devices=2, strategy="ddp") |
Trainer(accelerator="tpu_spawn", tpu_cores=8) |
Trainer(accelerator="tpu", devices=8) |
</div>
The new devices argument is now agnostic to all accelerators, but the previous arguments gpus, tpu_cores, ipus are still available and work the same as before. In addition, it is now also possible to set devices="auto" or accelerator="auto" to select the best accelerator available on the hardware.
from pytorch_lightning import Trainer
trainer = Trainer(accelerator="auto", devices="auto")
This release adds support for running not just Trainer.fit but any of the Trainer entry points!
python script.py fit
python script.py test
LightningCLI now supports registries for callbacks, optimizers, learning rate schedulers, LightningModules and LightningDataModules. This greatly improves the command line experience as only the class names and arguments are required as follows:
python script.py \
--trainer.callbacks=EarlyStopping \
--trainer.callbacks.patience=5 \
--trainer.callbacks.LearningRateMonitor \
--trainer.callbacks.logging_interval=epoch \
--optimizer=Adam \
--optimizer.lr=0.01 \
--lr_scheduler=OneCycleLR \
--lr_scheduler=anneal_strategy=linear
We've also added support for a manual mode where the CLI takes care of the instantiation but you have control over the Trainer calls:
cli = LightningCLI(MyModel, run=False)
cli.trainer.fit(cli.model)
As part of our commitment to extensibility, we have abstracted the checkpointing logic into a CheckpointIO plugin. This enables users to adapt Lightning to their own infrastructure.
from pytorch_lightning.plugins import CheckpointIO
class CustomCheckpointIO(CheckpointIO):
def save_checkpoint(self, checkpoint, path):
# put all logic related to saving a checkpoint here
def load_checkpoint(self, path):
# put all logic related to loading a checkpoint here
def remove_checkpoint(self, path):
# put all logic related to deleting a checkpoint here
PyTorch 1.10 introduces native Automatic Mixed Precision (AMP) support for torch.bfloat16 on CPU (was already supported for TPUs), enabling higher performance compared with torch.float16. Switch to bfloat16 training by setting the argument:
from pytorch_lightning import Trainer
trainer = Trainer(precision="bf16")
It is pretty common to share parameters within a model. However, TPUs don't retain shared parameters once moved on the devices. Lightning now supports automatic detection and re-assignement to alleviate this problem from TPUs.
Infinite training is now supported by setting Trainer(max_epochs=-1) for an unlimited number of epochs, or Trainer(max_steps=-1) for an endless epoch.
Note: you will want to avoid logging with
on_epoch=Truein case ofmax_steps=-1.
DeepSpeed is a deep learning training optimization library, providing the means to train massive billion parameter models at scale. Lightning now also supports the DeepSpeed ZeRO Stage 1 protocol that partitions your optimizer states across your GPUs to reduce memory.
from pytorch_lightning import Trainer
trainer = Trainer(gpus=4, strategy="deepspeed_stage_1", precision=16)
trainer.fit(model)
For even more memory savings and model sharding advice, check out stage 2 & 3 as well in our multi-GPU docs.
By overriding the LightningModule.configure_gradient_clipping hook, you can customize gradient clipping to your needs:
# Perform gradient clipping on gradients associated with discriminator (optimizer_idx=1) in GAN
def configure_gradient_clipping(
self,
optimizer,
optimizer_idx,
gradient_clip_val,
gradient_clip_algorithm
):
if optimizer_idx == 1:
# Lightning will handle the gradient clipping
self.clip_gradients(
optimizer,
gradient_clip_val=gradient_clip_val,
gradient_clip_algorithm=gradient_clip_algorithm
)
This means you can now implement state-of-the-art clipping algorithms with Lightning!
Added support for torch.use_deterministic_algorithms. Read more about how it works here. You can enable it by setting:
from pytorch_lightning import Trainer
trainer = Trainer(deterministic=True)
Lightning makes it easier to debug your code, so we've added support for torch.set_detect_anomaly. With this, PyTorch detects numerical anomalies like NaN or inf during forward and backward. Read more about anomaly detection here
from pytorch_lightning import Trainer
trainer = Trainer(detect_anomaly=True)
Are you having a hard time debugging DDP on your remote machine? Now you can debug DDP locally on the CPU:
trainer = Trainer(accelerator="cpu", strategy="ddp", devices=2)
When everything works, switch back to GPU by changing only the accelerator. Check our documentation for more useful debugging tricks.
Note that this will not provide any speed benefits.
Generates a summary of all layers in a LightningModule. This currently works with the new RichProgressBar callback.
from pytorch_lightning import Trainer
from pytorch_lightning.callbacks import ModelSummary
trainer = Trainer(callbacks=[ModelSummary(max_depth=1)])
An on_exception Callback hook has been added which allows the user to perform custom exception handling.
class MyCallback(Callback):
def on_exception(self, trainer, pl_module, exception):
# whatever you want!
...
The inter-batch parallelism feature aims at hiding the latency of host-to-device copy of input batches behind computationally intensive operations. In some use case, it can provide training speed up. This feature is experimental and subject to change, hence opt-in through an environment variable.
PL_INTER_BATCH_PARALLELISM=1 python train.py
If your training_step signature takes a dataloader_iter, Lightning would pass it directly. This can be useful for recommendation engine optimization.
PyTorch 1.10 introduces the meta tensors, tensors without the data. In this continuation, PyTorch Lightning provides an init_meta_context context manager and materialize_module function to handle large sharded models.
<a name="bc-changes"></a>
Here is a selection of important changes that are not backward compatible with versions < 1.5. The full list of changes and removals are listed in the changelog at the bottom.
The interpretation of the gpus Trainer argument when provided as a string has changed: Trainer(gpus="n") (string) no longer selects the GPU index n and instead selects the first n devices. In order to preserve the old behavior, you will have to change your code to Trainer(gpus=[n]) (list of indices) or Trainer(gpus="n,") (string with comma separated indices).
The argument distributed_backend has been removed from the Trainer in favor of the new accelerator and strategy arguments (#10017).
# BEFORE
trainer = Trainer(distributed_backend="ddp_spawn", gpus=2)
# NOW
trainer = Trainer(strategy="ddp_spawn", accelerator="gpu", devices=2)
max_steps Trainer argument has changed from None to -1 (#9460). You can no longer specify Trainer(max_steps=None) and if you did, you need to change the code to Trainer(max_steps=-1).accumulate_grad_batches has changed from 1 to None (#9652).The model weights now get loaded in all cases when the checkpoint path is provided in Trainer.{validate,test,predict}, regardless of whether the model instance is provided or not.
# model reference provided:
trainer.test(model, ckpt_path=None) # use provided model
trainer.test(model, ckpt_path="best") # load best model
trainer.test(model, ckpt_path="my_path") # load path
# model reference not provided
trainer.fit(model)
trainer.test(ckpt_path=None) # load best model (NEW BEHAVIOR!)
trainer.test(ckpt_path="my_path") # load path (NEW BEHAVIOR!)
Users who relied on trainer.test(ckpt_path=None) to load the latest model need to change their code to trainer.test(model) and pass the model reference directly.
All CLI commands now need to include the Trainer method to run as the first command, i.e., one of fit, validate, test, predict.
# BEFORE
python script.py --trainer.max_epochs=123
# NOW
python script.py fit --trainer.max_epochs=123
For questions and help regarding CLI, join our Lightning-CLI Slack channel.
optimizer_closure is now required when overriding the optimizer_step hook (#9360). If you relied on the previous behavior, we recommend to switch to Manual Optimization alltogether.on_before_optimizer_step hook previously ran before the entire optimization closure, including backward. This was unintended behavior and if you rely on this, move your code to the new on_before_backward` hook.Changes in Accelerators and Plugins were made without deprecation due to their experimental state. The API is expected to become stable in 1.6.
Removed attributes and methods:
Accelerator.{call_configure_sharded_model_hook, connect_training_type_plugin, connect_precision_plugin, on_reset_*_dataloader, on_train_epoch_end, on_save, post_optimizer_step, update_global_step}TrainingTypePlugin.{call_configure_sharded_model_hook, on_reset_*_dataloader, on_save, post_optimizer_step, update_global_step}PrecisionPlugin.{post_optimizer_step}ParallelPlugin.teardownChanged signatures:
setup hooks no longer have a model argument.Other changes:
Plugin class has been removed.HorovodPlugin.all_gather now returns a torch.Tensor instead of a list.DDPPlugin, DDPSpawnPlugin, DDPShardedPlugin, DDPSpawnShardedPlugin.<a name="changelog"></a>
LearningRateMonitor (#9786)ShardedTensor state dict hooks in LightningModule.__init__ if the PyTorch version supports ShardedTensor (#8944)on_keyboard_interrupt() and on_exception() for all entrypoints (fit, validate, test, predict) (#8819)training_step that takes dataloader_iter as an argument (#8807)state_key property to the Callback base class (#6886)TrainingEpochLoop.total_batch_idx (#8598)BatchProgress and integrated TrainingEpochLoop.is_last_batch (#9657)Tracker attributes (#9320)current progress counters when restarting an epoch loop that had already finished (#9371)reset_on_restart in the loop's reset hook instead of when loading a checkpoint (#9561)completed over processed in reset_on_restart (#9656)reset_on_epoch to reset_on_run (#9658)batch_size and rank_zero_only arguments for log_dict to match log (#8628)ResultCollection state_dict to the Loop state_dict and added support for distributed reload (#8641)handles_accumulate_grad_batches property to the training type plugins (#8856)WandbLogger when reusing a wandb run (#8714)log_graph argument for watch method of WandbLogger (#8662)LightningCLI additions:
LightningCLI(run=False|True) to choose whether to run a Trainer subcommand (#8751)LightningCLI via subcommands (#7508)multifile option to LightningCLI to enable/disable config saving to preserve multiple files structure (#9073)FastForwardSampler and CaptureIterableDataset injection to data loading utilities (#8366)DataFetcher to control fetching flow (#8890)SharedCycleIteratorState to prevent infinite loop (#8889)CaptureMapDataset for state management in map-style datasets (#8891)DataFetcher (#8891)DataFetcher in training loop (#8953)CheckpointIO plugin to expose checkpoint IO from training type plugin (#8743)CheckpointConnector to offload validation logic to the CheckpointIO plugin (#9045)remove_checkpoint to CheckpointIO plugin by moving the responsibility out of the ModelCheckpoint callback (#9373)XLACheckpointIO plugin (#9972)Closure and AbstractClosure classes (#8642)TrainingBatchLoop and extracted OptimizerLoop, splitting off automatic optimization into its own loop (#9191)TrainingBatchLoop.backward(); manual optimization now calls directly into Accelerator.backward() and automatic optimization handles backward in new OptimizerLoop (#9265)ManualOptimization logic from TrainingBatchLoop into its own separate loop class (#9266)OutputResult and ManualResult classes (#9437, #9424)OptimizerLoop.backward as protected (#9514)FitLoop.should_accumulate as protected (#9515)PredictionLoop as protected: on_predict_start, on_predict_epoch_end, on_predict_end, on_predict_model_eval (#9516)EvaluationLoop as protected: get_max_batches, on_evaluation_model_eval, on_evaluation_model_train, on_evaluation_start, on_evaluation_epoch_start, on_evaluation_epoch_end, on_evaluation_end, reload_evaluation_dataloaders (#9516)EvaluationEpochLoop as protected: on_evaluation_batch_start, evaluation_step, evaluation_step_end (#9516)yielding_training_step example (#9983)Python dataclass support for LightningDataModule (#8272)TensorBoardLogger (#9031)InterBatchParallelDataFetcher (#9020)DataLoaderIterDataFetcher (#9020)DataFetcher within Fit / Evaluation Loop (#9047)on_exception callback hook (#9183)ModelSummary callback (#9344)log_images, log_text and log_table to WandbLogger (#9545)PL_RECONCILE_PROCESS environment variable to enable process reconciliation regardless of cluster environment settings (#9389)get_device_stats to the Accelerator interface and added its implementation for GPU and TPU (#9586)OneCycleLR is used with "interval": "epoch" (#9666)DeviceStatsMonitor callback (#9712)enable_progress_bar to the Trainer constructor (#9664)pl_legacy_patch load utility for loading old checkpoints that have pickled legacy Lightning attributes (#9166)torch.use_deterministic_algorithms (#9121)torch.autograd.set_detect_anomaly through Trainer constructor argument detect_anomaly (#9848)enable_model_summary flag to Trainer (#9699)strategy argument to Trainer (#8597)init_meta_context, materialize_module utilities (#9920)TPUPrecisionPlugin (#10020)torch.bfloat16 support:
kfold example for loop customization (#9965)PrecisionPlugin.forward_context, making it the default implementation for all {train,val,test,predict}_step_context() methods (#9988)DDPSpawnPlugin.spawn() for spawning new processes of a given function (#10018, #10022)TrainingTypePlugin.{_setup_model, _setup_optimizer} methods (#9994, #10064)DataParallelPlugin._setup_model (#10010)DeepSpeedPlugin._setup_model_and_optimizers (#10009, #10064){DDPShardedPlugin,DDPShardedSpawnPlugin}._setup_model_and_optimizers (#10028, #10064)model argument to the optimizer_step methods in accelerators and plugins (#10023)DeepSpeedPlugin (#10164)DDPSpawnPlugin.spawn (#10162)pytorch_lightning.lite package (#10175)LightningLite documentation (#10043)LightningLite examples (#9987)_LiteDataLoader an iterator and add supports for custom dataloader (#10279)use_omegaconf argument to save_hparams_to_yaml plugin (#9170)ckpt_path argument for Trainer.fit() (#10061)auto_device_count method to Accelerators (#10222)devices="auto" (#10264)filename argument in ModelCheckpoint.format_checkpoint_name (#9818)gpus list to run on CPU (#10246)MisconfigurationException when its methods are called with ckpt_path="best" but a checkpoint callback isn't configured (#9841)Trainer(accelerator="ddp_cpu") now does not spawn a subprocess if num_processes is kept 1 along with num_nodes > 1 (#9603)ModuleNotFoundError instead of ImportError (#9867)pytorch_lightning.loggers.neptune.NeptuneLogger is now consistent with the new neptune-client API; the old neptune-client API is supported by NeptuneClient from the neptune-contrib repo (#6867)enums type hyperparameters to be saved in the haprams.yaml file by TensorBoard and CSV loggers has been fixed and made in line with how OmegaConf parses it (#9170)gpus Trainer argument has changed: gpus="n" (str) no longer selects the GPU index n and instead selects the first n devices (#8770)iteration_count and other index attributes in the loops has been replaced with progress dataclasses (#8477)trainer.lightning_module reference is now properly set at the very beginning of a run (#8536)Trainer functions reset_{train,val,test,predict}_dataloader, reset_train_val_dataloaders, and request_dataloader model argument is now optional (#8536)Callback as the key to avoid issues with unpickling (#6886)ResultCollection (#8622)LightningCLI changes:
LightningCLI.init_parser now returns the parser instance (#8721)LightningCLI.add_core_arguments_to_parser, LightningCLI.parse_arguments now take a parser argument (#8721)LightningCLI.instantiate_trainer now takes a config and a list of callbacks (#8721)LightningCLI.add_core_arguments_to_parser into LightningCLI.add_default_arguments_to_parser + LightningCLI.add_core_arguments_to_parser (#8721)setup hooks no longer have a model argument (#8536)update_global_step hook has been removed (#8856)self.log-ing in any LightningModule or Callback hook has been improved (#8498)self.log-ing without a Trainer reference now raises a warning instead of an exception (#9733)Trainer.request_dataloader now takes a RunningStage enum instance (#8858)rank_zero_warn to NotImplementedError in the {train, val, test, predict}_dataloader hooks that Lightning(Data)Module uses (#9161)block_ddp_sync_behaviour out of TrainingBatchLoop to loop utilities (#9192)optimizer_closure is now required when overriding the optimizer_step hook (#9360)LightningModule and LightningDataModule hyperparameters to raise an exception only if there are colliding keys with different values (#9496)seed_everything now fails when an invalid seed value is passed instead of selecting a random seed (#8787)TrainingTypePlugin collective APIs directly instead of going through the Accelerator reference (#9677, #9901)HorovodPlugin.all_gather to return a torch.Tensor instead of a list (#9696)current_epoch and global_step attributes now get restored irrespective of the Trainer task (#9413)amp_level with native amp_backend (#9755)pytorch_lightning.utilities.grads.grad_norm now raises an exception if parameter norm_type <= 0 (#9765)optimizer_step and clip_gradients hook from the Accelerator and TrainingTypePlugin into the PrecisionPlugin (#10143, #10029)NativeMixedPrecisionPlugin and its subclasses now take an optional GradScaler instance (#10055)MisconfigurationException instead of a warning if Trainer.{validate/test} is missing required methods (#10016)max_steps Trainer argument from None to -1 (#9460)log(on_step=False, on_epoch=False) (#10227)MisconfigurationException when total length of dataloader across ranks is zero, and give warning when total length is non-zero, but only local rank length is zero. (#9827)ByteCounter (#10123)on_load_checkpoint for LightningDataModule for all trainer_fn (#10238)subclass_mode=False (#10286)terminate_on_nan in favor of detect_anomaly(#9175)Trainer.terminate_on_nan public attribute access (#9849)LightningModule.summarize() in favor of pytorch_lightning.utilities.model_summary.summarize() (#8513)LightningModule.model_size (#8343)DataModule properties: train_transforms, val_transforms, test_transforms, size, dims (#8851)add_to_queue, get_from_queue from LightningModule in favor of corresponding methods in the DDPSpawnPlugin (#9118)LightningModule.get_progress_bar_dict and Trainer.progress_bar_dict in favor of pytorch_lightning.callbacks.progress.base.get_standard_metrics and ProgressBarBase.get_metrics (#8985)prepare_data_per_node flag on Trainer and set it as a property of DataHooks, accessible in the LightningModule and LightningDataModule (#8958)TestTubeLogger (#9065)on_{train/val/test/predict}_dataloader() from LightningModule and LightningDataModule (#9098)on_keyboard_interrupt callback hook in favor of new on_exception hook (#9260)process_position to the Trainer constructor in favor of adding the ProgressBar callback with process_position directly to the list of callbacks (#9222)flush_logs_every_n_steps as a Trainer argument, instead pass it to the logger init if supported (#9366)LightningLoggerBase.close, LoggerCollection.close in favor of LightningLoggerBase.finalize, LoggerCollection.finalize (#9422)progress_bar_refresh_rate to the Trainer constructor in favor of adding the ProgressBar callback with refresh_rate directly to the list of callbacks, or passing enable_progress_bar=False to disable the progress bar (#9616)LightningDistributed and moved the broadcast logic to DDPPlugin and DDPSpawnPlugin directly (#9691)stochastic_weight_avg to the Trainer constructor in favor of adding the StochasticWeightAveraging callback directly to the list of callbacks (#8989)barrier, broadcast, and all_gather in favor of calling the TrainingTypePlugin collective API directly (#9677)checkpoint_callback from the Trainer constructor in favor of enable_checkpointing (#9754)LightningModule.on_post_move_to_device method (#9525)pytorch_lightning.core.decorators.parameter_validation in favor of pytorch_lightning.utilities.parameter_tying.set_shared_parameters (#9525)weights_summary to the Trainer constructor in favor of adding the ModelSummary callback with max_depth directly to the list of callbacks (#9699)log_gpu_memory, gpu_metrics, and util funcs in favor of DeviceStatsMonitor callback (#9921)GPUStatsMonitor and XLAStatsMonitor in favor of DeviceStatsMonitor callback (#9924)Trainer(max_steps=None); To turn off the limit, set Trainer(max_steps=-1) (default) (#9460)AcceleratorConnector.is_slurm_managing_tasks attribute and marked it as protected (#10101)AcceleratorConnector.configure_slurm_ddp method and marked it as protected (#10101)resume_from_checkpoint to the Trainer constructor in favor of trainer.fit(ckpt_path=) (#10061)ClusterEnvironment.creates_children() in favor of ClusterEnvironment.creates_processes_externally (property) (#10106)PrecisionPlugin.master_params() in favor of PrecisionPlugin.main_params() (#10105)lr_sch_names from LearningRateMonitor (#10066)ProgressBar callback in favor of TQDMProgressBar (#10134)metrics (#8586)outputs argument in both the LightningModule.on_train_epoch_end and Callback.on_train_epoch_end hooks (#8587)TrainerLoggingMixin class (#8609)TrainerTrainingTricksMixin class (#8679)optimizer_idx from training_step as an accepted argument in manual optimization (#8576)on_save_checkpoint signature. The hook now takes a checkpoint positional parameter (#8697)on_load_checkpoint signature. The hook now takes a pl_module positional parameter (#8697)save_function property in ModelCheckpoint (#8680)model argument from ModelCheckpoint.save_checkpoint (#8688)sync_step argument from WandbLogger (#8763)Trainer.truncated_bptt_steps in favor of LightningModule.truncated_bptt_steps (#8826)LightningModule.write_predictions and LightningModule.write_predictions_dict (#8850)on_reset_*_dataloader hooks in TrainingType Plugins and Accelerators (#8858)GradInformation module in favor of pytorch_lightning.utilities.grads (#8831)TrainingTypePlugin.on_save and Accelerator.on_save (#9023){Accelerator,TrainingTypePlugin,PrecisionPlugin}.post_optimizer_step (#9746)connect_precision_plugin and connect_training_type_plugin from Accelerator (#9019)on_train_epoch_end from Accelerator (#9035)InterBatchProcessor in favor of DataLoaderIterDataFetcher (#9052)Plugin in base_plugin.py in favor of accessing TrainingTypePlugin and PrecisionPlugin directly instead (#9066)teardown from ParallelPlugin (#8943)profiled_functions argument from PyTorchProfiler (#9178)pytorch_lighting.utilities.argparse_utils module (#9166)Trainer.running_sanity_check in favor of Trainer.sanity_checking (#9209)BaseProfiler.output_filename arg from it and its descendants in favor of dirpath and filename (#9214)ModelCheckpoint.period in favor of ModelCheckpoint.every_n_epochs (#9213)auto_move_data decorator (#9231)LightningModule.datamodule in favor of Trainer.datamodule (#9233)DeepSpeedPlugin.cpu_offload* in favor of offload_optimizer, offload_parameters and pin_memory (#9244)AcceleratorConnector.is_using_torchelastic in favor of TorchElasticEnvironment.is_using_torchelastic() (#9729)pytorch_lightning.utilities.debugging.InternalDebugger (#9680)call_configure_sharded_model_hook property from Accelerator and TrainingTypePlugin (#9612)TrainerProperties mixin and moved property definitions directly into Trainer (#9495)ModelCheckpoint(monitor=None) callback (#9875)epoch from trainer.logged_metrics (#9904)should_rank_save_checkpoint property from Trainer (#9433)distributed_backend from Trainer (#10017)process_idx from the {DDPSpawnPlugin,TPUSpawnPlugin}.new_process methods (#10022){train,val,test,predict}_dataloader() on the LightningModule (#9764)pytorch_lightning.trainer.connectors.OptimizerConnector (#10120)move_metrics_to_cpu moving the loss to CPU while training on device (#9308)deepcopy-ed (#9349)Loop.on_run_end (#9386, #9915)BasePredictionWriter not returning the batch indices in a non-distributed setting (#9432)compute() output is a multielement tensor (#9582)DDPShardedPlugin (#9122)DDPPlugin, DDPSpawnPlugin, DDPShardedPlugin, DDPSpawnShardedPlugin (#9096)trainer.accumulate_grad_batches to be an int on init. The default value for it is now None inside Trainer (#9652)broadcast in DDPPlugin and DDPSpawnPlugin to respect the src input (#9691)self.log(on_epoch=True, reduce_fx=sum)) for the on_batch_start and on_train_batch_start hooks (#9791)self.log(on_epoch=True) for the on_batch_start and on_train_batch_start hooks (#9780)Trainer.fit only (#9413)val_dataloader in tuner/batch_size_scaling (#9857)LightningCLI in computer_vision_fine_tuning.py example (#9934)apply_to_collection (#9963)val_dataloader in tuner/batch_size_scaling for binsearch (#9975)TrainerDataLoadingMixin._worker_check (#9902)train_dataloader getting loaded twice when resuming from a checkpoint during Trainer.fit() (#9671)LearningRateMonitor logging with multiple param groups optimizer with no scheduler (#10044)Trainer patching dataloader methods on the LightningModule (#9764)on_before_optimizer_step getting called before the optimizer closure (including backward) has run (#10167)ModelCheckpoint getting moved to the wrong device in a special case where it becomes NaN (#10118)dirpath in BaseProfiler if it doesn't exist (#10073)log(on_step=True, on_epoch=True, sync_dist=True) wouldn't reduce the value on step (#10227)pl.utilities.seed.reset_seed converting the PL_SEED_WORKERS environment variable to bool (#10099)fast_dev_run > 0 (#10232)batch_size in ResultCollection not being reset to 1 on epoch end (#10242)distrib_type not being set when training plugin instances are being passed to the Trainer (#10251)@adamjstewart @akihironitta @alessiobonfiglio @ananthsub @aphedges @awaelchli @bamblebam @Benjamin-Etheredge @borchero @Borda @borisdayma @bryant1410 @carmocca @cowwoc @daniellepintz @danielykim @edward-io @eladsegal @EricWiener @ethanwharris @four4fish @gau-nernst @hankyul2 @HansolEom @himanshu-dutta @I-iBot @jjenniferdai @jstjohn @justusschock @kainoj @kaushikb11 @kingyiusuen @Knarik1 @low5545 @lsqshr @mauvilsa @michele-arrival @nasnoisaac @ninginthecloud @popfido @pre-commit-ci @PuneetDabral @qmpzzpmq @rohitgr7 @ronif @roshikouhai @s-rog @samlurye @SeanNaren @shnela @sidml @stancld @stfwn @tangbinh @tchaton @thepurpleowl @Tshimanga @twsl @victorjoos @VirajBagal @wayi1 @weiji14 @yifuwang @yopknopixx
If we forgot someone, let us know :]
Nothing published for this version
Nothing published for this version
Moved the gradient unscaling in NativeMixedPrecisionPlugin from pre_optimizer_step to post_backward
NativeMixedPrecisionPlugin from pre_optimizer_step to post_backward (#9606)lr_find to generate same results on multiple calls (#9704)reset metrics on validation epoch end (#9717)gradient_clip_val, gradient_clip_algorithm, track_grad_norm and terminate_on_nan Trainer arguments (#9595)@rohitgr7 @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed error reporting in DDP process reconciliation when processes are launched by an external agent
add_argparse_args raising TypeError when args are typed as typing.Generic in Python 3.6 (#9554)@ananthsub @akihironitta @awaelchli @carmocca @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed logging of nan parameters
replace_sampler missing the batch size under specific conditions (#9367)@asanakoy @awaelchli @borisdayma @carmocca @guotuofeng @justusschock @kaushikb11 @rohitgr7 @SeanNaren
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Removed deprecation warnings being called for on_{task}_dataloader
on_{task}_dataloader (#9279)EarlyStopping running on train epoch end when check_val_every_n_epoch>1 is set (#9156)on_before_optimizer_step hook (#9288)*_epoch_end hook wasn't overridden (#9261)_sync_dir was not initialized (#9267)save_hyperparameters (#9125)Timer.on_train_epoch_end and StochasticWeightAveraging.on_train_epoch_end to prevent unwanted deprecation warnings (#9347)@ananthsub @awaelchli @Borda @four4fish @justusschock @kaushikb11 @s-rog @SeanNaren @tangbinh @tchaton @xerus
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed reduction using self.log(sync_dict=True, reduce_fx={mean,max})
self.log(sync_dict=True, reduce_fx={mean,max}) (#9142)max_epochs if max_time was specified on the Trainer constructor (#9072)DDP "CUDA error: initialization error" due to a copy instead of deepcopy on ResultCollection (#9239)@ananthsub @bamblebam @carmocca @daniellepintz @ethanwharris @kaushikb11 @sohamtiwari3120 @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed a bug in the binary search mode of auto batch size scaling where exception was raised if the first trainer run resulted in OOM
log_gpu_memory='min_max' not working (#9013)@SkafteNicki @eladsegal
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed plateau scheduler stepping on incomplete epoch
CycleIterator and multiple loaders (#8889)StochasticWeightAveraging with a list of learning rates not applying them to each param group (#8747)_Metadata object in ResultMetricCollection (#8932)DDPPlugin._sync_dir in reconciliate_processes (#8939)@awaelchli @carmocca @justusschock @tchaton @yifuwang If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed recursive call for apply_to_collection(include_none=False)
apply_to_collection(include_none=False) (#8719)@Aiden-Jeon @ananthsub @awaelchli @edward-io If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed trainer.fit_loop.split_idx always returning None
trainer.fit_loop.split_idx always returning None (#8601)ResultCollection.extra (#8622)mpirun (#8610)training_step outputs not getting collected correctly for training_epoch_end (#8613)accelerator=ddp choice for CPU (#8645)@awaelchli, @borda, @carmocca, @kaushikb11, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated LightningModule.loaded_optimizer_states_dict
Today we are excited to announce Lightning 1.4, introducing support for TPU pods, XLA profiling, IPUs, and new plugins to reach 10+ billion parameters, including Deep Speed Infinity, Fully Sharded Data-Parallel and more!
https://devblog.pytorchlightning.ai/announcing-lightning-1-4-8cd20482aee9
extract_batch_size utility and corresponding tests to extract batch dimension from multiple batch types (#8357)LearningRateMonitor (#7987)dataclass support for pytorch_lightning.utilities.apply_to_collection (#7935)LightningModule.to_torchscript for saving to custom filesystems with fsspec (#7617)KubeflowEnvironment for use with the PyTorchJob operator in KubeflowModelPruning(prune_on_train_epoch_end=True|False) to choose when to apply pruning (#7704){,load_}state_dict to the progress tracking dataclasses (#8140)LightningDataModule positionally as the second argument to trainer.{validate,test,predict} (#7431)trainer.predict(ckpt_path) (#7430)clip_grad_by_value support for TPUs (#7025)is_overridden (#7918)sub_dir parameter to TensorBoardLogger (#6195)dataloader_idx to batch transfer hooks (#6241)include_none=bool argument to apply_to_collection (#7769)apply_to_collections to apply a function to two zipped collections (#7769)ddp_fully_sharded support (#7487)should_rank_save_checkpoint property to Training Plugins (#7684)log_grad_norm hook to LightningModule to customize the logging of gradient norms (#7873)save_config_filename init argument to LightningCLI to ease resolving name conflicts (#7741)save_config_overwrite init argument to LightningCLI to ease overwriting existing config files (#8059)on_before_optimizer_step hook (#8048){,load_}state_dict to ResultCollection (#7948){,load_}state_dict to Loops (#8197)Loop.restarting=False at the end of the first iteration (#8362)state_dict and load_state_dict utilities for CombinedLoader + utilities for dataloader (#8364)rank_zero_only to LightningModule.log function (#7966)metric_attribute to LightningModule.log function (#7966)Trainer(log_every_n_steps) is a value too high for the training dataloader (#7734)torch.nn.UninitializedParameter in ModelSummary (#7642)LightningModule.save_hyperparameters when LightningModule is a dataclass (#7992)optimizer_zero_grad and optimizer_step when using accumulate_grad_batches (#7980)logger boolean flag to save_hyperparameters (#7960)python -m package.script) (#8073)LightningCLI (#8093)PrecisionPlugin.{pre,post}_backward (#8328)on_load_checkpoint and on_save_checkpoint hooks to the PrecisionPlugin base class (#7831)max_depth parameter in ModelSummary (#8062)XLAStatsMonitor callback (#8235)restore function and restarting attribute to base Loop (#8247)FastForwardSampler and CaptureIterableDataset (#8307)save_hyperparameters in LightningDataModule (#3792)ModelCheckpoint(save_on_train_epoch_end) to choose when to run the saving logic (#8389)LSFEnvironment for distributed training with the LSF resource manager jsrun (#5102)accelerator='cpu'|'gpu'|'tpu'|'ipu'|'auto' (#7808)tpu_spawn_debug to plugin registry (#7933)LOCAL_RANK and NODE_RANK environment variable assignments (#7480)quantize_on_fit_end argument to QuantizationAwareTraining (#8464)devices flag to Trainer (#8440)prevent_trainer_and_dataloaders_deepcopy context manager on the LightningModule (#8472)Trainer's checkpoint_callback argument to allow only boolean values (#7539)on_evaluation_end hook (#7272)self.log(on_epoch=False) during epoch-only or single-call hooks (#7874)Trainer methods to be protected: call_setup_hook, call_configure_sharded_model, pre_dispatch, dispatch, post_dispatch, call_teardown_hook, run_train, run_sanity_check, run_evaluate, run_evaluation, run_predict, track_output_for_epoch_endmetrics_to_scalars to work with any collection or value (#7888)clip_grad_norm to use torch.nn.utils.clip_grad_norm_ (#7025)ModelCheckpoint now runs at the end of the training epoch by default (#8389)EarlyStopping now runs at the end of the training epoch by default (#8286)global_step, current_epoch, max/min_steps, max/min_epochs, batch_idx, and total_batch_idx to TrainLoop (#7437)hiddens and split_idx to TrainLoop (#7507)on_epoch guard from the "should stop" validation check (#7701)FitLoop, TrainingEpochLoop, TrainingBatchLoop (#7871, #8077)pytorch_lightning/trainer/training_loop.py (#7985)DataLoaderLoop, EvaluationLoop, EvaluationEpochLoop (#7990, #8077)pytorch_lightning/trainer/evaluation_loop.py (#8056)_run_* functions and separate evaluation loops (#8065)PredictionLoop, PredictionEpochLoop (#7700, #8077)pytorch_lightning/trainer/predict_loop.py (#8094)Loop API to better handle children state_dict and progress (#8334)core/step_result.py to trainer/connectors/logger_connector/result.py (#7736)LoggerConnector (#7882)trainer.{logged,progress_bar,callback}_metrics are now updated on-demand (#7882)Result object in favor of ResultMetric (#7882)self.log(batch_size=...) (#7891)EpochResultStore and HookResultStore in favor of ResultCollection (#7909)MetricsHolder (#7909)ignore_scalar_return_in_dp warning suppression to the DataParallelPlugin class (#7421)/epoch_* to the metric name (#7351)ValueError when a None value is self.log-ed (#7771)resolve_training_type_plugins to allow setting num_nodes and sync_batchnorm from Trainer setting (#7026)seed_everything(workers=True) in the LightningCLI (#7504)model.state_dict() in CheckpointConnector to allow training_type_plugin to customize the model's state_dict() (#7474)MLflowLogger now uses the env variable MLFLOW_TRACKING_URI as default tracking URI (#7457)Trainer arg and functionality from reload_dataloaders_every_epoch to reload_dataloaders_every_n_epochs (#5043)WandbLogger(log_model={True/'all'}) to log models as artifacts (#6231)run_name as an constructor argument (#7622)teardown() in Accelerator to allow training_type_plugin to customize teardown logic (#7579)Trainer.fit now raises an error when using manual optimization with unsupported features such as gradient_clip_val or accumulate_grad_batches (#7788)LightningModule overrides the same hooks (#7826)on_after_backward hook is now called on accumulating iterations. Use the on_before_optimizer_step hook to mimic the old behaviour (#8328)on_after_backward hook. Use the on_before_optimizer_step hook to mimic the old behaviour (#8328)TrainingTypePlugin.{pre,post}_backward hooks no longer take the optimizer, opt_idx, should_accumulate arguments (#8328)PrecisionPlugin.backward hooks no longer returns a value (#8328)PrecisionPlugin.backward hooks no longer takes a should_accumulate argument (#8328)on_before_backward hook (#7865)LightningCLI now aborts with a clearer message if config already exists and disables save config during fast_dev_run(#7963)LightningCLI config on setup and only on the main process (#8017)LightningCLI ArgumentParser when pickling (#8017)broadcast if distributed not initialized for the spawn plugins (#8017)Trainer(resume_from_checkpoint=...) now restores the model directly after LightningModule.setup(), which is before LightningModule.configure_sharded_model() (#7652)torch.cuda.set_device() to enable collective calls earlier in setup (#8312)replace_sampler when the DataLoader attributes are not included in the signature or the signature is missing optional arguments (#8519)DeviceDtypeModuleMixin and HyperparametersMixin mixin to core (#8396)default_root_dir as the log_dir when the logger is a LoggerCollection (#8187)LightningModule.loaded_optimizer_states_dict (#8229)trainer.{fit,valdiate,test,tune} (#7431)DataModule properties: has_prepared_data, has_setup_fit, has_setup_validate, has_setup_test, has_setup_predict, has_teardown_fit, has_teardown_validate, has_teardown_test, has_teardown_predict (#7657)TrainerModelHooksMixin in favor of pytorch_lightning.utilities.signature_utils (#7422)num_nodes and sync_batchnorm arguments in DDPPlugin and DDPSpawnPlugin (#7026)self.log(sync_dist_op) in favor of self.log(reduce_fx). (#7891)is_overridden(model=...) in favor of is_overridden(instance=...) (#7918)monitor argument in EarlyStopping callback to enforce monitor as a required argument (#7907)rank_zero_{warn,deprecation} directly from pytorch_lightning.utilities.distributed (#8085)CheckpointConnector.hpc_load() in favor of CheckpointConnector.restore() (#7652)ModelCheckpoint(every_n_val_epochs) in favor of ModelCheckpoint(every_n_epochs) (#8383)DDPPlugin.task_idx in favor of DDPPlugin.local_rank (#8203)Trainer.train_loop property in favor of Trainer.fit_loop (#8025)Trainer.disable_validation property in favor of not Trainer.enable_validation (#8291)mode parameter in ModelSummary in favor of max_depth (#8062)reload_dataloaders_every_epoch argument of Trainer in favor of reload_dataloaders_every_n_epochs (#5043)distributed_backend argument for Trainer (#8575)ProfilerConnector (#7654)pytorch_lightning.metrics.functional.classification (#7499)LightningDataParallel and LightningDistributedDataParallel from pytorch_lightning.overrides.data_parallel (#7510)get_model and accelerator_backend (#7502)val_loss key with ModelCheckpoint. Pass your monitor of choice to the ModelCheckpoint instance instead (#8293)self.log(tbptt_reduce_fx) and self.log(tbptt_pad_token). Please, open a discussion explaining your use-case if you relied on these. (#7644)model_utils, warning_utils, xla_device_utils and partially argparse_utils (#7503)RPCPlugin and RPCSequentialPlugin. If you were successfully using these plugins, please open a GitHub discussion about your use case (#8101)on_cpu, on_tpu, use_tpu, on_gpu, use_dp, use_ddp, use_ddp2, use_horovod, use_single_gpu (#7501)optimizer argument in LightningModule.manual_backward(); Toggling optimizers in manual optimization should be done using LightningModule.{un}toggle_optimizer() (#8287)PL_EXP_VERSION from DDP subprocesses (#7403)GPUStatsMonitor callbacks to use the correct GPU IDs if CUDA_VISIBLE_DEVICES set (#8260)lr_scheduler checkpointed state by calling update_lr_schedulers before saving checkpoints (#7877)None loss keys getting added in training_epoch_end when using manual optimization and not returning a loss (#7772)precision=64 with accelerator='ddp_spawn' would throw a pickle error (#6924)epoch value in logged_metrics when already logged by the user (#7982)dataloader_idx argument value when predicting with only one DataLoader (#7941)stage argument of Callback.{setup,teardown} as a keyword (#7973)validation sanity checking are cleaned on end (#8171)log_gpu_memory metrics not being added to logging when nothing else is logged (#8174)log with a Metric instance would raise an error if it was a nested attribute of the model (#8181)precision=64 would cause buffers with complex dtype to be cast to real (#8208)is_overridden returning true for wrapped functions with no changes (#8296)truncated_bptt_steps would throw an AttributeError when the target RNN has multiple hidden states (#8145)self.optimizers() not returning a single optimizer if it had been wrapped (#8326)on_after_backward hook not getting called when using manual optimization and no plugins (#8328)LightningModule.backward hook only getting called with the apex plugin when using manual optimization (#8328)on_*_batch_start/on_*_batch_end callbacks and model hooks (#7378)DDPPlugin when choosing accelerator="ddp_cpu" for the accelerator (#6208)LightningModule.untoggle_optimizer in training loop when running gradient accumulation with multiple optimizers (#8284)val_check_interval did not align with the number of training batches (#7724)move_data_to_device to return the batch if the object to function didn't return self (#8433)optimizer_states, ResultCollection.extra, ResultMetric attributes, and LoggerConnector metrics to cpu. Also, delete the DDP wrapper on teardown (#8490)SWA callback using LightningModule prevent_trainer_and_dataloaders_deepcopy to avoid OOM (#8472)ModelPruning callback on_save_checkpoint to avoid making a deepcopy potentially leading to OOM (#8472)DataLoaders which do not define all DataLoader attributes as __init__ parameters (#8519)lr_schedulers attribute (#8527)Trainer instances in sequence (#7403)accumulate_grad_batches not been recomputed during model reload (#5334)TypeError when wrapping optimizers in the HorovodPlugin and running Trainer.test (#7840)BackboneFinetuning restoration (#8501)lr_scheduler with metric (e.g. torch.optim.lr_scheduler.ReduceLROnPlateau) when using automatic_optimization = False (#7643)DeepSpeed breaking with no schedulers (#8580)@00sapo @AffineParameter @ajtritt @akihironitta @ananthsub @aniketmaurya @aslisabanci @awaelchli @bamblebam @Borda @borisdayma @carmocca @dalek-who @DavidMChan @davors72 @dcfidalgo @ddrevicky @deepsource-autofix @djthegr8 @edenlightning @edgarriba @eladsegal @ethanwharris @eugeneh101 @fepegar @gaoteng-git @gtauzin @i-aki-y @janhenriklambrechts @jiwidi @justusschock @karthikrangasai @kaushikb11 @loic-beheshti @Lucklyric @ManuelPalermo @mauvilsa @maxoppelt @neggert @nikvaessen @nisheethlahoti @pre-commit-ci @rohitgr7 @ruotianluo @satishjasthi @SeanNaren @shirayu @shuyingsunshine21 @sid-sundrani @Sileadim @simran2905 @stancld @t-vi @tchaton @theblackfly @theodumont @tilman151 @tomy0000000 @tshu-w @vatch123 @WrRan @yifuwang
If we forgot someone, let us know :]
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →