NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1157 most downloaded on PyPI
PyTorch Lightning is the lightweight PyTorch wrapper for ML researchers. Scale your models. Write less boilerplate.
Last release 25 days ago
10 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
223 releases · first in 2019
Nothing published for this version
Fixed metrics deprecation message at module import level
LightningModule that uses a torchmetrics 0.4 Metric (#8218)BaseFinetuning callback on a model that contains a ModuleDict (#8170)deadlock for DDP when only 1 process trigger an Exception. The mechanism will kill the processes when it happens (#8167)IterableDataset (#8172)@GabrielePicco @SeanNaren @ethanwharris @carmocca @tchaton @justusschock
One column per quarter.
Fixed backward compatibility of moved functions rank_zero_warn and rank_zero_deprecation
rank_zero_warn and rank_zero_deprecation (#8085)@kaushikb11 @carmocca
Fixed deprecation messages not showing due to incorrect stacklevel (#8002, #8005)
DistributedSampler when using a distributed plugin in a custom accelerator (#7814)PyTorchProfiler chrome traces names (#8009)EarlyStopping callback for TPU devices (#7959)@yifuwang @kaushikb11 @ajtritt @carmocca @tchaton
Fixed logs overwriting issue for remote filesystems
DataModule.prepare_data could only be called on the global rank 0 process (#7945)worker_init_fn to seed dataloaders correctly when using DDP (#7942)BaseFinetuning callback to properly handle parent modules w/ parameters (#7931)@awaelchli @Borda @kaushikb11 @Queuecumber @SeanNaren @senarvi @speediedan
Added warning to Training Step output
apply_to_collection and type signature of log_dict (#7851)training_output validation to after train_step_end (#7868)@Borda, @justusschock, @kandluis, @mauvilsa, @shuyingsunshine21, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed info message when max training time reached
Changed calling of untoggle_optimizer(opt_idx) out of the closure function
untoggle_optimizer(opt_idx) out of the closure function (#7563)ProgressBar pickling after calling trainer.predict (#7608)ProgressBar when trainer.fit is not called (#7674)@carmocca @kaushikb11 @ryanking13 @Lucklyric @ajtritt @yifuwang
If we forgot someone due to not matching commit email with GitHub account, let us know :]
DataModules now avoid duplicate {setup,teardown,prepare_data} calls for the same stage
DataModules now avoid duplicate {setup,teardown,prepare_data} calls for the same stage (#7238)wrong_type keyword argument in pytorch_lightning.utilities.apply_to_collection (#7433)DistribType for ddp_cpu (spawn) backend (#7492)check_val_every_n_epoch > 1 (#7032)@alanhdu @carmocca @justusschock @tkng
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed DeepSpeed with IterableDatasets
Trainer.current_epoch not getting restored after tuning (#7434)@akihironitta @awaelchli @leezu
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated outputs in both LightningModule.on_train_epoch_end and Callback.on_train_epoch_end hooks
Today we are excited to announce Lightning 1.3, containing highly anticipated new features including a new Lightning CLI, improved TPU support, integrations such as PyTorch profiler, new early stopping strategies, predict and validate trainer routines, and more.
https://medium.com/pytorch/pytorch-lightning-1-3-lightning-cli-pytorch-profiler-improved-early-stopping-6e0ffd8deb29
EarlyStopping callback to run at the end of the training epoch (#6944)setup hooks are run (#7202)teardown hook to ClusterEnvironment (#6942)trainer.test() or trainer.validate() with fast_dev_run=True (#6667)LightningCLI class to provide simple reproducibility with minimum boilerplate training CLI (#4492, #6862, #7156, #7299)gradient_clip_algorithm argument to Trainer for gradient clipping by value (#6123).ModelCheckpoint callback (#6146)TrainerStatus.{INITIALIZING,RUNNING,FINISHED,INTERRUPTED} (#7173)Trainer.validate() method to perform one evaluation epoch over the validation set (#4948)LightningEnvironment for Lightning-specific DDP (#5915)teardown() hook to LightningDataModule (#4673)auto_insert_metric_name parameter to ModelCheckpoint (#6277)self.log that enables users to give custom names when dealing with multiple dataloaders (#6274)teardown method to BaseProfiler to enable subclasses defining post-profiling steps outside of __del__ (#6370)setup method to BaseProfiler to enable subclasses defining pre-profiling steps for every process (#6633)Trainer.predict config validation (#6543)AbstractProfiler interface (#6621)PyTorchProfiler (#6349)outputs parameter to callback's on_validation_epoch_end & on_test_epoch_end hooks (#6120)configure_sharded_model hook (#6679)precision=64, enabling training with double precision (#6595)artifact_location argument to MLFlowLogger which will be passed to the MlflowClient.create_experiment call (#6677)model parameter to precision plugins' clip_gradients signature (#6764, #7231)is_last_batch attribute to Trainer (#6825)LightningModule.lr_schedulers() for manual optimization (#6567)MpModelWrapper in TPU Spawn (#7045)max_time Trainer argument to limit training time (#6823)on_predict_{batch,epoch}_{start,end} hooks (#7141)EarlyStopping parameters stopping_threshold and divergence_threshold (#6868)debug flag to TPU Training Plugins (PT_XLA_DEBUG) (#7219)UnrepeatedDistributedSampler and IndexBatchSamplerWrapper for tracking distributed predictions (#7215)trainer.predict(return_predictions=None|False|True) (#7215)BasePredictionWriter callback to implement prediction saving (#7127)trainer.tune(scale_batch_size_kwargs, lr_find_kwargs) arguments to configure the tuning algorithms (#7258)tpu_distributed check for TPU Spawn barrier (#7241)Callback and using resume_from_checkpoint (#7254)ignore param to save_hyperparameters (#6056)LightningModule.truncated_bptt_steps to be property (#7323)EarlyStopping callback from by default running EarlyStopping.on_validation_end if only training is run. Set check_on_train_epoch_end to run the callback at the end of the train epoch instead of at the end of the validation epoch (#7069)pytorch_lightning.callbacks.swa to pytorch_lightning.callbacks.stochastic_weight_avg (#6259)RunningStage and TrainerState usage (#4945, #7173)
RunningStage.SANITY_CHECKINGTrainerFn.{FITTING,VALIDATING,TESTING,PREDICTING,TUNING}trainer.evaluating to return True if validating or testingsetup() and teardown() stage argument to take any of {fit,validate,test,predict} (#6386)on_train_end functions (#6864)PyTorchProfiler to use torch.autograd.profiler.record_function to record functions (#6349)lr_scheduler.step() in manual optimization (#6825)ddp_spawn (#6762)pl.seed_everything will now also set the seed on the DistributedSampler (#7024)DDPShardedPlugin (#6937)trainer.tune() now returns the tuning result (#7258)LightningModule.from_datasets() now accepts IterableDataset instances as training datasets. (#7503)resume_from_checkpoint warning to an error when the checkpoint file does not exist (#7075)sync_batchnorm for training_type_plugin (#6536)save_function to accelerator (#6689)EarlyStopping callback (#6811)resolve_tags (#6746)save_hyperparameters to its own function (#7119)_DataModuleWrapper with __new__ (#7289)current_fx properties on lightning module in teardown (#7247)DataLoader.worker_init_fn with seed_everything (#6960)model.trainer call inside of dataloading mixin (#7317)outputs in both LightningModule.on_train_epoch_end and Callback.on_train_epoch_end hooks (#7339)Trainer.truncated_bptt_steps in favor of LightningModule.truncated_bptt_steps (#7323)outputs in both LightningModule.on_train_epoch_end and Callback.on_train_epoch_end hooks (#7339)LightningModule.grad_norm in favor of pytorch_lightning.utilities.grads.grad_norm (#7292)save_function property from the ModelCheckpoint callback (#7201)LightningModule.write_predictions and LightningModule.write_predictions_dict (#7066)TrainerLoggingMixin in favor of a separate utilities module for metric handling (#7180)TrainerTrainingTricksMixin in favor of a separate utilities module for NaN/Inf detection for gradients and parameters (#6834)period has been deprecated in favor of every_n_val_epochs in the ModelCheckpoint callback (#6146)trainer.running_sanity_check in favor of trainer.sanity_checking (#4945)Profiler(output_filename) in favor of dirpath and filename (#6621)PytorchProfiler(profiled_functions) in favor of record_functions (#6349)@auto_move_data in favor of trainer.predict (#6993)Callback.on_load_checkpoint(checkpoint) in favor of Callback.on_load_checkpoint(trainer, pl_module, checkpoint) (#7253)torchmetrics (#6505, #6530, #6540, #6547, #6515, #6572, #6573, #6584, #6636, #6637, #6649, #6659, #7131)LightningModule.datamodule getter and setter methods; access them through Trainer.datamodule instead (#7168)Trainer(gpus="i") (string) for selecting the i-th GPU; from v1.5 this will set the number of GPUs instead of the index (#6388)exp_save_path property from the LightningModule (#7266)EarlyStopping.on_validation_end if no validation is run (#7069)automatic_optimization as a property from the training loop in favor of LightningModule.automatic_optimization (#7130)*_epoch_end hooks (#6973)profiler argument of Trainer (#6164)ModelCheckpoint instance to Trainer(checkpoint_callback) (#6166)enable_pl_optimizer and automatic_optimization (#6163)pytorch_lightning.metrics.functional.classification removed to_onehot, to_categorical, get_num_classes, roc, multiclass_roc, average_precision, precision_recall_curve, multiclass_precision_recall_curvepytorch_lightning.metrics.functional.reduction removed reduce, class_reduceModelCheckpoint arguments prefix, mode="auto" (#6162)mode='auto' from EarlyStopping (#6167)epoch and step arguments from ModelCheckpoint.format_checkpoint_name(), these are now included in the metrics argument (#7344)Result object (#6016)LightningModule hparams setter (#6207)"log"/"progress_bar" magic keys. Use self.log instead (#6734)trainer.fit() return value of 1. It has no return now (#7237)logger_connector legacy code (#6733)reload_dataloaders_every_epoch=True and num_sanity_val_steps=0 (#7207)teardown to synchronize processes before execution finishes (#6814)local_rank instead of global_rank for main process assertion (#7061)WORLD_SIZE environment variable in DDP training when launching with torch distributed/torchelastic (#6942)Plugin.reduce method more consistent across all Plugins to reflect a mean-reduction by default (#6011)ModelCheckpoint(monitor=None) (#6109)ModelCheckpoint(save_top_k=0, save_last=True) not saving the last checkpoint (#6136).teardown(stage='fit') and .on_fit_{start,end}() getting called during trainer.test (#6386)all_gather on cpu tensors (#6416)trainer.tuner.{lr_find,scale_batch_size} not setting the Trainer state properly (#7258)pickle.PickleError to catch all pickle errors (#6917)LightningModule.training_epoch_end was different from the object passed to the on_train_end_epoch hook (#6969)train_batch_end would be listed even when using a single optimizer and no truncated backprop through time steps (#6969)self.device not returning the correct device in replicas of data-parallel (#6414)lr_find trying beyond num_training steps and suggesting a too high learning rate (#7076)Trainer.fit calls (#7077)self.log not being reset correctly (#7055)CombinedLoader in distributed settings for validation / testing (#7102)WandbLogger when the run was initiated externally (#7106)num_sanity_val_steps affecting reproducibility of training data shuffling (#7014)fitting/evaluating/predicting (#7188)trainer.tuner.scale_batch_size(max_trials=0) would not return the correct batch size result (#7262)precision=16 and manual_optimization (#7228)BaseFinetuning properly reloading optimizer_states when using resume_from_checkpoint (#6891)parameters_to_ignore not properly set to DDPWrapper (#7239)fast_dev_run=True with the built-in ArgumentParser (#7240)IterableDataset that fails to produce a batch at the beginning of an epoch (#7294)LightningModule.save_hyperparameters() when attempting to save an empty container (#7268)apex not properly instantiated when running with ddp (#7274)state not moved to GPU (#7277)WandbLogger (#6989)rank_zero_only.rank when training is launched with SLURM and torchelastic (#6802)gradient_clip_algorithm has no effect (#6928)unfreeze_and_add_param_group expects modules rather than module (#6822)lr_find call (#6784)set_default_tensor_type to torch.DoubleTensor with precision=64 (#7108)NeptuneLogger.log_text(step=None) (#7194)@akihironitta, @alessiobonfiglio, @amisev, @amogkam, @ananthsub, @ArvinZhuang, @ashleve, @asnorkin, @awaelchli, @BloodAxe, @bmahlbrand, @Borda, @borisdayma, @camruta, @carmocca, @ceshine, @dbonner, @dhkim0225, @EdwardJB, @EliaCereda, @EricCousineau-TRI, @ethanwharris, @FlorianMF, @hemildesai, @ifsheldon, @kaushikb11, @mauvilsa, @maxfrei750, @mesejo, @ramonemiliani93, @rohitgr7, @s-rog, @sadiqj, @scart97, @SeanNaren, @shuyingsunshine21, @SkafteNicki, @SpontaneousDuck, @stllfe, @tchaton, @THasthika, @vballoli
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixing missing packaging package in dependencies, which was affecting the only installation to a very blank system.
Fixing missing packaging package in dependencies, which was affecting the only installation to a very blank system.
Fixed the order to call for world ranks & the root_device property in TPUSpawnPlugin
Added TPUSpawn + IterableDataset error message
Trainer instantiation (#6941)sync_dist for tpus (#6950)AttributeError for require_backward_grad_sync` when running manual optimization with sharded plugin (#6915)--gpus default for parser returned by Trainer.add_argparse_args (#6898)EarlyStopping logic when min_epochs or min_steps requirement is not met (#6705)BaseFinetuning.flatten_modules() was duplicating leaf node parameters (#6879)rank_zero_only.rank when training is launched with SLURM and torchelastic:
@ananthsub @awaelchli @ethanwharris @justusschock @kandluis @kaushikb11 @liob @SeanNaren @skmatz
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed resolve a bug with omegaconf and xm.save
xm.save (#6741)TensorBoardLogger would give a warning and not log correctly to a symbolic link save_dir (#6730)@awaelchli, @ethanwharris, @karthikprasad, @kaushikb11, @mibaumgartner, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Changed the behavior of on_epoch_start to run at the beginning of validation & test epoch
on_epoch_start to run at the beginning of validation & test epoch (#6498)step dictionary returns in callback_metrics. Use self.log_dict instead. (#6682)DummyLogger.log_hyperparams raising a TypeError when running with fast_dev_run=True (#6398)ModelCheckpoint (#6654)trainer.test freeze on TPUs (#6654)Trainer.predict (#6657)@awaelchli, @carmocca, @ethanwharris, @kaushikb11, @rohitgr7, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added Autocast in validation, test and predict modes for Native AMP
all_gather would not work correctly with tpu_cores=8 (#6587)@awaelchli, @Borda, @ethanwharris, @justusschock, @kaushikb11
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Changed the default of find_unused_parameters back to True in DDP and DDP Spawn
find_unused_parameters back to True in DDP and DDP Spawn (#6438)broadcast_object_list and add reduce_decision (#6410)DummyLogger.log_hyperparams raising a TypeError when running with fast_dev_run=True (#6398)Tuner.scale_batch_size not finding the batch size attribute in the datamodule (#5968)Trainer.predict (#6541)@awaelchli, @kaushikb11, @Palzer, @SeanNaren, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed ModelPruning(make_pruning_permanent=True) pruning buffers getting removed when saved during training
ModelPruning(make_pruning_permanent=True) pruning buffers getting removed when saved during training (#6073)_stable_1d_sort to work when n >= N (#6177)AttributeError when logger=None on TPU (#6221)emit_nvtx (#6260)trainer.test from best_path hangs after calling trainer.fit (#6272)SingleTPU calling all_gather (#6296)LightningOptimizer doesn't delete optimizer hooks (#6305)Trainer not resetting lightning_optimizers when calling Trainer.fit() multiple times (#6372)@awaelchli, @carmocca, @Chizuchizu, @frankier, @SeanNaren, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added checkpoint parameter to callback's on_save_checkpoint hook
checkpoint parameter to callback's on_save_checkpoint hook (#6072)backward, step, zero_grad to zero_grad, backward, step (#6147)val_check_interval < 1.0 (#6075)detach(), cpu(), to() (#6216)WandbLogger from dropping values (#5931)@akihironitta, @borisdayma, @carmocca, @dvolgyes, @SeanNaren, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed incorrect yield logic for the amp autocast context manager
@awaelchli, @SeanNaren, @carmocca
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Function stat_scores_multiple_classes is deprecated in favor of stat_scores
DataType, AverageMethod and MDMCAverageMethod enum in metrics (#5657)Accuracy metric now generalizes to Top-k accuracy for (multi-dimensional) multi-class inputs using the top_k parameter (#4838)Accuracy metric now enables the computation of subset accuracy for multi-label or multi-dimensional multi-class inputs with the subset_accuracy parameter (#4838)HammingDistance metric to compute the hamming distance (loss) (#4838)max_fpr parameter to auroc metric for computing partial auroc metric (#3790)StatScores metric to compute the number of true positives, false positives, true negatives and false negatives (#4839)R2Score metric (#5241)LambdaCallback (#5347)BackboneLambdaFinetuningCallback (#5377)all_gather supports collection (#5221)image_gradients functional metric to compute the image gradients of a given input image. (#5056)MetricCollection (#4318).clone() method to metrics (#4318)IoU class interface (#4704)on_post_move_to_device hookLightningModule (#5467)Recall and Precision metrics (and their functional counterparts recall and precision) can now be generalized to Recall@K and Precision@K with the use of top_k parameter (#4842)ModelPruning Callback (#5618, #5825, #6045)PyTorchProfiler (#5560)predict(...) for high performence predictions (#5579)on_before_batch_transfer and on_after_batch_transfer data hooks (#3671)PredictLoop object (#5752)QuantizationAwareTraining callback (#5706, #6040)LightningModule.configure_callbacks to enable the definition of model-specific callbacks (#5621)dim to PSNR metric for mean-squared-error reduction (#5957)log_graph to CometLogger (#5295)sync_step to Wandb logger (#5351)StochasticWeightAveraging callback (#5640)LightningDataModule.from_datasets(...) (#5133)PL_TORCH_DISTRIBUTED_BACKEND env variable to select backend (#5981)Trainer flag to activate Stochastic Weight Averaging (SWA) Trainer(stochastic_weight_avg=True) (#6038)stat_scores metric now calculates stat scores over all classes and gains new parameters, in line with the new StatScores metric (#4839)computer_vision_fine_tunning example to use BackboneLambdaFinetuningCallback (#5377)automatic casting for LoggerConnector metrics (#5218)iou [func] to allow float input (#4704)compute() method will no longer automatically call reset() (#5409)torchvision>=0.5 and torchtext>=0.5 (#5418)callbacks argument in Trainer to allow Callback input (#5446)find_unused_parameters to False in DDP (#5185)ModelCheckpoint version suffixes to start at 1 (#5008)progress_bar_refresh_rate Trainer argument in Google COLAB notebooks to 20 (#5516)LightningModule.global_rank, LightningModule.local_rank and LightningModule.logger read-only properties (#5730)ModelCheckpoint callbacks to run after all others to guarantee all states are saved to the checkpoint (#5731)LightningModule-wrapper logic to new plugins and accelerator (#5734)self.log in callbacks (#5094)on_train_batch_end, on_batch_end & on_train_epoch_end, on_epoch_end hooks (#5688)setup_training and remove test_mode (#5388)num_training_batches when insufficient limit_train_batches (#5703)EpochResultStore (#5522)lr_finder to check for attribute if not running fast_dev_run (#5990)toggle_model (#5771)MlflowLogger limit parameter value length to 250 char (#5893)stat_scores_multiple_classes is deprecated in favor of stat_scores (#4839)legacy pkg (#5645)LightningDistributedDataParallel in favor of new wrapper module LightningDistributedModule (#5185)LightningDataParallel in favor of new wrapper module LightningParallelModule (#5670)argparse_utils >> argparsemodel_utils >> model_helperswarning_utils >> warningsxla_device_utils >> xla_device'val_loss' to set the ModelCheckpoint monitor (#6012).get_model() with explicit .lightning_module property (#6035)accelerator_backend in favor of accelerator (#6034)filepath (#5321)Fbeta, f1_score and fbeta_score metrics (#5322)TrainResult (#5323)EvalResult (#5633)LoggerStages (#5673)ddp_cpu only with num_processes>1 (#5297)ModelCheckpoint when it already exists (#4861)DDPHPCAccelerator hangs in DDP construction by calling init_device (#5157)num_workers for Windows example (#5375).fit() calls ignore max_steps iteration bound (#5936)MisconfigurationError on unknown mode (#5255)ModelCheckpoint race condition in file existence check (#5155)requires_grad state after return None with multiple optimizers (#5738)on_epoch_end hook at the end of validation, test epoch (#5986)process_dataloader call for TPUSpawn when in distributed mode (#6015)hparams.yaml saved twice when using TensorBoardLogger (#5953)fairscale compatible with PT 1.8 (#5996)process_dataloader is called when tpu_cores > 1 to use Parallel DataLoader (#6015)@alanhdu, @ananthsub, @awaelchli, @Borda, @borisdayma, @carmocca, @ddrevicky, @deng-cy, @ducthienbui97, @justusschock, @kartik4949, @kaushikb11, @manipopopo, @marload, @neighthan, @peblair, @prampey, @pranjaldatta, @rohitgr7, @SeanNaren, @sid-sundrani, @SkafteNicki, @tadejsv, @tchaton, @teddykoker, @titu1994, @yuntai
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Nothing published for this version
Separate epoch validation from step validation
toggle_optimizers not handling all optimizer parameters (#5775)@ananthsub, @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed TensorBoardLogger not closing SummaryWriter on finalize
TensorBoardLogger not closing SummaryWriter on finalize (#5696)num_classes argument in F1 metric (#5663)log_dir property (#5537)ModelCheckpoint when checking if a checkpoint file exists (#5144)@awaelchli @guillochon @noamzilo @rohitgr7 @SkafteNicki @sumanthratna
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Increased TPU check timeout from 20s to 100s
step param in Neptune logger's log_metric method (#5510)on_train_batch_end instead of epoch_end outputs (#4369)toggle_optimizer to reset requires_grad state (#5574)Metric's state_dict not included when child modules (#5614)trainer.test() (#5138)hparams present (#4559)@awaelchli @bryant1410 @lezwon @manipopopo @PiotrJander @psinger @rnett @SeanNaren @swethmandava @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed a visual bug in the progress bar display initialization
on_train_batch_end in a callback with multiple optimizers (#5521)reinit_scheduler_properties with correct optimizer (#5519)val_check_interval with fast_dev_run (#5540)@awaelchli, @carmocca, @rohitgr7
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Changed deprecated enable_pl_optimizer=True
enable_pl_optimizer=True (#5244)transfer_batch_to_device for DDP with len(devices_ids) == 1 (#5195)not should_accumulate() during training (#5417)@ananthsub, @SeanNaren, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added a check for optimizer attached to lr_scheduler
lr_scheduler (#5338)filepaths to resume_from_checkpoint (#4402)resume_from_checkpoint while testing (#5161)log_momentum for adaptive optimizers in LearningRateMonitor (#5333)fast_dev_run (#5277)WORLD if None (#5125)trainer.test returning non-test metrics (#5214)--num-nodes on DDPSequentialPlugin (#5327)weights_summary (#5296)Trainer.test not using the latest best_model_path (#5161)hparams not using underlying filesystem (#5250)LightningOptimizer AMP bug (#5191)_flatten_dict (#5354)@8greg8, @haven-jeon, @kandluis, @marload, @rohitgr7, @tadejsv, @tarepan, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Updated DALIClassificationLoader to not use deprecated arguments
sync_dist=True (#5080)enable_pl_optimizer=False by default to temporarily fix AMP issues (#5163)ModelCheckpoint if it already exists (#4861)TensorRunningAccum (#5106)DALIClassificationLoader to not use deprecated arguments (#4925)torch.no_grad (#5124)@8greg8, @ananthsub, @borisdayma, @gan3sh500, @rohitgr7, @SeanNaren, @tchaton, @VinhLoiIT
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Add deprecated metric utility functions back to functional (#5067, #5068)
None in DDPAccelerator (#4915)LightningOptimizer to expose optimizer attributes (#5095)name key is used in the lr_scheduler dict (#5057)to_onnx and to_torchscript (#4378)lr_scheduler dict (#5057)@Borda, @carmocca, @hemildesai, @rohitgr7, @s-rog, @tarepan, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated prefix argument in ModelCheckpoint
Lightning 1.1 is out! You can now train models with twice the parameters and zero code changes with the new sharded model training! We also have a new plugin for sequential model parallelism, more logging options, and a lot of improvements! Release highlights: https://bit.ly/3gyLZpP
Learn more about sharded training: https://bit.ly/2W3hgI0
ModelCheckpoints (#4383)ConfusionMatrix class interface (#4348)current_score to ModelCheckpoint.on_save_checkpoint (#4721)self.log in train and evaluation for epoch end hooks (#4913)hparams (#4647)prefix argument in loggers (#4557)PrecisionRecallCurve, ROC, AveragePrecision class metric (#4549)Apex and NativeAMP as Precision plugins (#4355)DALI MNIST example (#3721)sharded plugin for DDP for multi-GPU training memory optimizations (#4773)experiment_id to the NeptuneLogger (#3462)Pytorch Geometric integration example with Lightning (#4568)all_gather method to LightningModule which allows gradient-based tensor synchronizations for use-cases such as negative sampling. (#5012)self.log in most functions (#4969)ModelCheckpoint (#4977)multiclass_roc and multiclass_precision_recall_curve, use roc and precision_recall_curve instead (#4549)fast_dev_run=True (#3903)reinit arg to True anymore and creates a run only when needed (#4648)automatic_optimization to be a model attribute (#4602)Simple Profiler report to order by percentage time spent + num calls (#4880)fast_dev_run to accept integer representing num_batches (#4629)prefix argument in ModelCheckpoint (#4765)self.hparams = ... (#4813)mode='auto' from ModelCheckpoint and EarlyStopping (#4695)reorder parameter of the auc metric (#5004)LoggerConnector to have logged metrics on root device in DP (#4138)gather_all (#4907)PYTHONPATH for DDP test model (#4528)@ananyahjha93, @awaelchli, @blatr, @Borda, @borisdayma, @carmocca, @ddrevicky, @george-gca, @gianscarpe, @irustandi, @janhenriklambrechts, @jeremyjordan, @justusschock, @lezwon, @rohitgr7, @s-rog, @SeanNaren, @SkafteNicki, @tadejsv, @tchaton, @williamFalcon, @zippeurfou
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added casting to python types for numpy scalars when logging hparams
hparams (#4647)F1 class metric (#4656)step=trainer.global_step in LearningRateMonitor independently of logging_interval (#4376)state_dict (#4685)Fbeta >> FBeta (#4656)PYTHONWARNINGS (#4700)hparams dict casting when omegaconf is available (#4770)batch_arg_name to all calls to _adjust_batch_sizebug (#4812)torchtext data to GPU (#4785)@awaelchli, @jonashaag, @jungwhank, @M-Salti, @moi90, @pgagarinov, @s-rog, @Samyak2, @SkafteNicki, @teddykoker, @ydcjeff
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added lambda closure to manual_optimizer_step
manual_optimizer_step (#4618)persistent default mode to False (#4685)sync_dist=True on CPU (#4626)setup callback hook to correctly pass the LightningModule through (#4608)hparams inside (#4662)split_idx set by LoggerConnector in on_trainer_init to Trainer (#4697)@ananthsub, @Borda, @SeanNaren, @SkafteNicki, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added metrics aggregation in Horovod and fixed early stopping
manual_optimizer_step which work with AMP Native and accumulated_grad_batches (#4485)persistent(mode) method to metrics, to enable and disable metric states being added to state_dict (#4482)fsspec to tuner (#4458)hpc_load (#4526)lightning_getattr, lightning_hasattr not finding the correct attributes in datamodule (#4347)manual_optimization_step (#4485)MisconfigurationException with warning in ModelCheckpoint Callback (#4560)is_picklable by catching AttributeError (#4508)@dscarmo, @jtamir, @kazhang, @maxjeblick, @rohitgr7, @SkafteNicki, @tarepan, @tchaton, @tgaddair, @williamFalcon
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated passing ModelCheckpoint instance to checkpoint_callback Trainer argument
tpu_device_exists to ensure process does not hang indefinitely (#4340)Trainer step (#4405)on_after_backward is called only when optimizer_step is being called (#4439)track_and_norm_grad into training loop and called only when optimizer_step is being called (#4439)ModelCheckpoint instance to checkpoint_callback Trainer argument (#4336)auto_select_gpus=True with gpus=-1 (#4209)limit_train_batches=0 (#4371)on_after_backward (#4439)@ananthsub, @awaelchli, @borisdayma, @carmocca, @justusschock, @lezwon, @rohitgr7, @SeanNaren, @SkafteNicki, @ssaru, @tchaton, @ydcjeff
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated filepath in ModelCheckpoint
dirpath and filename parameter in ModelCheckpoint (#4213)strict option to the scheduler dictionary (#3586)fsspec support for profilers (#4162)Trainer.add_argparse_args (#4344)Trainer's profiler parameter (#3656)configure_optimizers returns (#3587)validation_step (#4130)replace_sampler_ddp=True with a distributed sampler already added (#4273)WandbLogger.log_hyperparams (#4320)filepath in ModelCheckpoint (#4213)reorder parameter of the auc metric (#4237)Trainer's profiler parameter (#3656)ddp_accelerator (#4323)WandbLogger not uploading checkpoint artifacts at the end of training (#4341)@ananthsub, @awaelchli, @carmocca, @ddrevicky, @louis-she, @mauvilsa, @rohitgr7, @SeanNaren, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added persistent flag to Metric.add_state
Metric.add_state (#4195)checkpoint_connector.hpc_save in SLURM (#4217)hparams assign in init (#4189)@Borda, @EspenHa, @teddykoker
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixes the last major bugs for validation logging. Also removes duplicate charts for metric / metric_loss. Doing this minor release because correct val
Fixes the last major bugs for validation logging. Also removes duplicate charts for metric / metric_loss. Doing this minor release because correct validation metrics logging is critical.
to_torchscript (#4142)on_load_checkpoint before loading state_dict (#4057)validation_step() (#4169)hparams saving - save the state when save_hyperparameters() is called [in __init__] (#4163)hparams to yaml (#4158)@Borda, @NumesSanguis, @rohitgr7, @williamFalcon
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Obligatory post 1.0 minor release. Main fix is to make Lightning module fully compatible with Jit (had some edge-cases we had not covered).
Obligatory post 1.0 minor release. Main fix is to make Lightning module fully compatible with Jit (had some edge-cases we had not covered).
Removed deprecated trainer flags: overfit_pct, log_save_interval, row_log_interval
...
Trainer (#4022)LightningModule.toggle_optimizer (#4058)LightningModule.manual_backward (#4063)Accelerator (#4066)output argument from *_batch_end hooks (#3965, #3966)output argument from *_epoch_end hooks (#3967)overfit_pct, log_save_interval, row_log_interval (#3969)trainer argument in LightningModule.backward [#4056)current_epoch property update to reflect true epoch number inside LightningDataModule, when reload_dataloaders_every_epoch=True. (#3974)on_load_checkpoint hook is called (#3996)@ananyahjha93, @Borda, @edenlightning, @hbredin, @rohitgr7, @SkafteNicki, @teddykoker, @williamFalcon
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Results objects are deprecated (we hated them too haha)
This release is a buffer in case 1.0 breaks any compatibility for people who upgrade. 0.10.0 has all the bug fixes and features of 1.0 but is 100% backward compatible. The 1.0 release following in the next 24 hours.
The major changes are:
To log:
def any_step(...):
self.log('something', i_computed)
Separately, return whatever you want from methods:
def training_step(...):
return loss
or
def training_step(...):
return {'loss': loss, 'whatever': [1, 'want']}
LightningModule.to_torchscript to support exporting as ScriptModule (#3258)hparams (#2874)ModelCheckpoint.to_yaml method (#3048)ModelCheckpoint monitor to be None, meaning it will always save ([3630)broadcast to TPUBackend (#3814)XLADeviceUtils class to check XLA device type (#3274)xxx_step to backend (#3118)forward (#3119)__step (#3120)___step (#3123)test_mode for if so we can split later (#3129)___step_end hooks (#3130)run_evaluation (#3156)_evaluate fx (#3197)Trainer.fit hook clean up (#3198)Trainer.tune() (#3293)run_pretrain_routine -> setup_training (#3294)prepare_data to data connector (#3307)lr_finder (#3434)checkpoint_on to simplify (#3571)LightningDataModule (#3684).log to lightning module (#3686, #3699, #3701, #3704, #3715)torchelastic from DDP (#3810)init_slurm_connection causing hostname errors (#3856)LearningRateLogger to LearningRateMonitor (#3251)fsspec instead of gfile for all IO (#3320)
torch.load for fsspec load in DDP spawn backend (#3787)torch.load for fsspec load in cloud_io loading (#3692)to_disk() to use remote filepaths with fsspec (#3930)fsspec open (#3801)fsspec is inconsistant when doing fs.ls (#3805)GPUStatsMonitor to improve training speed (#3257)remove_bg bool to ignore_index optional int (#3098)save_top_k and save_last to None in ModelCheckpoint (#3680)row_log_interval and log_save_interval are now based on training loop's global_step instead of epoch-internal batch index (#3667)ModelCheckpoint monitor to be None (#3633)None model checkpoint default (#3669)best_model_path if checkpoint_callback is None (#2962)raise .. from .. to explicitly chain exceptions (#3750)TrainResult and EvalResult, use self.log and self.write from the LightningModule to log metrics and write predictions. training_step can now only return a scalar (for the loss) or a dictionary with anything you want. (#3681)early_stop_callback Trainer argument (#3845)row_log_interval >> log_every_n_steps and log_save_interval >> flush_logs_every_n_steps (#3748)EmbeddingSimilarity metric (#3349, [#3358)ModelCheckpoint with save_top_k=-1 option not tracking the best models when a monitor metric is available (#3735)Accuracy metric for zero target tensor (#3764)reduction to class_reduction in classification metrics (#3322)class_reduction similar to sklearn for classification metrics (#3322)on_train_batch_start hook to end epoch early (#3700)num_sanity_val_steps is clipped to limit_val_batches (#2917)GpuUsageLogger to work on different platforms (#3008)auto_lr_find parameter (#3151)batch_outputs with optimizer frequencies (#3229)LightningModule.datamodule when using auto_scale_batch_size (#3266)experiment_id from MLFlow only once instead of each training loop (#3394)overfit_batches which now correctly disables shuffling for the training loader. (#3501)row_log_interval > 1 (#3489)ModelCheckpoint name formatting ([3164)t() to transpose() as XLA devices do not support .t() on 1-dim tensor (#3252)gather_all_tensors cross GPUs in DDP (#3319)training_epoch_end hook is used (#3673)overfit_batches > 0 and distributed_backend = "ddp" (#3534)DDPSpawnBackend when using seed_everything in main process (#3335)ModelCheckpoint period to actually save every period epochs (#3630)val_progress_bar total with num_sanity_val_steps (#3751)current_epoch to dumped_params (#3261)current_epoch and global_step properties mismatch between Trainer and LightningModule (#3785)tbptt_reduce_fx when non-floating tensors are logged (#3796)TrainerEvaluationLoopMixin activates model.train() at the end (#3858)overfit_batches when using with multiple val/test_dataloaders (#3857)training_step to return None (#3862)load_from_checkpoint (#2776)batch_sizes when Dataloader returns a dict with multiple tensors (#3668)validation_step (#3947)@abrahambotros, @akihironitta, @ananthsub, @ananyahjha93, @awaelchli, @Borda, @c00k1ez, @carmocca, @f4hy, @GimmickNG, @jbschiratti, @justusschock, @LeeJZh, @lezwon, @Lucas-Steinmann, @maxjeblick, @monney, @mpariente, @nateraw, @nrupatunga, @patrickorlando, @PhilJd, @rohitgr7, @s-rog, @ShomyLiu, @SkafteNicki, @Sordie, @teddykoker, @tgaddair, @Vozf, @williamFalcon, @XDynames, @ydcjeff
If we forgot someone due to not matching the commit email with GitHub account, let us know :]
Fixed shell injection vulnerability in subprocess call
The newest PyTorch Lightning release includes final API clean-up with better data decoupling and shorter logging syntax.
Were happy to release PyTorch Lightning 0.9 today, which contains many great new features, more bugfixes than any release we ever had, but most importantly it introduced our mostly final API changes! Lightning is being adopted by top researchers and AI labs around the world, and we are working hard to make sure we provide a smooth experience and support for all the latest best practices.
CSVLogger (#2721)Trainer(num_sanity_val_steps=-1) to check all validation data before training (#2246)EvalResult support for train and val. loop (#2615, #2651)LightningDataModule (#2668)sklearn metrics: AveragePrecision, BalancedAccuracy, CohenKappaScore, DCG, Hamming, Hinge, Jaccard, MeanAbsoluteError, MeanSquaredError, MeanSquaredLogError, MedianAbsoluteError, R2Score, MeanPoissonDeviance, MeanGammaDeviance, MeanTweedieDeviance, ExplainedVariance (#2562)limit_{mode}_batches (int) to work with infinite dataloader (IterableDataset) (#2840)hparams (#2846)Trainer (#2541)strict=False for load_from_checkpoint (#2819)transfer_batch_to_device to the LightningDataModule (#3038)accelerator module:
.comet.config file for CometLogger (#1913)setup and teardown (#2850)gfile to support remote directories (#2164)**DictConfig for hparam serialization (#2519)ckpt_path, which will now be set by weights_save_path (#2681)data_loaderon_sanity_check_start and loading load_from_metricspytorch_lightning.loggingshow_progress_bar, num_tpu_cores, use_amp, print_nan_gradsnum_accumulation_stepsaccumulate_grad_batches for last batch (#2853)dtype and device properties not getting updated in submodules (#2657)fast_dev_run to run for all dataloaders (#2581)save_dir in loggers getting ignored by default value of weights_save_path when user did not specify weights_save_path (#2681)weights_save_path getting ignored when logger=False is passed to Trainer (#2681)LoggerCollection (#2723)torchtext.data.Field and include_lengths is True (#2689)accumulate_grad_batches > 1 (#2738)CUDA_VISIBLE_DEVICES (#2739, #2796)num_classes warning in metrics (#2781)hparams compatibility (#2821)ModelCheckpoint not saving the latest information when save_last=True (#2881)non_blocking=True when transferring a batch object that does not support it (#2910)val_step argument to metrics (#2986)Trainer.test() to stall in DDP mode (#2997)@ananthsub, @ananyahjha93, @awaelchli, @bkhakshoor, @Borda, @ethanwharris, @f4hy, @groadabike, @ibeltagy, @justusschock, @lezwon, @nateraw, @neighthan, @nsarang, @PhilJd, @pwwang, @rohitgr7, @romesco, @ruotianluo, @shijianjian, @SkafteNicki, @tgaddair, @thschaaf, @williamFalcon, @xmotli02, @ydcjeff, @yukw777, @zerogerc
If we forgot someone due to not matching commit email with GitHub account, let us know :]
The point of this release is more bug fixes ahead of v 1.0.0. We now have CI tests on TPU thanks to @zcain117 from Google! :slightly_smiling_face: Thi
The point of this release is more bug fixes ahead of v 1.0.0. We now have CI tests on TPU thanks to @zcain117 from Google! :slightly_smiling_face: This means we fixed many TPU bugs we hadn’t caught before because we had no tests. In addition, we fixed:
TensorBoardLogger and CometLogger pickleable (#2518)MLflowLogger creating multiple run folders (#2502)argparse default value bug (#2526).fit() returning last not best weights in "ddp_spawn" (#2565).test() (#2512, #2570)@anthonytec2, @awaelchli, @bernardomig, @Borda, @EspenHa, @HHousen, @InCogNiTo124, @rohitgr7, @williamFalcon
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added reduce ddp results on eval
IterableDataset has __len__ defined (#2437)@awaelchli, @Borda, @olineumann, @williamFalcon
### Fixed - Fixed AMP wrong call (593837e1da24ff6c942b24ed803fc1496a304609) - Fixed batch typo
Fixing critical bugs in newly added hooks and hparams assignment. The recommended data following:
Fixing critical bugs in newly added hooks and hparams assignment.
The recommended data following:
prepare_data to download and process the dataset.setup to do splits, and build your model internalsload_from_checkpoint path detected as URL bug (#2244)hparams - remove frame inspection on self.hparams (#2253)Deprecated tags_csv in favor of hparams_file
Highlights of this release are adding support for TorchElastic enables distributed PyTorch training jobs to be executed in a fault-tolerant and elastic manner; auto-scaling of batch size; new transfer learning example; an option to provide seed to random generators to ensure reproducibility.
Trainer.fit() and Trainer.test() to reflect that also a list of dataloaders can be passed in (#1723).training_epoch_end (#1724)NeptuneLogger to work with distributed_backend=ddp (#1753)load_from_ckpt (#1797)torchelastic (#1811, #1818)store_true for bool args (#1822, #1842)non-blocking for device transfers to GPU (#1843)batch_size < num_gpus (#1609)namedtuple to tuple when transferring the batch to target device (#1589)hparams as a keyword argument to LightningModule when loading from checkpoint (#1639)tags_csv in favor of hparams_file (#1271)on_load_checkpoint() when resuming from a checkpoint (#1666)_reset_eval_dataloader() for IterableDataset (#1560)root_gpu property (#1669)global_step affects other loggers (#1492)version_ when it shouldn't (#1748)hparam logging with metrics (#1647)@ashwinb, @awaelchli, @Borda, @cmpute, @festeh, @jbschiratti, @justusschock, @kepler, @kumuji, @nanddalal, @nathanbreitsch, @olineumann, @pitercl, @rohitgr7, @S-aiueo32, @SkafteNicki, @tgaddair, @tullie, @tw991, @williamFalcon, @ybrovman, @yukw777
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed CPU DDP breaking change and DDP change
We made a few changes to Callbacks to test ops on detached GPU tensors to avoid CPU transfer. However, it made callbacks unpicklable which will crash DDP.
This release fixes that core issue
@justusschock, @quinor, @williamFalcon
We had a few (subtle) bugs that affected DDP and a few key things in 0.7.2 so we released 0.7.3 to fix them because they are critical for DDP. sorry a
We had a few (subtle) bugs that affected DDP and a few key things in 0.7.2 so we released 0.7.3 to fix them because they are critical for DDP. sorry about that! still, no API changes, but please do skip straight to 0.7.3 upgrade for those fixes
rank_zero_warn for warning only in rank 0 (#1428)DistributedSampler for DDP training (#1425)run_training_batch (#1431)@alsrgv, @Borda, @williamFalcon
Monir bug fix with print issues and data_loader
Monir bug fix with print issues and data_loader (#1080)
Deprecated max_nb_epochs and min_nb_epochs
This release focused on a ton of bug fixes, small optimizations to training but most importantly, clean new docs!
We have released New documentation, please bear with us as we fix broken links and patch in missing pieces. This project moved to new org PyTorchLightning, so no longer the root sits on WilliamFalcon/PyTorchLightning. We have added own custom Tensorboard logger as default logger. We have upgrade Continues Integration to speed up the automatic testing. We have fixed GAN training - supporting multiple optimizers.
resume_from_checkpoint argument (#516)ReduceLROnPlateau scheduler (#320)O2 in conjunction with Data Parallel (#493)save_top_k) to save the top k models in the ModelCheckpoint class (#128)on_train_start and on_train_end hooks to ModelHooks (#598)TensorBoardLogger (#607)map_location argument to load_from_metrics and load_from_checkpoint (#625)val_percent_check=0 (#649)NeptuneLogger class (#648)WandbLogger class (#627)step_idx to step, epoch_idx to epoch, max_num_epochs to max_epochs and min_num_epochs to min_epochs (#589)Trainer atributes: (#567)
total_batch_nb to total_batches,nb_val_batches to num_val_batches,nb_training_batches to num_training_batches,max_nb_epochs to max_epochs,min_nb_epochs to min_epochs,nb_test_batches to num_test_batches,nb_val_batches to num_val_batches (#567)TensorBoardLogger (#609)max_nb_epochs and min_nb_epochs (#567)on_sanity_check_start hook in ModelHooks (#598)save_best_only argument from ModelCheckpoint, use save_top_k=1 instead (#128)gpus=0 or gpus=[] (#561)print_nan_gradients when some parameters do not require gradient (#579)val_check_interval < 1.0 in Trainer (#492)CometLogger object that would cause it to not work properly (#481)-1 from on_batch_start following an early exit or when the batch was None (#509)truncated_bptt > 1 (#532)IterableDataset (#547](https://github.com/PyTorchLightning/pytorch-lightning/pull/547)).item was called on non-tensor objects (#602)Trainer.train would crash on an uninitialized variable if the trainer was run after resuming from a checkpoint that was already at max_epochs (#608)num_training_batches and num_test_batches would sometimes be rounded down to zero (#649)num_training_batches (#653).copy method (#701)log_gpu_memory=True in Python 3.6 (#715)on_train_end was not called when early stopping (#723)@akhti, @alumae, @awaelchli, @Borda, @borisdayma, @ctlaltdefeat, @dreamgonfly, @elliotwaite, @fdiehl, @goodok, @haossr, @HarshSharma12, @Ir1d, @jakubczakon, @jeffling, @kuynzereb, @MartinPernus, @matthew-z, @MikeScarp, @mpariente, @neggert, @rwesterman, @ryanwongsa, @schwobr, @tullie, @vikmary, @VSJMilewski, @williamFalcon, @YehCF
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →