NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1157 most downloaded on PyPI
PyTorch Lightning is the lightweight PyTorch wrapper for ML researchers. Scale your models. Write less boilerplate.
Last release 26 days ago
10 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
223 releases · first in 2019
The main focus of this release was on adding flexibility and generalization to support broad research cases.
The main focus of this release was on adding flexibility and generalization to support broad research cases.
Next release will be Dec 7th (every 30 days).
@lorenzoFabbri @tullie @myleott @ashwinb @shootingsoul @vreis These features were added to support FAIR, FAIAR and broader ML across other FB teams.
In general, we can expose any part that isn't exposed yet where someone might want to override the lightning implementation.
Trainer(truncated_bptt_steps=2)
# return iterabledataset
def train_dataloader(...):
ds = IterableDataset(...)
return Dataloader(ds)
# set validation to a fix number of batches
# (checks val every 100 train epochs)
Trainer(val_check_interval=100)
def backward(self, use_amp, loss, optimizer):
"""
Override backward with your own implementation if you need to
:param use_amp: Whether amp was requested or not
:param loss: Loss is already scaled by accumulated grads
:param optimizer: Current optimizer being used
:return:
"""
if use_amp:
with amp.scale_loss(loss, optimizer) as scaled_loss:
scaled_loss.backward()
else:
loss.backward()
def configure_ddp(self, model, device_ids):
"""
Override to init DDP in a different way or use your own wrapper.
Must return model.
:param model:
:param device_ids:
:return: DDP wrapped model
"""
model = LightningDistributedDataParallel(
model,
device_ids=device_ids,
find_unused_parameters=True
)
return model
def init_ddp_connection(self, proc_rank, world_size):
"""
Connect all procs in the world using the env:// init
Use the first node as the root address
"""
# use slurm job id for the port number
# guarantees unique ports across jobs from same grid search
try:
# use the last 4 numbers in the job id as the id
default_port = os.environ['SLURM_JOB_ID']
default_port = default_port[-4:]
# all ports should be in the 10k+ range
default_port = int(default_port) + 15000
except Exception as e:
default_port = 12910
# if user gave a port number, use that one instead
try:
default_port = os.environ['MASTER_PORT']
except Exception:
os.environ['MASTER_PORT'] = str(default_port)
# figure out the root node addr
try:
root_node = os.environ['SLURM_NODELIST'].split(' ')[0]
except Exception:
root_node = '127.0.0.2'
root_node = self.trainer.resolve_root_node_address(root_node)
os.environ['MASTER_ADDR'] = root_node
dist.init_process_group('nccl', rank=proc_rank, world_size=world_size)
def configure_apex(self, amp, model, optimizers, amp_level):
"""
Override to init AMP your own way
Must return a model and list of optimizers
:param amp:
:param model:
:param optimizers:
:param amp_level:
:return: Apex wrapped model and optimizers
"""
model, optimizers = amp.initialize(
model, optimizers, opt_level=amp_level,
)
return model, optimizers
training_endtraining_step (performed on each GPU with a portion of the batch),
to do something with the outputs of all batches on the node (ie: negative sampling).Trainer(distributed_backend='ddp2')
def training_step(...):
# x is 1/nb_gpus of the full batch
out = model(x)
return {'out': out}
def training_end(self, outputs):
# all_outs has outs from ALL gpus
all_outs = outputs['out']
loss = softmax(all_outs)
return {'loss': loss}
Thank you to the amazing contributor community! Especially @neggert and @Borda for reviewing PRs and taking care of a good number of Github issues. The community is thriving and has really embraced making Lightning better.
Great job everyone!
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
All trainers now have a default logger, early stopping and checkpoint object. To modify the behavior, pass in your own versions of those.
All trainers now have a default logger, early stopping and checkpoint object. To modify the behavior, pass in your own versions of those.
Trainer(distributed_backend='ddp2')
training_step and validation_end now return two separate dicts, one for the progress bar and one for logging.
Added options to memory printing: 'min_max' logs only the max/min memory use. 'all' logs all the GPUs on the root node.
This release has breaking API changes. See #124 for all details. Syntax changes are: ` in trainer options use: train, test, val for data: val_dataload
This release has breaking API changes. See #124 for all details. Syntax changes are:
in trainer options use: train, test, val
for data: val_dataloader, test_dataloader, train_dataloader
data_batch -> batch
prog -> progress
gradient_clip -> gradient_clip_val
add_log_row_interval -> row_log_interval
This release does the following:
This release does the following:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
validation_step, val_dataloader are now optional.
Nothing published for this version
Nothing published for this version
Nothing published for this version
0.4.0 is the first public release after a short period testing with public users. Thanks for all the help ironing out bugs to get Lightning to run on
0.4.0 is the first public release after a short period testing with public users. Thanks for all the help ironing out bugs to get Lightning to run on everything from notebooks to local to server machines.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Added a decorator to do lazy loading internally:
Simplified data loader.
Added a decorator to do lazy loading internally:
Old:
@property
def tng_dataloader(self):
if self._tng_dataloader is None:
self._tng_dataloader = DataLoader(...)
return self.tng_dataloder
Now:
@ptl.data_loader
def tng_dataloader(self):
return DataLoader(...)
Full tests that run multiple models in different configs
Fully tested!
Includes:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →