NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1101 most downloaded on PyPI
PyTorch native Metrics
Last release 6 months ago
09 Mar 2026
Release timing varies
gaps range from 1 weeks to 6 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
6 years old
73 releases · first in 2021
One column per quarter.
Fixed torch.sort currently does not support bool dtype on CUDA
torch.sort currently does not support bool dtype on CUDA (#665)MAP metric (#673)@OlofHarrysson, @tkupek, @twsl
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Migrate MAP metrics from pycocotools to PyTorch
torch.topk instead of torch.argsort in retrieval precision for speedup (#627)average=weighted on GPU (#606)forward in compositional metrics (#645)@Callidior, @SkafteNicki, @tkupek, @twsl, @zuoxingdong
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated torchmetrics.functional.self_supervised.embedding_similarity in favour of new pairwise submodule
We are excited to announce that Torchmetrics v0.6 is now publicly available. TorchMetrics v0.6 does not focus on specific domains but adds a ton of new metrics to several domains, thus increasing the number of metrics in the repository to over 60! Not only have v0.6 added metrics within already covered domains, but we also add support for two new: Pairwise metrics and detection.
https://devblog.pytorchlightning.ai/torchmetrics-v0-6-more-metrics-than-ever-e98c3983621e
TorchMetrics v0.6 offers a new set of metrics in its functional backend for calculating pairwise distances. Given a tensor X with shape [N,d] (N observations, each in d dimensions), a pairwise metric calculates [N,N] matrix of all possible combinations between the rows of X.
TorchMetrics v0.6 now includes a detection package that provides for the MAP metric. The implementation essentially wraps pycocotools around securing that we get the correct value, but with the benefit of now being able to scale to multiple devices (as any other metric in TorchMetrics).
In the audio package, we have two new metrics: Perceptual Evaluation of Speech Quality (PESQ) and Short Term Objective Intelligibility (STOI). Both metrics can be used to assert speech quality.
In the retrieval package, we also have two new metrics: R-precision and Hit-rate. R-precision corresponds to recall at the R-th position of the query. The hit rate is the ratio of the total number of hits returned as a result of a query (hits) to the total number of hits returned.
The text package also receives an update in the form of two new metrics: Sacre BLEU score and character error rate. Sacre BLUE score provides and more systematic way of comparing BLUE scores across tasks. The character error rate is similar to the word error rate but instead calculates if a given algorithm has correctly predicted a sentence based on a character-by-character comparison.
The regression package got a single new metric in the form of the Tweedie deviance score metric. Deviance scores are generally a better measure of fit than measures such as squared error when trying to model data coming from highly screwed distributions.
Finally, we have added five new metrics for simple aggregation: SumMetric, MeanMetric, MinMetric, MaxMetric, CatMetric. All five metrics take in a single input (either native python floats or torch.Tensor) and keep track of the sum, average, min, etc. These new aggregation metrics are especially useful in combination with self.log from lightning if you want to log something other than the average of the metric you are tracking.
RetrievalRPrecision (#577)RetrievalHitRate (#576)SacreBLEUScore (#546)CharErrorRate (#575)MAP (mean average precision) metric to new detection package (#467)nDCG metric (#437)average argument to AveragePrecision metric for reducing multi-label and multi-class problems (#477)MultioutputWrapper (#510)higher_is_better as constant attribute (#544)higher_is_better to rest of codebase (#584)SumMetric, MeanMetric, CatMetric, MinMetric, MaxMetric (#506)pairwise_cosine_similaritypairwise_euclidean_distancepairwise_linear_similaritypairwise_manhatten_distanceAveragePrecision will now as default output the macro average for multilabel and multiclass problems (#477)half, double, float will no longer change the dtype of the metric states. Use metric.set_dtype instead (#493)AverageMeter to MeanMetric (#506)is_differentiable from property to a constant attribute (#551)ROC and AUROC will no longer throw an error when either the positive or negative class is missing. Instead, return 0 scores and give a warningtorchmetrics.functional.self_supervised.embedding_similarity in favour of new pairwise submoduledtype property (#493)F1 with average='macro' and ignore_index!=None (#495)pit by using the returned first result to initialize device and type (#533)SSIM metric using too much memory (#539)device property was not properly updated when the metric was a child of a module (#542)@an1lam, @Borda, @karthikrangasai, @lucadiliello, @mahinlma, @Obus, @quancs, @SkafteNicki, @stancld, @tkupek
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Added device and dtype properties
device and dtype properties (#462)TextTester class for robustly testing text metrics (#450)nDCG metric (#437)rouge-score as dependency for text package (#443)jiwer as dependency for text package (#446)bert-score as dependency for text package (#473)SpearmanCorrCoef metric (#448)BootStrapper metrics not working on GPU (#462)SSIM metric (#474)@justusschock, @karthikrangasai, @kingyiusuen, @Obus, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
This release includes general improvements to the library and new metrics within the NLP domain.
This release includes general improvements to the library and new metrics within the NLP domain.
https://devblog.pytorchlightning.ai/torchmetrics-v0-5-nlp-metrics-f4232467b0c5
Natural language processing is arguably one of the most exciting areas of machine learning, with models such as BERT, ROBERTA, GPT-3 etc., really pushing what automated text translation, recognition, and generation systems are capable of.
With the introduction of these models, many metrics have been proposed that measure how well these models perform. TorchMetrics v0.5 includes 4 such metrics: BERT score, BLEU, ROUGE and WER.
MetricTracker wrapper metric for keeping track of the same metric over multiple epochs (#238)nDCG metric for target with values larger than 1 (#349)nDCG metric (#378)None as reduction option in CosineSimilarity metric (#400)AveragePrecision (#386)psnr and ssim from functional.regression.* to functional.image.* (#382)image_gradient from functional.image_gradients to functional.image.gradients (#381)R2Score from regression.r2score to regression.r2 (#371)torch.argmax instead of torch.topk when k=1 for better performance (#419)r2score >> r2_score and kldivergence >> kl_divergence in functional (#371)bleu_score from functional.nlp to functional.text.bleu (#360)threshold has to be in (0,1) range to support logit input (#351, #401)preds could not be bigger than num_classes to support logit input (#357)regression.psnr and regression.ssim (#382):functional.mean_relative_errornum_thresholds argument in BinnedPrecisionRecallCurveaverage='macro' would lead to wrong result if a class was missing (#303)weighted, multi-class AUROC computation to allow for 0 observations of some class, as contribution to final AUROC is 0 (#376)_forward_cache and _computed attributes are also moved to the correct device if metric is moved (#413)IoU metric when using ignore_index argument (#328)@BeyondTheProof, @Borda, @CSautier, @discort, @edwardclem, @gagan3012, @hugoperrin, @karthikrangasai, @paul-grundmann, @quancs, @rajs96, @SkafteNicki, @vatch123
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Extend typing (#330, #332, #333, #335, #314)
is_sync logic to Metric (#339)Deprecated torchmetrics.functional.mean_relative_error
https://devblog.pytorchlightning.ai/torchmetrics-v0-4-introducing-multimedia-metrics-e6380a3ad354
The first highlight of v0.4.0 is a set of 3 new metrics for calculating for evaluating audio data: Scale-invariant signal-to-distortion ratio, Scale-invariant signal-to-noise ratio, and signal-to-noise ratio. All these metrics take a predicted audio tensor and a target tensor, both with the shape [...,time] and calculate the metric over the time axis.
Version v0.4.0 also includes a completely new image package. Since its initial 0.2.0 release, Torchmetrics has had both PSNR and SSIM in its regression module, metrics that can be used to evaluate image quality. With the image module, we are adding three new metrics for evaluating the quality of generative models (such as GANS): Inception score (IS), Fréchet inception distance (FID) and kernel inception distance (KID).
In addition to the new audio and image package, we also want to highlight a couple of features:
.softmax(dim=-1) anymore!sync and sync_context methods that allow the user full control over when metric states are synced. Note that we still automatically do this whenever calling the compute method.is_differentiable property has been adopted by many more of our metrics!Big thanks to all community members for their contributions and feedback. A special thanks to @quancs for leading the development of the new audio package.
add_metrics method to MetricCollection for adding additional metrics after initialization (#221)dist_reduce_fx="cat" to reduce communication cost (#217)AUROC when num_classes is not provided for multiclass input (#244)Accuracy, Precision, Recall, FBeta, F1, StatScore, Hamming, ConfusionMatrix metrics (#200)MeanAbsolutePercentageError(MAPE) metric. (#248)squared argument to MeanSquaredError for computing RMSE (#249)is_differentiable property to ConfusionMatrix, F1, FBeta, Hamming, Hinge, IOU, MatthewsCorrcoef, Precision, Recall, PrecisionRecallCurve, ROC, StatScores (#253)sync and sync_context methods for manually controlling when metric states are synced (#302)KLDivergence metric (#247)reset method is called (#260)precision, recall, precision_recall, fbeta, f1, accuracy, and specificity (#204)torch.jit.unused to MetricCollection forward (#307)thresholds argument to binned metrics for manually controlling the thresholds (#322)torchmetrics.functional.mean_relative_error (#248)num_thresholds argument in BinnedPrecisionRecallCurve (#322)is_multiclass (#319)dtype of modular metrics after reset has been called (#243)matthews_corrcoef to correctly match formula (#321)@AnselmC, @arvindmuralie77, @bhadreshpsavani, @Borda, @GiannisVagionakis, @hassiahk, @IgorHoholko, @johannespitz, @justusschock, @maximsch2, @pranjaldatta, @quancs, @simran2905, @SkafteNicki, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
add_metrics method to MetricCollection for adding additional metrics after initialization (#221)dist_reduce_fx="cat" to reduce communication cost (#217)AUROC when num_classes is not provided for multiclass input (#244)Accuracy, Precision, Recall, FBeta, F1, StatScore, Hamming, ConfusionMatrix metrics (#200)squared argument to MeanSquaredError for computing RMSE (#249)is_differentiable property to ConfusionMatrix, F1, FBeta, Hamming, Hinge, IOU, MatthewsCorrcoef, Precision, Recall, PrecisionRecallCurve, ROC, StatScores (#253)sync and sync_context methods for manually controlling when metric states are synced (#302)reset method is called (#260)precision, recall, precision_recall, fbeta, f1, accuracy, and specificity (#204)torch.jit.unused to MetricCollection forward (#307)thresholds argument to binned metrics for manually controlling the thresholds (#322)functional.mean_relative_error, use functional.mean_absolute_percentage_error (#248)num_thresholds argument in BinnedPrecisionRecallCurve (#322)is_multiclass (#319)dtype of modular metrics after reset has been called (#243)matthews_corrcoef to correctly match formula (#321)Added is_differentiable property:
is_differentiable property:
AUC, AUROC, CohenKappa and AveragePrecision (#178)PearsonCorrCoef, SpearmanCorrcoef, R2Score and ExplainedVariance (#225)MetricCollection should return metrics with prefix on items(), keys() (#209)compute before update will now give an warning (#164)numpy as dependency (#212)load_state_dict() (#202)PSNR not working with DDP (#214)AUROC metric for large input (#230)@bhadreshpsavani, @hlin09, @maximsch2, @SkafteNicki, @tchaton
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Cleaning remaining inconsistency and fix PL develop integration (#191, #192, #193, #194)
Cleaning remaining inconsistency and fix PL develop integration (#191, #192, #193, #194)
Information retrieval (IR) metrics are used to evaluate how well a system is retrieving information from a database or from a collection of documents.
Information retrieval (IR) metrics are used to evaluate how well a system is retrieving information from a database or from a collection of documents. This is the case with search engines, where a query provided by the user is compared with many possible results, some of which are relevant and some are not.
When you query a search engine, you hope that results that could be useful are ranked higher on the results page. However, each query is usually compared with a different set of documents. For this reason, we had to implement a mechanism to allow users to easily compute the IR metrics in cases where each query is compared with a different number of possible candidates.
For this reason, IR metrics feature an additional argument called indexes that say to which query a prediction refers to. In the end, all query-document pairs are grouped by query index and then the final result is computed as the average of the metric over each group.
In total 6 new metrics have been added for doing information retrieval:
Special thanks go to @lucadiliello, for implementing all IR.
In addition to expanding our collection to the field of information retrieval, this release also includes new metrics for the classification domain:
The current implementation of the AveragePrecision and PrecisionRecallCurve has the drawback that it saves all predictions and targets in memory to correctly calculate the metric value. These metrics now receive a binned version that calculates the value at fixed thresholds. This is less precise than original implementations but also much more memory efficient.
Special thanks go to @SkafteNicki, for letting all this happen.
https://devblog.pytorchlightning.ai/torchmetrics-v0-3-0-information-retrieval-metrics-and-more-c55265e9b94f
BootStrapper to easily calculate confidence intervals for metrics (#101)RetrievalMAP (PL^5032)RetrievalMRR (#119)RetrievalPrecision (#139)RetrievalRecall (#146)RetrievalNormalizedDCG (#160)RetrievalFallOut (#161)CohenKappa (#69)MatthewsCorrcoef (#98)PearsonCorrcoef (#157)SpearmanCorrcoef (#158)Hinge (#120)average='micro' as an option in AUROC for multilabel problems (#110)ROC metric (#114)half precision (#77, #135)AverageMeter for ad-hoc averages of values (#138)prefix argument to MetricCollection (#70)__getitem__ as metric arithmetic operation (#142)is_differentiable to metrics and test for differentiability (#154)average, ignore_index and mdmc_average in Accuracy metric (#166)postfix arg to MetricCollection (#188)ExplainedVariance from storing all preds/targets to tracking 5 statistics (#68)confusionmatrix for multilabel data to better match multilabel_confusion_matrix from sklearn (#134)reset method to use detach.clone() instead of deepcopy when resetting to default (#163)MetricCollection will now always be in deterministic order (#173)MetricCollection pass metrics as arguments (#176)is_multiclass -> multiclass (#162)_stable_1d_sort to work when n>=N (PL^6177)_computed attribute not being correctly reset (#147)@alanhdu, @arvindmuralie77, @bhadreshpsavani, @Borda, @ethanwharris, @lucadiliello, @maximsch2, @SkafteNicki, @thomasgaudelet, @victorjoos
If we forgot someone due to not matching commit email with GitHub account, let us know :]
BootStrapper to easily calculate confidence intervals for metrics (#101)average='micro' as an option in AUROC for multilabel problems (#110)ROC metric (#114)half precision (#77,
#135
)AverageMeter for ad-hoc averages of values (#138)prefix argument to MetricCollection (#70)__getitem__ as metric arithmetic operation (#142)is_differentiable to metrics and test for differentiability (#154)average, ignore_index and mdmc_average in Accuracy metric (#166)postfix arg to MetricCollection (#188)ExplainedVariance from storing all preds/targets to tracking 5 statistics (#68)confusionmatrix for multilabel data to better match multilabel_confusion_matrix from sklearn (#134)reset method to use detach.clone() instead of deepcopy when resetting to default (#163)MetricCollection will now always be in deterministic order (#173)MetricCollection pass metrics as arguments (#176)is_multiclass -> multiclass (#162)_stable_1d_sort to work when n>=N (PL^6177)_computed attribute not being correctly reset (#147)TorchMetrics is a collection of 25+ PyTorch metrics implementations and an easy-to-use API to create custom metrics. It offers:
TorchMetrics is a collection of 25+ PyTorch metrics implementations and an easy-to-use API to create custom metrics. It offers:
You can use TorchMetrics in any PyTorch model, or with in PyTorch Lightning to enjoy additional features:
Similar to torch.nn, most metrics have both a module-based and a functional version. The functional version implements the basic operations required for computing each metric. They are simple python functions that as input take torch.tensors and return the corresponding metric as a torch.tensor.
import torch
# import our library
import torchmetrics
# simulate a classification problem
preds = torch.randn(10, 5).softmax(dim=-1)
target = torch.randint(5, (10,))
acc = torchmetrics.functional.accuracy(preds, target)
Nearly all functional metrics have a corresponding module-based metric that calls it a functional counterpart underneath. The module-based metrics are characterized by having one or more internal metrics states (similar to the parameters of the PyTorch module) that allow them to offer additional functionalities:
import torch
# import our library
import torchmetrics
# initialize metric
metric = torchmetrics.Accuracy()
n_batches = 10
for i in range(n_batches):
# simulate a classification problem
preds = torch.randn(10, 5).softmax(dim=-1)
target = torch.randint(5, (10,))
# metric on current batch
acc = metric(preds, target)
print(f"Accuracy on batch {i}: {acc}")
# metric on all batches using custom accumulation
acc = metric.compute()
print(f"Accuracy on all data: {acc}")
And many more!
@Borda, @SkafteNicki, @williamFalcon, @teddykoker, @justusschock, @tadejsv, @edenlightning, @ydcjeff, @ddrevicky, @ananyahjha93, @awaelchli, @rohitgr7, @akihironitta, @manipopopo, @Diuven, @arnaudgelas, @s-rog, @c00k1ez, @tgaddair, @elias-ramzi, @cuent, @jpcarzolio, @bryant1410, @shivdhar, @Sordie, @krzysztofwos, @abhik-99, @bernardomig, @peblair, @InCogNiTo124, @j-dsouza, @pranjaldatta, @ananthsub, @deng-cy, @abhinavg97, @tridao, @prampey, @abrahambotros, @ozen, @ShomyLiu, @yuntai, @pwwang
If we forgot someone due to not matching commit email with GitHub account, let us know :]
MetricCollection (#19)Your coding agent can read these notes before it upgrades. Set up the MCP server →