NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1101 most downloaded on PyPI
PyTorch native Metrics
Last release 6 months ago
09 Mar 2026
Release timing varies
gaps range from 1 weeks to 6 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
6 years old
73 releases · first in 2021
One column per quarter.
Fixed n-d slicing deprecation warning
average="macro" (#3042)pkg_resources with packaging (#3329)Metric base class (#3316)logAUC (#3295)_safe_divide by creating tensor directly on device (#3284)@adaliaramon, @bhimrazy, @GdoongMathew, @Isalia20, @KyleMylonakisProtopia, @VijayVignesh1
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: v1.8.0...v1.9.0
Fixed BinaryPrecisionRecallCurve now returns NaN for precision when no predictions meet a threshold
BinaryPrecisionRecallCurve now returns NaN for precision when no predictions meet a threshold (#3227)precision_at_fixed_recall and recall_at_fixed_precision to correctly return NaN thresholds when recall/precision conditions are not met (#3226)@iamkulbhushansingh
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.8.1...v1.8.2
Added reduction='none' to vif metric
reduction='none' to vif metric (#3196)sigmoid normalization in BinaryPrecisionRecallCurve (#3182)@iamkulbhushansingh, @PussyCat0700, @simonreise
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.8.0...v1.8.1
The upcoming TorchMetrics v1.8.0 release introduces three flagship metrics, each designed to address critical evaluation needs in real-world applicati
The upcoming TorchMetrics v1.8.0 release introduces three flagship metrics, each designed to address critical evaluation needs in real-world applications.
Video Multi-Method Assessment Fusion (VMAF) brings a perceptual video-quality score that closely mirrors human judgment, powering streaming services such as Netflix and YouTube to optimize encoding ladders for consistent viewer experiences and enabling video-restoration labs to quantify improvements achieved by denoising and super-resolution algorithms.
Continuous Ranked Probability Score (CRPS) enables comprehensive evaluation of full predictive distributions rather than point estimates; meteorological centers leverage CRPS to benchmark probabilistic precipitation and temperature forecasts, improving public weather alerts, while energy companies apply it to assess uncertainty in load-demand predictions and refine grid management and trading strategies.
Lip Vertex Error (LVE) measures the discrepancy between predicted and ground-truth lip landmarks to quantify audio-visual synchronization. Localization studios use LVE to validate lip-sync accuracy during film dubbing, while AR/VR developers integrate it into avatar pipelines to ensure natural mouth movements in real-time virtual meetings and social experiences.
VMAF metric to new video domain (#2991)CRPS in regression domain (#3024)aggregation_level argument to DiceScore (#3018)reduction="none" to LearnedPerceptualImagePatchSimilarity (#3053)str input for functional interface of bert_score (#3056)BERTScore to evaluate hypotheses against multiple references (#3069)Lip Vertex Error (LVE) in multimodal domain (#3090)antialias argument to FID metric (#3177)mixed input format to segmentation metrics (#3176)data_range argument in PSNR metric to be a required argument (#3178)zero_division argument from DiceScore (#3018)@nkaenzig, @rittik9, @simonreise, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.7.0...v1.8.0
Improved numerical stability of pearson's correlation coefficient
dist_reduce_fx when reduction=None for distributed training (#3162, #3166)_pearson_corrcoef_update (#3168)@AymenKallala, @gratus907, @Isalia20, @rittik9
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.7.3...v1.7.4
Fixed: ensure WrapperMetric resets wrapped_metric state
WrapperMetric resets wrapped_metric state (#3123)top_k in multiclass_accuracy (#3117)pycocotools 2.0.10 (#3131)@rittik9
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.7.2...v1.7.3
Enhance: improve performance of _rank_data
_rank_data (#3103)UnboundLocalError in MatthewsCorrCoef (#3059)byte dtype with custom encoders (#3064)ignore_index in MultilabelExactMatch (#3085)@ahmedhshahin, @gratus907, @rittik9, @ZhiyuanChen
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.7.1...v1.7.2
Enhance Support Adding a MetricCollection to Another MetricCollection in add_metrics Function
MetricCollection to Another MetricCollection in add_metrics Function (#3032)MeanIOU (#2892)MulticlassAccuracy when top_k>1 (#3039)@Isalia20, @rittik9, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.7.0...v1.7.1
The upcoming release of TorchMetrics is set to deliver a range of innovative features and enhancements across multiple domains, further solidifying it
The upcoming release of TorchMetrics is set to deliver a range of innovative features and enhancements across multiple domains, further solidifying its position as a leading tool for machine learning metrics. In the image domain, significant additions include the ARNIQA and DeepImageStructureAndTextureSimilarity metrics, which provide new insights into image quality and similarity. Additionally, the CLIPScore metric now supports more models and processors, expanding its versatility in image-text alignment tasks.
Beyond image analysis, the regression package welcomes the JensenShannonDivergence metric, offering a powerful tool for comparing probability distributions. The clustering package also sees a notable update with the introduction of the ClusterAccuracy metric, which helps evaluate the performance of clustering algorithms more effectively.
In the realm of classification, the Equal Error Rate (EER) metric has been added, providing a crucial measure for assessing the performance of classification models, particularly in scenarios where false positives and false negatives have different costs. Furthermore, the MeanAveragePrecision metric now includes a functional interface, enhancing its usability and flexibility for users.
These updates collectively enhance the capabilities of TorchMetrics, making it an even more comprehensive and indispensable resource for machine learning practitioners and researchers.
ARNIQA metric (#2953)DeepImageStructureAndTextureSimilarity (#2993)CLIPScore (#2978)JensenShannonDivergence metric to regression package (#2992)ClusterAccuracy metric to cluster package (#2777)Equal Error Rate (EER) to classification package (#3013)MeanAveragePrecision metric (#3011)num_classes optional for one-hot inputs in MeanIoU (#3012)Dice from classification (#3017)IndexError in MultiClassAccuracy when using top_k with single sample (#3021)@Isalia20, @LorenzoAgnolucci, @nathanpainchaud, @rittik9, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.6.0...v1.7.0
Fixed logic in how metric states referencing is handled in MetricCollection
MetricCollection (#2990)@SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.6.2...v1.6.3
Added zero_division argument to DiceScore in segmentation package
zero_division argument to DiceScore in segmentation package (#2860)cache_session to DNSMOS metric to control caching behavior (#2974)disable option to nan_strategy in basic aggregation metrics (#2943)num_classes optional for classification in case of micro averaging (#2841)Clip_Score to calculate similarities between same modalities (#2875)DiceScore when there is zero overlap between predictions and targets (#2860)MeanAveragePrecision for average="micro" when 0 label is not present (#2968)PearsonCorrCoef when input is constant (#2975)MetricCollection.update gives identical results (#2944)kwargs in PIT metric for permutation wise mode (#2977)_final_aggregation function for PearsonCorrCoef (#2980)@baskrahmer, @czmrand, @rbedyakin, @rittik9, @SkafteNicki, @wooseopkim
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.6.1...v1.6.2
Enabled specifying weights path for FID
Device2Host caused by comm with device and host (#2840)top_k for multiclassf1score with one-hot encoding (#2839)@Isalia20, @nkaenzig, @podgorki, @rittik9, @yuvalkirstain, @zhaozheng09
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.6.0...v1.6.1
Deprecated Dice from classification metrics
The latest release of TorchMetrics introduces several significant enhancements and new features that will greatly benefit users across various domains. This update includes the addition of new metrics and methods that enhance the library's functionality and usability.
One of the key additions is the NISQA audio metric, which provides advanced capabilities for evaluating audio quality. In the classification domain, the new LogAUC and NegativePredictiveValue metrics offer improved tools for assessing model performance, particularly in imbalanced datasets. For regression tasks, the NormalizedRootMeanSquaredError metric has been introduced, providing a normalized measure of prediction accuracy that is less sensitive to outliers.
In the field of image segmentation, the new Dice metric enhances the evaluation of segmentation models by providing a robust measure of overlap between predicted and ground truth masks. Additionally, the merge_state method has been added to the Metric class, allowing for more efficient state management and aggregation across multiple devices or processes.
Furthermore, this release includes support for the propagation of the autograd graph in Distributed Data-Parallel (DDP) settings, enabling more efficient and scalable training of models across multiple GPUs. These enhancements collectively make TorchMetrics a more powerful and versatile tool for machine learning practitioners, enabling more accurate and efficient model evaluation across a wide range of applications.
NISQA (#2792)LogAUC (#2377)NegativePredictiveValue (#2433)NormalizedRootMeanSquaredError (#2442)Dice (#2725)merge_state to Metric (#2786)KLDivergence (#2800)num_outputs in R2Score (#2800)Dice + GeneralizedDice for 2d index tensors (#2832)rouge_score with accumulate='best' (#2830)@Borda, @cw-tan, @philgzl, @rittik9, @SkafteNicki
1.5.0Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.5.0...v1.6.0
Fixed iou scores in detection for either empty predictions/targets leading to wrong scores
numpy 2+ support (#2804)MetricCollection compatibility with torch.jit.script (#2813)np.Inf for numpy 2.0+ (#2826)@adamjstewart, @Borda, @SkafteNicki, @StalkerShurik, @yurithefury
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.5.1...v1.5.2
Changing _modules dict type in Pytorch 2.5 preventing to fail collections metrics
_modules dict type in Pytorch 2.5 preventing to fail collections metrics (#2793)@bfolie
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.5.0...v1.5.1
Deprecated num_outputs in R2Score
Shape metrics are quantitative methods used to assess and compare the geometric properties of objects, often in datasets that represent shapes. One such metric is the Procrustes Disparity, which measures the sum of the squared differences between two datasets after applying a Procrustes transformation. This transformation involves scaling, rotating, and translating the datasets to achieve optimal alignment. The Procrustes Disparity is particularly useful when comparing datasets that are similar in structure but not perfectly aligned, allowing for more meaningful comparison by minimizing differences due to orientation or size.
HausdorffDistance (#2122)DNSMOS (#2525)ProcrustesDistance (#2723)MetricInputTransformer wrapper (#2392)input_format argument to segmentation metrics (#2572)multi-output support for MAE metric (#2605)truncation argument to BERTScore (#2776)InfoLM class to dynamically set higher_is_better (#2674)num_outputs in R2Score (#2705)IoU metric for single empty prediction tensors (#2780)PSNR calculation for integer type input images (#2788)@Astraightrain, @grahamannett, @lgienapp, @matsumotosan, @quancs, @SkafteNicki
1.4.0Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.0...v1.5.0
Fixed for Pearson changes inputs
PESQ metric where NoUtterancesError prevented calculating on a batch of data (#2753)MatthewsCorrCoef (#2743)@Borda, @SkafteNicki, @veera-puthiran-14082
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.2...v1.4.3
Fixed wrong aggregation in segmentation.MeanIoU
Chrf implementation (#2701)segmentation.MeanIoU (#2698)scipy (#2733)prefix/postfix works in MultitaskWrapper (#2722)torch.unique with dim=None (#2650)@Borda, @petertheprocess, @rittik9, @SkafteNicki, @vkinakh
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.1...v1.4.2
Calculate the text color of ConfusionMatrix plot based on luminance
ConfusionMatrix plot based on luminance (#2590)_safe_divide to allow Accuracy to run on the GPU (#2640)Chrf implementation due to licensing issues with the upstream package (#2668)MetricCollection when using compute groups and compute is called more than once (#2571)panoptic_quality(..., return_per_class=True) output (#2548)BootstrapWrapper not being reset correctly (#2574)ClasswiseWrapper and MetricCollection with custom _filter_kwargs method (#2575)_cumsum helper function in multi-gpu (#2636)MeanAveragePrecision.coco_to_tm (#2588)@Borda, @gxy-gxy, @i-aki-y, @ndrwrbgs, @relativityhd, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.0...v1.4.1
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.0...v1.4.0.post0
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.4.0...v1.4.0.post0
In Torchmetrics v1.4, we are happy to introduce a new domain of metrics to the library: segmentation metrics. Segmentation metrics are used to evaluat
In Torchmetrics v1.4, we are happy to introduce a new domain of metrics to the library: segmentation metrics. Segmentation metrics are used to evaluate how well segmentation algorithms are performing, e.g., algorithms that take in an image and pixel-by-pixel decide what kind of object it is. These kind of algorithms are necessary in applications such as self driven cars. Segmentations are closely related to classification metrics, but for now, in Torchmetrics, expect the input to be formatted differently; see the documentation for more info. For now, MeanIoU and GeneralizedDiceScore have been added to the subpackage, with many more to follow in upcoming releases of Torchmetrics. We are happy to receive any feedback on metrics to add in the future or the user interface for the new segmentation metrics.
Torchmetrics v1.3 adds new metrics to the classification and image subpackage and has multiple bug fixes and other quality-of-life improvements. We refer to the changelog for the complete list of changes.
SensitivityAtSpecificity metric to classification subpackage (#2217)QualityWithNoReference metric to image subpackage (#2288)MeanIoU (#1236)GeneralizedDiceScore (#1090)PanopticQuality metric (#2381)pretty-errors for improving error prints (#2431)torch.float weighted networks for FID and KID calculations (#2483)zero_division argument to selected classification metrics (#2198)__getattr__ and __setattr__ of ClasswiseWrapper more general (#2424)ERGAS metric (#2498)BootStrapper wrapper not working with kwargs provided argument (#2503)MeanAveragePrecision when requested (#2501)binary_average_precision when only negative samples are provided (#2507)@baskrahmer, @Borda, @ChristophReich1996, @daniel-code, @furkan-celik, @i-aki-y, @jlcsilva, @NielsRogge, @oguz-hanoglu, @SkafteNicki, @ywchan2005
1.3.0If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.3.0...v1.4.0
Fixed negative variance estimates in certain image metrics
top_k>1 and average="macro" for classification metrics (#2423)PrecisionRecallCurve.plot methods (#2437)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.3.1...v1.3.2
@Borda, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed how backprop is handled in LPIPS metric
LPIPS metric (#2326)MultitaskWrapper not being able to be logged in lightning when using metric collections (#2349)Perplexity metric (#2346)FeatureShare not being moved to the correct device (#2348)MeanAveragePrecision with custom max det thresholds (#2367)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.3.0...v1.3.1
@Borda, @fschlatt, @JonasVerbickas, @nsmlzl, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.3.0...v1.3.0.post0
Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.3.0...v1.3.0.post0
Deprecated metric._update_called
TorchMetrics v1.3 is out now! This release introduces seven new metrics in the different subdomains of TorchMetrics, adding some nice features to already established metrics. In this blogpost, we present the new metrics with short code samples.
We are happy to see the continued adoption of TorchMetrics in over 19,000 Github repositories projects, and we are proud to release that we have passed 1,800 GitHub stars.
The retrieval domain has received one new metric in this release: RetrievalAUROC. This metric calculates the Area Under the Receiver Operation Curve for document retrieval data. It is similar to the standard AUROC metric from classification but also supports the additional indexes argument that all retrieval metrics support.
from torch import tensor
from torchmetrics.retrieval import RetrievalAUROC
indexes = tensor([0, 0, 0, 1, 1, 1, 1])
preds = tensor([0.2, 0.3, 0.5, 0.1, 0.3, 0.5, 0.2])
target = tensor([False, False, True, False, True, False, True])
r_auroc = RetrievalAUROC()
r_auroc(preds, target, indexes=indexes)
# tensor(0.7500)
The image subdomain is receiving two new metrics in v1.3, which brings the total number image-specific metrics in TorchMetrics to 21! As with other metrics, these two new metrics work by comparing a predicted image tensor to a ground truth image, but they focus on different properties for their metric calculation.
The first metrics is SpatialCorrelationCoefficient. As the name indicates this metric focuses on how well the spatial structure of the predicted image correlates with the ground truth image.
import torch
torch.manual_seed(42)
from torchmetrics.image import SpatialCorrelationCoefficient as SCC
preds = torch.randn([32, 3, 64, 64])
target = torch.randn([32, 3, 64, 64])
scc = SCC()
scc(preds, target)
# tensor(0.0023)
The second metrics is SpatialDistortionIndex compares the spatial structure of the images, and is especially useful for evaluating multi spectral images
import torch
from torchmetrics.image import SpatialDistortionIndex
preds = torch.rand([16, 3, 32, 32])
target = {
'ms': torch.rand([16, 3, 16, 16]),
'pan': torch.rand([16, 3, 32, 32]),
}
sdi = SpatialDistortionIndex()
sdi(preds, target)
# tensor(0.0090)
A new wrapper metric called FeatureShare has also been added. This can be seen as a specialized version of MetricCollection that can be combined with metrics that use a neural network as part of their metric calculation. For example, FrechetInceptionDistance , InceptionScore, KernelInceptionDistance all, by default, use an inception network for their metric calculations. When these metrics were combined inside a MetricCollection, the underlying neural network was still called three times, which is quite redundant and wastes resources. In principle, it should be possible only to call it once and then propagate the value to all metrics, which is exactly what the FeatureShare wrapper solves.
import torch
from torchmetrics.wrappers import FeatureShare
from torchmetrics import MetricCollection
from torchmetrics.image import FrechetInceptionDistance, KernelInceptionDistance
def fs_wrapper():
fs = FeatureShare([FrechetInceptionDistance(), KernelInceptionDistance(subset_size=10, subsets=2)])
fs.update(torch.randint(255, (50, 3, 64, 64), dtype=torch.uint8), real=True)
fs.update(torch.randint(255, (50, 3, 64, 64), dtype=torch.uint8), real=False)
fs.compute()
def mc_wrapper():
mc = MetricCollection([FrechetInceptionDistance(), KernelInceptionDistance(subset_size=10, subsets=2)])
mc.update(torch.randint(255, (50, 3, 64, 64), dtype=torch.uint8), real=True)
mc.update(torch.randint(255, (50, 3, 64, 64), dtype=torch.uint8), real=False)
mc.compute()
# lets compare (using ipython timeit function)
% timeit fs_wrapper()
# 8.38 s ± 564 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
% timeit mc_wrapper()
# 13.8 s ± 232 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
This will most likely be significantly faster than the alternative metric collection, as show in the code example.
In v1.2, several new arguments were added to MeanAveragePrecision metric from the detection package. This metric has seen a further small improvement in that the argument extended_summary=True also returns confidence scores. The confidence scores are the score assigned by the model on how confident a given predicted bounding box belongs to a certain class.
from torch import tensor
from torchmetrics.detection import MeanAveragePrecision
# enable extended summary
map_metric = MeanAveragePrecision(extended_summary=True)
preds = [
{
"boxes": torch.tensor([[0.5, 0.5, 1, 1]]),
"scores": torch.tensor([1.0]),
"labels": torch.tensor([0]),
}
]
target = [
{"boxes": torch.tensor([[0, 0, 1, 1]]), "labels": torch.tensor([0])}
]
map_metric.update(preds, target)
result = map_metric.compute()
# new confidence score can be found in the "score" key
confidence_scores = result["scores"]
# in this case confidence_score will have shape (10, 101, 1, 4, 3)
# because
# * We are by default evaluating for 10 different IoU thresholds
# * We evaluate the PR-curve based on 101 linearly spaced locations
# * We only have 1 class (see the labels tensor)
# * There are 4 area sizes we evaluate on (small, medium, large and all)
# * By default `max_detection_thresholds=[1,10,100]` meaning we evaluate for 3 values
From v1.3 all retrieval metrics now support an argument called aggregation that determines how the metric should be aggregated over different documents. The supported options are "mean", "median", "max", "min" with the default value being "mean" which is fully backward compatible with earlier versions of TorchMetrics.
from torch import tensor
from torchmetrics.retrieval import RetrievalHitRate
indexes = tensor([0, 0, 0, 1, 1, 1, 1])
preds = tensor([0.2, 0.3, 0.5, 0.1, 0.3, 0.5, 0.2])
target = tensor([True, False, False, False, True, False, True])
hr2 = RetrievalHitRate(aggregation="max")
hr2(preds, target, indexes=indexes)
# tensor(1.000)
Finally, the SacreBLEU metric from the text domain now supports even more tokenizers: "ja-mecab", "ko-mecab", "flores101", "flores200”.
Users should be aware that from v1.3, TorchMetrics now only supports v1.10 of Pytorch and up (before v1.8). We always try to provide support for Pytorch releases for up to two years.
There have been several bug fixes related to numerical stability in several metrics. For this reason, we always recommend that users use the most recent version of Torchmetrics for the best experience.
Thank you!
As always, we offer a big thank you to all of our community members for their contributions and feedback. Please open an issue in the repo if you have any recommendations for the next metrics we should tackle.
If you want to ask a question or join us in expanding Torchmetrics, please join our discord server, where you can ask questions and get guidance in the #torchmetrics channel.
🔥 Check out the documentation and code! 🚀
SacreBLEU metric (#2068)MultiTaskWrapper directly with lightnings log_dict method (#2213)FeatureShare wrapper to share submodules containing feature extractors between metrics (#2120)SpatialDistortionIndex (#2260)CriticalSuccessIndex (#2257)Spatial Correlation Coefficient (#2248)average argument to multiclass versions of PrecisionRecallCurve and ROC (#2084)extended_summary=True in MeanAveragePrecision (#2212)RetrievalAUROC metric (#2251)aggregate argument to retrieval metrics (#2220)segmentation.utils for future segmentation metrics (#2105)PrecisionRecallCurve to be consistent with scikit-learn (#2183)metric._update_called (#2141)specicity_at_sensitivity in favour of specificity_at_sensitivity (#2199)Running metrics (#2256)FID metric (#2277)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.2.0...v1.3.0
@Borda, @HoseinAkbarzadeh, @matsumotosan, @miskfi, @oguz-hanoglu, @SkafteNicki, @stancld, @ywchan2005
1.2.0If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added error if NoTrainInceptionV3 is being initialized without torch-fidelity not being installed
NoTrainInceptionV3 is being initialized without torch-fidelity not being installed (#2143)v2.1 (#2142)SpectralAngleMapper and UniversalImageQualityIndex to be tensors (#2089)arange and repeat for deterministic bincount (#2184)lpips third-party package as dependency of LearnedPerceptualImagePatchSimilarity metric (#2230)LearnedPerceptualImagePatchSimilarity metric (#2144)UniversalImageQualityIndex metric (#2222)MeanAveragePrecision with pycocotools backend when too little max_detection_thresholds are provided (#2219)LearnedPerceptualImagePatchSimilarity functional metric (#2234)Metric._reduce_states(...) when using dist_sync_fn="cat" (#2226)CosineSimilarity where 2d is expected but 1d input was given (#2241)MetricCollection when using compute groups and compute is called more than once (#2211)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.2.0...v1.2.1
@Borda, @jankng, @kyle-dorman, @SkafteNicki, @tanguymagne
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Torchmetrics v1.2 is out now! The latest release includes 11 new metrics within a new subdomain: *Clustering*. In this blog post, we briefly explain w
Torchmetrics v1.2 is out now! The latest release includes 11 new metrics within a new subdomain: Clustering. In this blog post, we briefly explain what clustering is, why it’s a useful measure and newly added metrics that can be used with code samples.
Clustering is an unsupervised learning technique. The term unsupervised here refers to the fact that we do not have ground truth targets as we do in classification. The primary goal of clustering is to discover hidden patterns or structures within data without prior knowledge about the meaning or importance of particular features. Thus, clustering is a form of data exploration compared to supervised learning, where the goal is “just” to predict if a data point belongs to one class.
The key goal of clustering algorithms is to split data into clusters/sets where data points from the same cluster are more similar to each other than any other points from the remaining clusters. Some of the most common and widely used clustering algorithms are K-Means, Hierarchical clustering, and Gaussian Mixture Models (GMM).
An objective quality evaluation/measure is required regardless of the clustering algorithm or internal optimization criterion used. In general, we can divide all clustering metrics into two categories: extrinsic metrics and intrinsic metrics.
Extrinsic metrics are characterized by requirements of some ground truth labeling, even if used for an unsupervised method. This may seem counter-intuitive at first as we, by clustering definition, do not use such ground truth labeling. However, most clustering algorithms are still developed on datasets with labels available, so these metrics use this fact as an advantage.
In contrast, intrinsic metrics do not need any ground truth information. These metrics estimate inter-cluster consistency (cohesion of all points assigned to a single set) compared to other clusters (separation). This is often done by comparing the distance in the embedding space.
MeanAveragePrecision, the most widely used metric for object detection in computer vision, now supports two new arguments: average and backend.
The average argument controls averaging over multiple classes. By the core definition, the default way is macro averaging, where the metric is calculated for each class separately and then averaged together. This will continue to be the default in Torchmetrics, but now we also support the setting average="micro". Every object under this setting is essentially considered to be the same class, and the returned value is, therefore, calculated simultaneously over all objects.
The second argument - backend, is important, as it indicates what computational backend will be used for the internal computations. Since MeanAveragePrecision is not a simple metric to compute, and we value the correctness of our metric, we rely on some third-party library to do the internal computations. By default, we rely on users to have the official pycocotools installed, but with the new argument, we will also be supporting other backends.
MutualInformationScore (#2008)RandScore (#2025)NormalizedMutualInfoScore (#2029)AdjustedRandScore (#2032)CalinskiHarabaszScore (#2036)DunnIndex (#2049)HomogeneityScore (#2053)CompletenessScore (#2053)VMeasureScore (#2053)FowlkesMallowsIndex (#2066)AdjustedMutualInfoScore (#2058)DaviesBouldinScore (#2071)backend argument to MeanAveragePrecision (#2034)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.1.0...v1.2.0
v1.1.0@matsumotosan, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed tie breaking in ndcg metric
BootStrapper when very few samples were evaluated that could lead to crash (#2052)RecallAtFixedPrecision for large batch sizes (#2042)MetricCollection used with custom metrics have prefix/postfix attributes (#2070)@GlavitsBalazs, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added average argument to MeanAveragePrecision
average argument to MeanAveragePrecision (#2018)PearsonCorrCoef is updated on single samples at a time (#2019)MetricCollection when used with multiple metrics that return dicts with same keys (#2027)class_metrics=True resulting in wrong values (#1924)higher_is_better, is_differentiable for some metrics (#2028)@adamjstewart, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
In version v1.1 of Torchmetrics, in total five new metrics have been added, bringing the total number of metrics up to 128! In particular, we have two
In version v1.1 of Torchmetrics, in total five new metrics have been added, bringing the total number of metrics up to 128! In particular, we have two new exciting metrics for evaluating your favorite generative models for images.
Introduced in the famous StyleGAN paper back in 2018 the Perceptual path length metric is used to quantify how smoothly a generator manages to interpolate between points in its latent space. Why does the smoothness of the latent space of your generative model matter? Assume you find an image at some point in your latent space that generates an image you like, but you would like to see if you could find a better one if you slightly change the latent point it was generated from. If your latent space could be smoother, this because very hard because even small changes to the latent point can lead to large changes in the generated image.
CLIP image quality assessment (CLIPIQA) is a very recently proposed metric in this paper. The metrics build on the OpenAI CLIP model, which is a multi-modal model for connecting text and images. The core idea behind the metric is that different properties of an image can be assessed by measuring how similar the CLIP embedding of the image is to the respective CLIP embedding of a positive and negative prompt for that given property.
VisualInformationFidelity has been added to the image package. The first proposed in this paper can be used to automatically assess the quality of images in a perceptual manner.
EditDistance have been added to the text package. A very classical metric for text that simply measures the amount of characters that need to be substituted, inserted, or deleted, to transform the predicted text into the reference text.
SourceAggregatedSignalDistortionRatio has been added to the audio package. Metric was originally proposed in this paper and is an improvement over the classical Signal-to-Distortion Ratio (SDR) metric (also found in torchmetrics) that provides more stable gradients during training when trying to train models for style source separation.
VisualInformationFidelity to image package (#1830)EditDistance to text package (#1906)top_k argument to RetrievalMRR in retrieval package (#1961)"segm" and "bbox" detection in MeanAveragePrecision at the same time (#1928)PerceptualPathLength to image package (#1939)MeanSquaredError (#1937)extended_summary to MeanAveragePrecision such that precision, recall, iou can be easily returned (#1983)ClipScore if long captions are detected and truncate (#2001)CLIPImageQualityAssessment to multimodal package (#1931)metric_state to all metrics for users to investigate currently stored tensors in memory (#2006)Full Changelog: https://github.com/Lightning-AI/torchmetrics/compare/v1.0.0...v1.1.0
v1.0.0@bojobo, @lucadiliello, @quancs, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added warning to MeanAveragePrecision if too many detections are observed
MeanAveragePrecision if too many detections are observed (#1978)multidim_average="samplewise" in classification metrics (#1977)@borda, @SkafteNicki^n, @Vivswan
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added warning to PearsonCorrCoeff if input has a very small variance for its given dtype
PearsonCorrCoeff if input has a very small variance for its given dtype (#1926)Metric (#1963)CalibrationError where calculations for double precision input was performed in float precision (#1919)prefix/postfix arguments in MetricCollection and ClasswiseWrapper being duplicated (#1918)score argument (#1948)@borda, @SkafteNicki^n
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixes corner case when using MetricCollection together with aggregation metrics
MetricCollection together with aggregation metrics (#1896)max_fpr in AUROC metric when only one class is present (#1895)IntersectionOverUnion metric (#1892)MeanMetric and broadcasting of weights when Nans are present (#1898)MeanAveragePrecision (#1913)@fansuregrin, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated domain metrics import from package root (#1685, #1694, #1696, #1699, #1703)
We are happy to announce that the first major release of Torchmetrics, version v1.0, is publicly available. We have
worked hard on a couple of new features for this milestone release, but for v1.0.0, we have also managed to implement
over 100 metrics in torchmetrics.
The big new feature of v1.0 is a built-in plotting feature. As the old saying goes: "A picture is worth a thousand words". Within machine learning, this is definitely also true for many things.
Metrics are one area that, in some cases, is definitely better showcased in a figure than as a list of floats. The only requirement for getting started with the plotting feature is installing matplotlib. Either install with pip install matplotlib or pip install torchmetrics[visual] (the latter option also installs Scienceplots and uses that as the default plotting style).
The basic interface is the same for any metric. Just call the new .plot method:
metric = AnyMetricYouLike()
for _ in range(num_updates):
metric.update(preds[i], target[i])
fig, ax = metric.plot()
The plot method by default does not require any arguments and will automatically call metric.compute internally on
whatever metric states have been accumulated.
prefix and postfix arguments to ClasswiseWrapper (#1866)compute_with_cache to control caching behaviour after compute method (#1754)ComplexScaleInvariantSignalNoiseRatio for audio package (#1785)Running wrapper for calculate running statistics (#1752)RelativeAverageSpectralError and RootMeanSquaredErrorUsingSlidingWindow to image package (#816)SpecificityAtSensitivity Metric (#1432).plot() method (#1328, #1481, #1480, #1490, #1581, #1585, #1593, #1600, #1605, #1610, #1609, #1621, #1624, #1623, #1638, #1631, #1650, #1639, #1660, #1682, #1786).plot() method (#1434)classes to output from MAP metric (#1419)MinkowskiDistance to regression package (#1362)pairwise_minkowski_distance to pairwise package (#1362)PanopticQuality (#929, #1527)PSNRB metric (#1421)ClassificationTask Enum and use in metrics (#1479)ignore_index option to exact_match metric (#1540)top_k to RetrievalMAP (#1501)torch.cumsum operator (#1499).plot() method (#1485)data_range (#1606)ModifiedPanopticQuality metric to detection package (#1627)PrecisionAtFixedRecall metric to classification package (#1683)IntersectionOverUnionGeneralizedIntersectionOverUnionCompleteIntersectionOverUnionDistanceIntersectionOverUnionMultitaskWrapper to wrapper package (#1762)RelativeSquaredError metric to regression package (#1765)MemorizationInformedFrechetInceptionDistance metric to image package (#1580)permutation_invariant_training to allow using a 'permutation-wise' metric function (#1794)update_count and update_called from private to public methods (#1370)EnumStr raising ValueError for invalid value (#1479)PrecisionRecallCurve with large number of samples (#1493)__iter__ method from raising NotImplementedError to TypeError by setting to None (#1538)FID metric will now raise an error if too few samples are provided (#1655)torch.float64 (#1628)LPIPS implementation to no more rely on third-party package (#1575)scipy to torch (#1708)PearsonCorrCoeff to be more robust in certain cases (#1729)MeanAveragePrecision to pycocotools backend (#1832)MetricTracker for MultioutputWrapper and nested structures (#1608)PearsonCorrCoef (#1649)jsonargparse and LightningCLI (#1651)MultiOutputWrapper (#1675)MSSSIM (#1674)max_det_threshold in MAP detection (#1712)register_buffer (#1728)MeanAveragePrecision for iou_type="segm" (#1763)prefix and postfix in nested MetricCollection (#1773)ax plotting logging in `MetricCollection (#1783)RougeScore (#1789)CompositionalMetric (#1761)SpectralDistortionIndex metric (#1808)MatthewsCorrCoef (#1812, #1863)PearsonCorrCoef (#1819)average="macro" in classification metrics (#1821)ignore_index = num_classes + 1 in Multiclass-jaccard (#1860)@alexkrz, @AndresAlgaba, @basveeling, @Bomme, @Borda, @Callidior, @clueless-skywatcher, @Dibz15, @EPronovost, @fkroeber, @ItamarChinn, @marcocaccin, @martinmeinke, @niberger, @Piyush-97, @quancs, @relativityhd, @shenoynikhil, @shhs29, @SkafteNicki, @soma2000-lang, @srishti-git1110, @stancld, @twsl, @ValerianRey, @venomouscyanide, @wbeardall
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Fixed evaluation of R2Score with the near constant target
R2Score with the near constant target (#1576)dtype conversion when the metric is submodule (#1583)top_k>1 and ignore_index!=None in StatScores based metrics (#1589)PearsonCorrCoef when running in DDP mode but only on a single device (#1587)MAP when big areas are calculated (#1607)@borda, @FarzanT, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/metrics/compare/v0.11.3...v0.11.4
Fixed classification metrics for byte input
byte input (#1521)ignore_index in MulticlassJaccardIndex (#1386)@SkafteNicki, @vincentvaroquauxads
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/metrics/compare/v0.11.2...v0.11.3
Fixed compatibility between XLA in _bincount function
_bincount function (#1471)MetricTracker wrapper (#1472)multilabel in ExactMatch (#1474)@7shoe, @borda, @SkafteNicki, @ValerianRey
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/metrics/compare/v0.11.1...v0.11.2
Fixed type checking on the maximize parameter at the initialization of MetricTracker
maximize parameter at the initialization of MetricTracker (#1428)SSIM metric (#1454)nltk.punkt in RougeScore if a machine is not online (#1456)MultioutputWrapper (#1460)dtype checking in PrecisionRecallCurve for target tensor (#1457)@borda, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Full Changelog: https://github.com/Lightning-AI/metrics/compare/v0.11.0...v0.11.1
Removed deprecated BinnedAveragePrecision, BinnedPrecisionRecallCurve, RecallAtFixedPrecision
We are happy to announce that Torchmetrics v0.11 is now publicly available. In Torchmetrics v0.11 we have primarily focused on the cleanup of the large classification refactor from v0.10 and adding new metrics. With v0.11 are crossing 90+ metrics in Torchmetrics nearing the milestone of having 100+ metrics.
In Torchmetrics we are not only looking to expand with new metrics in already established metric domains such as classification or regression, but also new domains. We are therefore happy to report that v0.11 includes two new domains: Multimodal and nominal.
If there is one topic within machine learning that is hot right now then it is generative models and in particular image-to-text generative models. Just recently stable diffusion v2 was released, able to create even more photorealistic images from a single text prompt than ever
In Torchmetrics v0.11 we are adding a new domain called multimodal to support the evaluation of such models. For now, we are starting out with a single metric, the CLIPScore from this paper that can be used to evaluate such image-to-text models. CLIPScore currently achieves the highest correlation with human judgment, and thus a high CLIPScore for an image-text pair means that it is highly plausible that an image caption and an image are related to each other.
If you have ever taken any course in statistics or introduction to machine learning you should hopefully have heard about data can be of different types of attributes: nominal, ordinal, interval, and ratio. This essentially refers to how data can be compared. For example, nominal data cannot be ordered and cannot be measured. An example, would it be data that describes the color of your car: blue, red, or green? It does not make sense to compare the different values. Ordinal data can be compared but does have not a relative meaning. An example, would it be the safety rating of a car: 1,2,3? We can say that 3 is better than 1 but the actual numerical value does not mean anything.
In v0.11 of TorchMetrics, we are adding support for classic metrics on nominal data. In fact, 4 new metrics have already been added to this domain:
CramersVPearsonsContingencyCoefficientTschuprowsTTheilsUAll metrics are measures of association between two nominal variables, giving a value between 0 and 1, with 1 meaning that there is a perfect association between the variables.
In addition to metrics within the two new domains v0.11 of Torchmetrics contains other smaller changes and fixes:
TotalVariation metric has been added to the image package, which measures the complexity of an image with respect to its spatial variation.
MulticlassExactMatch metric has been added to the classification package, which for example can be used to measure sentence level accuracy where all tokens need to match for a sentence to be counted as correct
KendallRankCorrCoef have been added to the regression package for measuring the overall correlation between two variables
LogCoshError have been added to the regression package for measuring the residual error between two variables. It is similar to the mean squared error close to 0 but similar to the mean absolute error away from 0.
Finally, Torchmetrics now only supports v1.8 and higher of Pytorch. It was necessary to increase from v1.3 to secure because we were running into compatibility issues with an older version of Pytorch. We strive to support as many versions of Pytorch, but for the best experience, we always recommend keeping Pytorch and Torchmetrics up to date.
MulticlassExactMatch to classification metrics (#1343)TotalVariation to image package (#978)CLIPScore to new multimodal package (#1314)KendallRankCorrCoef (#1271)LogCoshError (#1316)CramersV (#1298)PearsonsContingencyCoefficient (#1334)TschuprowsT (#1334)TheilsU (#1337)distributed_available_fn to metrics to allow checks for custom communication backend for making dist_sync_fn actually useful (#1301)normalize argument to Inception, FID, KID metrics (#1246)BinnedAveragePrecision, BinnedPrecisionRecallCurve, RecallAtFixedPrecision (#1251)LabelRankingAveragePrecision, LabelRankingLoss and CoverageError (#1251)KLDivergence and AUC (#1251)pairwise_euclidean_distance (#1352)@borda, @justusschock, @ragavvenkatesan, @shenoynikhil, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed bug in Metrictracker.best_metric when return_step=False
Metrictracker.best_metric when return_step=False (#1306)Metrictracker.best_metric when return_step=False (#1306)@SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Changed in-place operation to out-of-place operation in pairwise_cosine_similarity
pairwise_cosine_similarity (#1288)average='micro' (#1286)structural_similarity_index_measure was used with autocast (#1291)spearman_corrcoef when used with autocast (#1303)@SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed broken clone method for classification metrics
nltk.punkt when lsum not in rouge_keys (#1258)MAP metric between bool and float32 (#1150)@dreaquil, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
This also means that the Binned* metrics that currently exist in TorchMetrics are being deprecated as their functionality is now captured by this argu…
TorchMetrics v0.10 is now out, significantly changing the whole classification package. This blog post will go over the reasons why the classification package needs to be refactored, what it means for our end users, and finally, what benefits it gives. A guide on how to upgrade your code to the recent changes can be found near the bottom.
We have for a long time known that there were some underlying problems with how we initially structured the classification package. Essentially, classification tasks can e divided into either binary, multiclass, or multilabel, and determining what task a user is trying to run a given metric on is hard just based on the input. The reason a package such as sklearn can do this is to only support input in very specific formats (no multi-dimensional arrays and no support for both integer and probability/logit formats).
This meant that some metrics, especially for binary tasks, could have been calculating something different than expected if the user were to provide another shape but the expected. This is against the core value of TorchMetrics, that our users, of course should trust that the metric they are evaluating is given the excepted result.
Additionally, classification metrics were missing consistency. For some, metrics num_classes=2 meant binary, and for others num_classes=1 meant binary. You can read more about the underlying reasons for this refactor in this and this issue.
The solution we went with was to split every classification metric into three separate metrics with the prefix binary_* , multiclass_* and multilabel_* . This solves a number of the above problems out of the box because it becomes easier for us to match our users' expectations for any given input shape. It additionally has some other benefits both for us as developers and ends users
The input arguments for the classification package are now much more standardized. Here are a few examples:
binary_* metrics are now required for all multiclass_* metrics and renamed to num_labels for all multilabel_* metrics.ignore_index argument is now supported by ALL classification metrics and supports any value and not only values in the [0,num_classes] range (similar to torch loss functions). Below is shown an example:validate_args to all classification metrics to allow users to skip validation of inputs making the computations completely faster. By default, we will still do input validation because it is the safest option for the user. Still, if you are confident that the input to the metric is correct, then you can now disable this, checking for a potential speed-up (more on this later).Some of the most useful metrics for evaluating classification problems are metrics such as ROC, AUROC, AveragePrecision, etc., because they not only evaluate your model for a single threshold but a whole range of thresholds, essentially giving you the ability to see the trade-off between Type I and Type II errors. However, a big problem with the standard formulation of these metrics (which we have been using) is that they require access to all data for their calculation. Our implementation has been extremely memory-intensive for these kinds of metrics.
In v0.10 of TorchMetrics, all these metrics now have an argument called thresholds. By default, it is None and the metric will still save all targets and predictions in memory as you are used to. However, if this argument is instead set to a tensor - torch.linspace(0,1,100) it will instead use a constant-memory approximation by evaluating the metric under those provided thresholds.
Setting thresholds=None has an approximate memory footprint of O(num_samples) whereas using thresholds=torch.linspace(0,1,100) has an approximate memory footprint of O(num_thresholds). In this particular case, users will save memory when the metric is computed on more than 100 samples. This feature can save memory by comparing this to modern machine learning, where evaluation is often done on thousands to millions of data points.
This also means that the Binned* metrics that currently exist in TorchMetrics are being deprecated as their functionality is now captured by this argument.
By splitting each metric into 3 separate metrics, we reduce the number of calculations needed. We, therefore, expected out-of-the-box that our new implementations would be faster. The table below shows the timings of different metrics with the old and new implementations (with and without input validation). Numbers in parentheses denote speed-up over old implementations.
The following observations can be made:
multiclass_confusion_matrix goes from a speedup of 3.36x to 4.81 when input validation is disabled. A clear advantage for users that are familiar with the metrics and do not need validation of their input at every update.InfoLM (#915)Perplexity metric (#922)ConcordanceCorrCoef metric to regression package (#1201)normalize to LPIPS metric (#1216)PESQ metric (#1227)PearsonCorrCoef and SpearmanCorrCoef (#1200)FID metric to be done in an online fashion to save memory (#1199)SSIM and MSSSIM update to be online to reduce memory usage (#1231)ssim when return_full_image=True where the score was still reduced (#1204)ClasswiseWrapper such that compute gave wrong result (#1225)@Borda, @bryant1410, @geoffrey-g-delhomme, @justusschock, @lucadiliello, @nicolas-dufour, @Queuecumber, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
InfoLM (#915)Perplexity metric (#922)ConcordanceCorrCoef metric to regression package (#1201)normalize to LPIPS metric (#1216)PESQ metric (#1227)PearsonCorrCoef and SpearmanCorrCoef (#1200)FID metric to be done in online fashion to save memory (#1199)SSIM and MSSSIM update to be online to reduce memory usage (#1231)BinnedAveragePrecision, BinnedPrecisionRecallCurve, BinnedRecallAtFixedPrecision (#1163)
BinnedAveragePrecision -> use AveragePrecision with thresholds argBinnedPrecisionRecallCurve -> use AveragePrecisionRecallCurve with thresholds argBinnedRecallAtFixedPrecision -> use RecallAtFixedPrecision with thresholds argLabelRankingAveragePrecision, LabelRankingLoss and CoverageError (#1167)
LabelRankingAveragePrecision -> MultilabelRankingAveragePrecisionLabelRankingLoss -> MultilabelRankingLossCoverageError -> MultilabelCoverageErrorKLDivergence and AUC from classification package (#1189)
KLDivergence moved to regression packageAUC use torchmetrics.utils.compute.aucssim when return_full_image=True where the score was still reduced (#1204)ClasswiseWrapper such that compute gave wrong result (#1225)Nothing published for this version
Added global option sync_on_compute to disable automatic synchronization when compute is called
sync_on_compute to disable automatic synchronization when compute is called (#1107)ClasswiseWrapper (#1129)JaccardIndex multi-label compute (#1125)gaussian_kernel is False, add test (#1149)@KeVoyer1, @krshrimali, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed mAP calculation for areas with 0 predictions
MeanAveragePrecision (#1097)average="none" in AvaragePrecision metric (#1116)@23pointsNorth, @kouyk, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Added specific RuntimeError when metric object is on the wrong device
RuntimeError when metric object is on the wrong device (#1056)BLEUScore and SacreBLEUScore instead of using uniform weights only. (#1075)TypeError when providing superclass arguments as kwargs (#1069)@jlcsilva, @SkafteNicki, @stancld
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Removed deprecated compute_on_step argument (#962, #967, #979 ,#990, #991, #993, #1005, #1004, #1007)
TorchMetrics v0.9 is now out, and it brings significant changes to how the forward method works. This blog post goes over these improvements and how they affect both users of TorchMetrics and users that implement custom metrics. TorchMetrics v0.9 also includes several new metrics and bug fixes.
Blog: TorchMetrics v0.9 — Faster forward
Since the beginning of TorchMetrics, Forward has served the dual purpose of calculating the metric on the current batch and accumulating in a global state. Internally, this was achieved by calling update twice: one for each purpose, which meant repeating the same computation. However, for many metrics, calling update twice is unnecessary to achieve both the local batch statistics and accumulating globally because the global statistics are simple reductions of the local batch states.
In v0.9, we have finally implemented a logic that can take advantage of this and will only call update once before making a simple reduction. As you can see in the figure below, this can lead to a single call of forward being 2x faster in v0.9 compared to v0.8 of the same metric.
With the improvements to forward, many metrics have become significantly faster (up to 2x)
It should be noted that this change mainly benefits metrics (for example, confusionmatrix) where calling update is expensive.
We went through all existing metrics in TorchMetrics and enabled this feature for all appropriate metrics, which was almost 95% of all metrics. We want to stress that if you are using metrics from TorchMetrics, nothing has changed to the API, and no code changes are necessary.
RetrievalPrecisionRecallCurve and RetrievalRecallAtFixedPrecision to retrieval package (#951)full_state_update that determines forward should call update once or twice (#984,#1033)Dice to classification package (#1021)segm as IOU for mean average precision (#822)reduction argument to average in Jaccard score and added additional options (#874)compute_on_step argument (#962, #967, #979 ,#990, #991, #993, #1005, #1004, #1007)dict for a few metrics (#1012)torch.double support in stat score metrics (#1023)FID calculation for non-equal size real and fake input (#1028)KLDivergence could output Nan (#1030)mdmc_average in Accuracy (#1036)MetricCollection (#1052)@Borda, @burglarhobbit, @charlielito, @gianscarpe, @MrShevan, @phaseolud, @razmikmelikbekyan, @SkafteNicki, @tanmoyio, @vumichien
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Fixed multi-device aggregation in PearsonCorrCoef
PearsonCorrCoef (#998)MetricCollection and prefix/postfix arg (#1007)safe_matmul (#1011, #1014)@ben-davidson-6, @Borda, @SkafteNicki, @tanmoyio
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Reimplemented the signal_distortion_ratio metric, which removed the absolute requirement of fast-bss-eval
signal_distortion_ratio metric, which removed the absolute requirement of fast-bss-eval (#964)BinnedPrecisionRecallCurve when thresholds argument is not provided (#968)CalibrationError to work on logit input (#985)@DuYicong515, @krshrimali, @quancs, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Deprecated argument compute_on_step
We are excited to announce that TorchMetrics v0.8 is now available. The release includes several new metrics in the classification and image domains and some performance improvements for those working with metrics collections.
Common wisdom dictates that you should never evaluate the performance of your models using only a single metric but instead a collection of metrics. For example, it is common to simultaneously evaluate the accuracy, precision, recall, and f1 score in classification. In TorchMetrics, we have for a long time provided the MetricCollection object for chaining such metrics together for an easy interface to calculate them all at once. However, in many cases, such a collection of metrics shares some of the underlying computations that have been repeated for every metric in the collection. In Torchmetrics v0.8 we have introduced the concept of compute_groups to MetricCollection that will, as default, be auto-detected and group metrics that share some of the same computations.
Thus, if you are using MetricCollections in your code, upgrading to TorchMetrics v0.8 should automatically make your code run faster without any code changes.
TorchMetrics v0.8 includes several new metrics within the classification and image domain, both for the functional and modular API. We refer to the documentation for the full description of all metrics if you want to learn more about them.
SpectralAngleMapper or SAM was added to the image package. This metric can calculate the spectral similarity between given reference spectra and estimated spectra.CoverageError was added to the classification package. This metric can be used when you are working with multi-label data. The metric works similar to the sklearn counterpart and computes how far you need to go through ranked scores such that all true labels are covered.LabelRankingAveragePrecision and LabelRankingLoss were added to the classification package. Both metrics are used in multi-label ranking problems, where the goal is to give a better rank to the labels associated with each sample. Each metric gives a measure of how well your model is doing this.ErrorRelativeGlobalDimensionlessSynthesis or ERGAS was added to the image package. This metric can be used to calculate the accuracy of Pan sharpened images considering the normalized average error of each band of the resulting image.UniversalImageQualityIndex was added to the image package. This metric can assess the difference between two images, which considers three different factors when computed: loss of correlation, luminance distortion, and contrast distortion.ClasswiseWrapper was added to the wrapper package. This wrapper can be used in combinations with metrics that return multiple values (such as classification metrics with the average=None argument). The wrapper will unwrap the result into a dict with a label for each value.WeightedMeanAbsolutePercentageError to regression package (#948)CoverageError (#787)LabelRankingAveragePrecision and LabelRankingLoss (#787)SpectralAngleMapper (#885)ErrorRelativeGlobalDimensionlessSynthesis (#894)UniversalImageQualityIndex (#824)SpectralDistortionIndex (#873)MetricCollection in MetricTracker (#718)StructuralSimilarityIndexMeasure (#818)MetricCollection (#709)ClasswiseWrapper for better logging of classification metrics with multiple output values (#832)**kwargs argument for passing additional arguments to base class (#833)ignore_index for the Accuracy metric (#362)adaptive_k for the RetrievalPrecision metric (#910)reset_real_features argument image quality assessment metrics (#722)compute_on_cpu to all metrics (#867)num_classes in jaccard_index a required argument (#853, #914)permutation_invariant_training (#864)None (#891)MetricTracker.best_metric will now give a warning when computing on metric that do not have a best (#913)compute_on_step (#792)dist_sync_on_step, process_group, dist_sync_fn direct argument (#833)WER and functional.werSSIM and functional.ssimPSNR and functional.psnrFBeta and functional.fbetaF1 and functional.f1Hinge and functional.hingeIoU and functional.iouMatthewsCorrcoefPearsonCorrcoefSpearmanCorrcoefMAP and functional.pairwise.manhattenPESQ and functional.audio.pesqPIT and functional.audio.pitSDR and functional.audio.sdr and functional.audio.si_sdrSNR and functional.audio.snr and functional.audio.si_snrSTOI and functional.audio.stoiMAP metric in specific cases (#950)ClasswiseWrapper with the prefix argument of MetricCollection (#843)BestScore on GPU (#912)ROUGEScore (#944)@ankitaS11, @ashutoshml, @Borda, @hookSSi, @justusschock, @lucadiliello, @quancs, @rusty1s, @SkafteNicki, @stancld, @vumichien, @weningerleon, @yassersouri
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Fixed unsafe log operation in TweedieDeviace for power=1
TweedieDeviace for power=1 (#847)ConfusionMatrix, AUROC and AveragePrecision on GPU when running in deterministic mode (#900)signal_distortion_ratio (#899)update method with tensor where requires_grad=True (#902)@mtailanian, @quancs, @SkafteNicki
If we forgot someone due to not matching commit email with GitHub account, let us know :]
## [0.7.2] - 2022-02-10 ### Fixed - Minor patches in JOSS paper.
Used torch.bucketize in calibration error when torch>1.8 for faster computations
torch.bucketize in calibration error when torch>1.8 for faster computations (#769)@Borda, @ramonemiliani93, @SkafteNicki, @twsl
If we forgot someone due to not matching commit email with GitHub account, let us know :]
If different naming was used before v0.7, it is deprecated and completely removed in v0.8.
We are excited to announce that TorchMetrics v0.7 is now publicly available. This release is pretty significant. It includes several new metrics (mainly for NLP), naming and import changes, general improvements to the API, and some other great features. TorchMetrics thus now has over 60+ metrics, and the package is more user-friendly than ever.
Text package is a part of TorchMetrics as of v0.5. With the growing capability of language generation models, there is also a real need to have reliable evaluation metrics. With several added metrics and unified API, TorchMetrics makes the usage of various metrics even easier! TorchMetrics v0.7 newly includes a couple of machine translation metrics such as chrF, chrF++, Translation Edit Rate, or Extended Edit Distance. Furthermore, it also supports other metrics - Match Error Rate, Word Information Lost, Word Information Preserved, and SQuAD evaluation metrics. Last but not least, we also made possible the evaluation of the ROUGE score using multiple references.
Importantly, all text metrics assume preds, target input order with these explicit keyword arguments. If different naming was used before v0.7, it is deprecated and completely removed in v0.8.
TorchMetrics v0.7 brings more extensive and minor changes to how metrics should be imported. The import changes directly impact v0.7, meaning that you will most likely need to change the import statement for some specific metrics. All naming changes follow our standard deprecation process, meaning that in v0.7, any metric that is renamed will still work but raise an error asking to use the new metric name. From v0.8, the old metric names will no longer be available.
MatchErrorRate (#619)WordInfoLost and WordInfoPreserved (#630)SQuAD (#623)CHRFScore (#641)TranslationEditRate (#646)ExtendedEditDistance (#668)MultiScaleSSIM into image metrics (#679)SDR) to audio package (#565)MinMaxMetric to wrappers (#556)ignore_index to retrieval metrics (#676)ROUGEScore (#680)BLEUScore input stay consistent with all the other text metrics (#640)TER, BLEUScore, SacreBLEUScore, CHRFScore now the expected input order is predictions first and target second (#696)torch.float to torch.long in ConfusionMatrix to accommodate larger values (#715)preds, target input argument's naming across all text metrics (#723, #727)
bert, bleu, chrf, sacre_bleu, wip, wil, cer, ter, wer, mer, rouge, squadfunctional.wer -> functional.word_error_rateWER -> WordErrorRateMatthewsCorrcoef -> MatthewsCorrCoefPearsonCorrcoef -> PearsonCorrCoefSpearmanCorrcoef -> SpearmanCorrCoefaudio.STOI to audio.ShortTimeObjectiveIntelligibilityfunctional.audio.stoi to functional.audio.short_time_objective_intelligibilityfunctional.audio.pesq -> functional.audio.perceptual_evaluation_speech_qualityaudio.PESQ -> audio.PerceptualEvaluationSpeechQualityfunctional.sdr -> functional.signal_distortion_ratiofunctional.si_sdr -> functional.scale_invariant_signal_distortion_ratioSDR -> SignalDistortionRatioSI_SDR -> ScaleInvariantSignalDistortionRatiofunctional.snr -> functional.signal_distortion_ratiofunctional.si_snr -> functional.scale_invariant_signal_noise_ratioSNR -> SignalNoiseRatioSI_SNR -> ScaleInvariantSignalNoiseRatiofunctional.f1 -> functional.f1_scoreF1 -> F1Scorefunctional.fbeta -> functional.fbeta_scoreFBeta -> FBetaScorefunctional.hinge -> functional.hinge_lossHinge -> HingeLossfunctional.psnr -> functional.peak_signal_noise_ratioPSNR -> PeakSignalNoiseRatiofunctional.pit -> functional.permutation_invariant_trainingPIT -> PermutationInvariantTrainingfunctional.ssim -> functional.scale_invariant_signal_noise_ratioSSIM -> StructuralSimilarityIndexMeasureMAP to MeanAveragePrecision metric (#754)image.FID -> image.FrechetInceptionDistanceimage.KID -> image.KernelInceptionDistanceimage.LPIPS -> image.LearnedPerceptualImagePatchSimilarityembedding_similarity metric (#638)concatenate_texts from wer metric (#638)newline_sep and decimal_places from rouge metric (#638)kwargs are present in update signature (#707)@ashutoshml, @Borda, @cuent, @Fariborzzz, @getgaurav2, @janhenriklambrechts, @justusschock, @karthikrangasai, @lucadiliello, @mahinlma, @mathemusician, @mona0809, @mrleu, @puhuk, @quancs, @SkafteNicki, @stancld, @twsl
If we forgot someone due to not matching commit email with GitHub account, let us know :]
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →