NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3942 most downloaded on PyPI
State-of-the-art speaker diarization toolkit
Last release 3 months ago
30 Jun 2026
Release timing varies
gaps range from 8 days to 12 months
Nearly every release is documented
notes for 20 of 22 stable releases
1 version withdrawn
withdrawn after publishing
6 years old
23 releases · first in 2020
One column per quarter.
fix(pipeline): add missing support for subfolder in Pipeline.from_pretrained
subfolder in Pipeline.from_pretrainedNothing published for this version
improve(telemetry): reduce number of sent packages @litdarya
--average-case option to optimize command @antoinelaurentTask.prepare_data to support saving preprocessors that produce int values in metadata @lylyhanfeat(sample): add transcription of sample file
feat(cli): add --revision option to most CLI commands
--revision option to most CLI commandsCalibration.safe_transform method (supports NaNs as well as any shape)Model.from_pretrained to support lightning 2.6+pyannote-database dependency to 6.1+BREAKING(util): make Binarize.__call__ return string tracks (instead of int ) @benniekiss
Binarize.__call__ return string tracks (instead of int) @benniekisstorch, torchcodec, and torchaudio versions to avoid segmentation faultPipeline.cuda() convenience method @tkanarskypreload option to base Pipeline.__call__ to force preloading audio in memory (@antoinelaurent)permutate faster thanks to vectorized cost functionFull Changelog: 4.0.1...4.0.2
feat: allow passing preloaded pipeline config to get_pipeline
get_pipelinepyannoteai-sdk dependency to 0.3.0OpenTelemetry dependenciesBREAKING(cli): remove deprecated pyannote-audio-train CLI
pyannote/speaker-diarization-community-1 pretrained pipeline relies on VBx clustering instead of agglomerative hierarchical clustering (as suggested by BUT Speech@FIT researchers Petr Pálka and Jiangyu Han).
pyannote/speaker-diarization-community-1 pretrained pipeline returns a new exclusive speaker diarization, on top of the regular speaker diarization.
This is a feature which is backported from our latest commercial model that simplifies the reconciliation between fine-grained speaker diarization timestamps and (sometimes not so precise) transcription timestamps.
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1", token="huggingface-access-token")
output = pipeline("/path/to/conversation.wav")
print(output.speaker_diarization) # regular speaker diarization
print(output.exclusive_speaker_diarization) # exclusive speaker diarization
Metadata caching and optimized dataloaders make training on large scale datasets much faster.
This led to a 15x speed up on pyannoteAI internal large scale training.
Change one line of code to use pyannoteAI premium models and enjoy more accurate speaker diarization.
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
- "pyannote/speaker-diarization-community-1", token="huggingface-access-token")
+ "pyannote/speaker-diarization-precision-2, token="pyannoteAI-api-key")
diarization = pipeline("/path/to/conversation.wav")
Pipelines can now be stored alongside their internal models in the same repository, streamlining fully offline use.
Accept pyannote/speaker-diarization-community-1 pipeline user agreement
Clone the pipeline repository from Huggingface (if prompted for a password, use a Huggingface access token with correct permissions)
$ git lfs install
$ git clone https://hf.co/pyannote/speaker-diarization-community-1 /path/to/directory/pyannote-speaker-diarization-community-1
Enjoy!
# load pipeline from disk (works without internet connection)
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained('/path/to/directory/pyannote-speaker-diarization-community-1')
# run the pipeline locally on your computer
diarization = pipeline("audio.wav")
With the optional telemetry feature in pyannote.audio, you can choose to send anonymous usage metrics to help the pyannote team improve the library.
sox and soundfile audio I/O backends (only ffmpeg or in-memory audio is supported)Python < 3.10use_auth_token to token{pipeline_name}@{revision} syntax in Model.from_pretrained(...) and Pipeline.from_pretrained(...) -- use new revision keyword argument insteadOverlappedSpeechDetection task (part of SpeakerDiarization task)OverlappedSpeechDetection and Resegmentation unmaintained pipelines (part of SpeakerDiarization)huggingface_hub caching directory (PYANNOTE_CACHE is no longer used)Inference now only supports already instantiated modelsmultilabel training in SpeakerDiarization taskwarm_up option in SpeakerDiarization taskweigh_by_cardinality option in SpeakerDiarization taskvad_loss option in SpeakerDiarization taskpyannote-audio-train CLItorchaudio to torchcodec for audio I/Ok-means clusteringwav2vec_frozen option to freeze/unfreeze wav2vec in SSeRiouSS architectureSpeakerDiarization taskhidden option to ProgressHookFilterByNumberOfSpeakers protocol files filterCalibration class to calibrate logits/distances into probabilitiesDetectionErrorRate, SegmentationErrorRate, DiarizationPrecision, and DiarizationRecall metricsSSeRiouSS architecture (@clement-pages)SpeakerDiarization training with manual optimization (@clement-pages)uvlightning from pytorch-lightningio.BytesIO bufferdict (@benniekiss)ToTaToNet architecture (@clement-pages)PixIT training with manual optimization (@clement-pages)SpeechSeparation and SpeakerDiarization docstring (@razi-tm).Upcoming major releases of pyannote.{core,database,metrics,pipeline} dependencies will break 3.x branch. Version 3.4.0 pins those dependencies to comp
Upcoming major releases of pyannote.{core,database,metrics,pipeline} dependencies will break 3.x branch.
Version 3.4.0 pins those dependencies to compatible versions.
fix: (really) fix support for numpy==2.x (@metal3d)
setup: drop support for Python 3.8
numpy==2.x (@ibevers)speechbrain==1.x (@Adel-Moumen)pyannote.audio does speech separation: multi-speaker audio in, one audio channel per speaker out!
pyannote.audio does speech separation: multi-speaker audio in, one audio channel per speaker out!
pip install pyannote.audio[separation]==3.3.0
PixIT joint speaker diarization and speech separation task (with @joonaskalda)ToTaToNet joint speaker diarization and speech separation model (with @joonaskalda)SpeechSeparation pipeline (with @joonaskalda)backendsoundfile backendmax_speakers is set to 1feat(task): add option to cache task training metadata to speed up training (with @clement-pages)
receptive_field, num_frames and dimension to models (with @Bilal-Rahou)fbank_only property to WeSpeaker modelsPowerset.permutation_mapping to help with permutation in powerset space (with @FrenchKrab)pyannote.audio.sample.SAMPLE_FILEreduce option to diarization_error_rate metric (with @Bilal-Rahou)Waveform and SampleRate preprocessorstorch.Tensor support in ArtifactHookPowerset docstring (with @lukasstorck)numpy.ndarray waveform (with @Purfview)diarization_error_rate metricModel and nn.Module attributes in Pipeline.to(device)torchaudio >= 2.2.0Model.example_output in favor of num_frames method, receptive_field property, and dimension propertypyannote/speaker-diarization-3.1 (by @simonottenhauskenbun)Providing num_speakers to `pyannote/speaker-diarization-3.1` now works as expected.
Providing num_speakers to pyannote/speaker-diarization-3.1 now works as expected.
num_speakers in pyannote/speaker-diarization-3.1 pipeline`pyannote/speaker-diarization-3.1` no longer requires unpopular ONNX runtime
pyannote/speaker-diarization-3.1 no longer requires unpopular ONNX runtime
TimingHook for profiling processing timeArtifactHook for saving internal stepsHooks"soft" option to Powerset.to_multilabelSpeakerDiarizationAgglomerativeClustering to honor num_clusters when providedmax_speakers or detected num_speakers in SpeakerDiarization pipelinefbank on GPU when requestedWeSpeakerPretrainedSpeakerEmbedding to ONNXWeSpeakerPretrainedSpeakerEmbeddingonnxruntime dependency.
You can still use ONNX hbredin/wespeaker-voxceleb-resnet34-LM but you will have to install onnxruntime yourself.logging_hook (use ArtifactHook instead)onset and offset parameter in SpeakerDiarizationMixin.speaker_count
You should now binarize segmentations before passing them to speaker_countfix(pipeline): fix WeSpeaker GPU support
feat(pipeline): send pipeline to device with pipeline.to(device)
pipeline.to(device)return_embeddings option to SpeakerDiarization pipelinesegmentation_batch_size and embedding_batch_size mutable in SpeakerDiarization pipeline (they now default to 1)SpeakerDiarization taskSegmentation task to SpeakerDiarizationpipeline.to(device))SpeakerSegmentation pipeline (use SpeakerDiarization pipeline)segmentation_duration parameter from SpeakerDiarization pipeline (defaults to duration of segmentation model)FINCHClustering and HiddenMarkovModelClusteringpyannote.audio.core.io.Audio is instantiated:
Audio() by Audio(mono="downmix");Audio(mono=True) by Audio(mono="downmix");Audio(mono=False) by Audio().Model.introspection
If, for some weird reason, you wrote some custom code based on that,
you should instead rely on Model.example_output.BREAKING(pipeline): rewrite speaker diarization pipeline
feat: pretrained pipelines (and models) on Huggingface model hub
fix: make sure master branch is used to load pretrained models
Nothing published for this version
last release before complete rewriting
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →