NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1618 most downloaded on PyPI
An audio package for PyTorch
Last release 6 months ago
23 Mar 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 42 of 43 stable releases
1 version withdrawn
withdrawn after publishing
7 years old
44 releases · first in 2019
This release is compatible with torch 2.11 and is compatible with future versions of torch. No new features were added.
This release is compatible with torch 2.11 and is compatible with future versions of torch. No new features were added.
The C++ and CUDA extensions that were previously marked for deprecation are preserved and will remain in torchaudio (lfilter, RNNTLoss, CUCT, forced_a…
This release is compatible with torch 2.10. No new features were added.
The C++ and CUDA extensions that were previously marked for deprecation are
preserved and will remain in torchaudio (lfilter, RNNTLoss, CUCT, forced_align, and overdrive.).
This release marks the finalization of the TorchAudio migration described in https://github.com/pytorch/audio/issues/3902.
One column per quarter.
This is a patch release, which is compatible with PyTorch 2.9.1. There are no new features added.
This is a patch release, which is compatible with PyTorch 2.9.1. There are no new features added.
Most APIs marked as drop are now removed.
Most APIs marked as drop are now removed.
load() and save() to TorchCodecWe are consolidating the decoding and encoding capabilities of PyTorch in TorchCodec. torchaudio.load() and torchaudio.save() still exist, but their underlying implementation now relies on TorchCodec.
We're still working towards preserving forced_align, lfilter, overdrive, RNNT, and CUCTC.
Most APIs marked as "Drop" are now explicitly deprecated, raising deprecation warnings in the docs, and when using them from Python. They will be remo…
Copy pasting update from https://github.com/pytorch/audio/issues/3902#issuecomment-3160818888:
Most APIs marked as "Drop" are now explicitly deprecated, raising deprecation warnings in the docs, and when using them from Python. They will be removed in the next 2.9 version.
load() and save() to TorchCodecAs we mentioned, we are consolidating the decoding and encoding capabilities of PyTorch in TorchCodec.
torchaudio.load() and torchaudio.save() are some of the most popular TorchAudio APIs, so for convenience we are providing torchaudio.load_with_torchcodec() and torchaudio.save_with_torchcodec(), which can largely be used as drop-in replacements. However, we do encourage users to directly migrate to TorchCodec's AudioDecoder() and AudioEncoder().
In future versions, torchaudio.load() and torchaudio.save() will still exist, but their underlying implementation will be relying on torchaudio.load_with_torchcodec() and torchaudio.save_with_torchcodec().
We hope for this migration to be as smooth as possible - most users should just need to pip install torchcodec, and things should still work as-is.
TorchCodec doesn't support Windows yet, but we're working hard on it. Please bear with us.
We mentioned that we were exploring options to retain the C++-backed APIs, which are currently slated for deletion. Specifically: forced_align, lfilter, overdrive, RNNT, and CUCTC.
While this isn't something I can assert with 100% certainty, we are now more confident that we'll be able to preserve these extensions by porting them to Pytorch's new "stable ABI" operators. We are actively working on it.
This release is compatible with PyTorch 2.7.1 There are no new features added.
This release is compatible with PyTorch 2.7.1 There are no new features added.
[!NOTE] We are in the process of refactoring TorchAudio and transitioning it into a maintenance phase. This process will include removing some user-facing features. Our main goals are to reduce redundancies with the rest of the PyTorch ecosystem, make it easier to maintain, and create a version of TorchAudio that is more tightly scoped to its strengths: processing audio data for ML. Please see our community message for more details.
This release is compatible with PyTorch 2.7. There are no new features added.
This release is compatible with PyTorch 2.7. There are no new features added.
[!NOTE] We are in the process of refactoring TorchAudio and transitioning it into a maintenance phase. This process will include removing some user-facing features. Our main goals are to reduce redundancies with the rest of the PyTorch ecosystem, make it easier to maintain, and create a version of TorchAudio that is more tightly scoped to its strengths: processing audio data for ML. Please see our community message for more details.
This release is compatible with PyTorch 2.6. There are no new features added.
This release is compatible with PyTorch 2.6. There are no new features added.
The following fixes / improvement were made:
Nothing published for this version
This release is compatible with PyTorch 2.5. There are no new features added.
This release is compatible with PyTorch 2.5. There are no new features added.
This release contains one improvement:
This release is compatible with PyTorch 2.4.1 patch release. There are no new features added.
This release is compatible with PyTorch 2.4.1 patch release. There are no new features added.
This release is compatible with PyTorch 2.4. There are no new features added.
This release is compatible with PyTorch 2.4. There are no new features added.
This release contains 2 fixes:
This release is compatible with PyTorch 2.3.1 patch release. There are no new features added.
This release is compatible with PyTorch 2.3.1 patch release. There are no new features added.
This release is compatible with PyTorch 2.3.0 patch release. There are no new features added.
This release is compatible with PyTorch 2.3.0 patch release. There are no new features added.
This release contains minor documentation and code quality improvements (#3734, #3748, #3757, #3759)
This release is compatible with PyTorch 2.2.2 patch release. There are no new features added.
This release is compatible with PyTorch 2.2.2 patch release. There are no new features added.
This release is compatible with PyTorch 2.2.1 patch release. There are no new features added.
This release is compatible with PyTorch 2.2.1 patch release. There are no new features added.
Add path-like object support to StreamReader/Writer https://github.com/pytorch/audio/pull/3608
trio top-level module, dedicated for core I/O operations (https://github.com/pytorch/audio/pull/3676, https://github.com/pytorch/audio/pull/3680, https://github.com/pytorch/audio/pull/3681, https://github.com/pytorch/audio/pull/3682) Please refer to https://pytorch.org/audio/2.2.0/torio.html for the details.This is a patch release, which is compatible with PyTorch 2.1.2. There are no new features added.
This is a patch release, which is compatible with PyTorch 2.1.2. There are no new features added.
This is a minor release, which is compatible with PyTorch 2.1.1 and includes bug fixes, improvements and documentation updates.
This is a minor release, which is compatible with PyTorch 2.1.1 and includes bug fixes, improvements and documentation updates.
torchaudio.functional.apply_codec (also deprecated, see below)
TorchAudio v2.1 introduces the new features and backward-incompatible changes;
torchaudio.io.AudioEffector can apply filters, effects and encodings to waveforms in online/offline fashion.torchaudio.functional.forced_align computes alignment from an emission and torchaudio.pipelines.MMS_FA provides access to the model trained for multilingual forced alignment in MMS: Scaling Speech Technology to 1000+ languages project.forced_align function, and https://pytorch.org/audio/2.1/tutorials/forced_alignment_for_multilingual_data_tutorial.html for how one can use MMS_FA to align transcript in multiple languages.torchaudio.pipelines.SQUIM_SUBJECTIVE and torchaudio.pipelines.SQUIM_OBJECTIVE models to estimate the various speech quality and intelligibility metrics. This is helpful when evaluating the quality of speech generation models, such as TTS.torchaudio.models.decoder.CUCTCDecoder takes emission stored in CUDA memory and performs CTC beam search on it in CUDA device. The beam search is fast. It eliminates the need to move data from CUDA device to CPU when performing automatic speech recognition. With PyTorch's CUDA support, it is now possible to perform the entire speech recognition pipeline in CUDA.torchaudio.io.StreamWriter (#3135)torchaudio.io.StreamReader.get_out_stream_info (#3155)torchaudio.io.StreamReader filter graph (#3183, #3479)torchaudio.io.StreamWriter (#3194)torchaudio.io.StreamReader (#3216)torchaudio.io.StreamWriter (#3207)420p10le support to torchaudio.io.StreamReader CPU decoder (#3332)torchaudio.io.AudioEffector (#3163, #3372, #3374)torchaudio.transforms.SpecAugment (#3309, #3314)torchaudio.functional.forced_align (#3348, #3355, #3533, #3536, #3354, #3365, #3433, #3357)torchaudio.functional.merge_tokens (#3535, #3614)torchaudio.functional.frechet_distance (#3545)torchaudio.models.SquimObjective for speech enhancement (#3042, 3087, #3512)torchaudio.models.SquimSubjective for speech enhancement (#3189)torchaudio.models.decoder.CUCTCDecoder (#3096)torchaudio.pipelines.SquimObjectiveBundle for speech enhancement (#3103)torchaudio.pipelines.SquimSubjectiveBundle for speech enhancement (#3197)torchaudio.pipelines.MMS_FA Bundle for forced alignment (#3521, #3538)torchaudio.io.AudioEffector (#3226)torchaudio.models.decoder.CUCTCDecoder (#3297)In this release, the following third party libraries are removed from TorchAudio binary distributions. TorchAudio now search and link these libraries at runtime. Please install them to use the corresponding APIs.
libsox is used for various audio I/O, filtering operations.
Pre-built binaries are avaialble via package managers, such as conda, apt and brew. Please refer to the respective documetation.
The APIs affected include;
torchaudio.load ("sox" backend)torchaudio.info ("sox" backend)torchaudio.save ("sox" backend)torchaudio.sox_effects.apply_effects_tensortorchaudio.sox_effects.apply_effects_filetorchaudio.functional.apply_codec (also deprecated, see below)Changes related to the removal: #3232, #3246, #3497, #3035
flashlight-text is the core of CTC decoder.
Pre-built packages are available on PyPI. Please refer to https://github.com/flashlight/text for the detail.
The APIs affected include;
torchaudio.models.decoder.CTCDecoderChanges related to the removal: #3232, #3246, #3236, #3339
A custom built libkaldi was used to implement torchaudio.functional.compute_kaldi_pitch. This function, along with libkaldi integration, is removed in this release. There is no replcement.
Changes related to the removal: #3368, #3403
To make I/O operations more flexible, TorchAudio introduced the backend dispatcher in v2.0, and users could opt-in to use the dispatcher. In this release, the backend dispatcher becomes the default mechanism for selecting the I/O backend.
You can pass backend argument to torchaudio.info, torchaudio.load and torchaudio.save function to select I/O backend library per-call basis. (If it is omitted, an available backend is automatically selected.)
If you want to use the global backend mechanism, you can set the environment variable, TORCHAUDIO_USE_BACKEND_DISPATCHER=0.
Please note, however, that this the global backend mechanism is deprecated and is going to be removed in the next release.
Please see #2950 for the detail of migration work.
torchaudio.io.StreamReader accepted a byte-string wrapped in 1D torch.Tensor object. This is no longer supported.
Please wrap the underlying data with io.BytesIO instead.
The optional arguments of add_[audio|video]_stream methods of torchaudio.io.StreamReader and torchaudio.io.StreamWriter are now keyword-only arguments.
Previously TorchAudio supported FFmpeg 4 (>=4.1, <=4.4). In this release, TorchAudio supports FFmpeg 4, 5 and 6 (>=4.4, <7). With this change, support for FFmpeg 4.1, 4.2 and 4.3 are dropped.
torchaudio.functional.apply_codec (#3397)In previous versions, TorchAudio shipped custom built libsox, so that it can perform in-memory decoding and encoding.
Now, in-memory decoding and encoding are handled by FFmpeg binding, and with the switch to dynamic libsox linking, torchaudio.functional.apply_codec no longer process audio in in-memory fashion. Instead it writes to temporary file.
For in-memory processing, please use torchaudio.io.AudioEffector.
lstsq when solving InverseMelScale (#3280)Previously, torchaudio.transform.InverseMelScale ran SGD optimizer to find the inverse of mel-scale transform. This approach has number of issues as listed in #2643.
This release switches to use torch.linalg.lstsq.
The infer method of torchaudio.models.RNNTBeamSearch has been updated to accept series of previous hypotheses.
bundle = torchaudio.pipelines.EMFORMER_RNNT_BASE_LIBRISPEECH
decoder: RNNTBeamSearch = bundle.get_decoder()
hypothesis = None
while streaming:
...
hypo, state = decoder.infer(
features,
length,
beam_width,
state=state,
hypothesis=hypothesis,
)
...
hypothesis = hypo
# Previously this had to be hypothesis = hypo[0]
torchaudio.functional.apply_codec function (#3386)Due to the removal of custom libsox binding, torchaudio.functional.apply_codec no longer supports in-memory processing. Please migrate to torchaudio.io.AudioEffector.
Please refer to for the detailed usage of torchaudio.io.AudioEffector.
get_trellis in forced alignment tutorial (#3172)torchaudio.io.StreamWriter (#3373)lfilter (#3432)torchaudio.io.StreamWriter is not opened (#3152)torchaudio.io.StreamReader (#3157, #3170, #3186, #3184, #3188, #3320, #3296, #3328, #3419, #3209)torchaudio.io.StreamWriter (#3205, #3319, #3296, #3328, #3426, #3428)n_fft (#3442)torch.norm to torch.linalg.vector_norm (#3522)torch.nn.utils.weight_norm to nn.utils.parametrizations.weight_norm (#3523)This is a minor release, which is compatible with PyTorch 2.0.1 and includes bug fixes, improvements and documentation updates. There is no new featur
This is a minor release, which is compatible with PyTorch 2.0.1 and includes bug fixes, improvements and documentation updates. There is no new feature added.
Full Changelog: https://github.com/pytorch/audio/compare/v2.0.1...v2.0.2
Removed deprecated/unused/undocumented functions from datasets.utils (#2926, #2927) The following functions are removed from datasets.utils
TorchAudio 2.0 release includes:
info, load, save functionsThe release adds several data augmentation operators under torchaudio.functional and torchaudio.transforms:
torchaudio.functional.add_noisetorchaudio.functional.convolvetorchaudio.functional.deemphasistorchaudio.functional.fftconvolvetorchaudio.functional.preemphasistorchaudio.functional.speedtorchaudio.transforms.AddNoisetorchaudio.transforms.Convolvetorchaudio.transforms.Deemphasistorchaudio.transforms.FFTConvolvetorchaudio.transforms.Preemphasistorchaudio.transforms.Speedtorchaudio.transforms.SpeedPerturbationThe operators can be used to synthetically diversify training data to improve the generalizability of downstream models.
For usage details, please refer to the documentation for torchaudio.functional and torchaudio.transforms, and tutorial “Audio Data Augmentation”.
The release adds two self-supervised learning models for speech and audio.
Besides the model architectures, torchaudio also supports corresponding pre-trained pipelines:
torchaudio.pipelines.WAVLM_BASEtorchaudio.pipelines.WAVLM_BASE_PLUStorchaudio.pipelines.WAVLM_LARGEtorchaudio.pipelines.WAV2VEC_XLSR_300Mtorchaudio.pipelines.WAV2VEC_XLSR_1Btorchaudio.pipelines.WAV2VEC_XLSR_2BFor usage details, please refer to factory function and pre-trained pipelines documentation.
Release 2.0 introduces new versions of I/O functions torchaudio.info, torchaudio.load and torchaudio.save, backed by a dispatcher that allows for selecting one of backends FFmpeg, SoX, and SoundFile to use, subject to library availability. Users can enable the new logic in Release 2.0 by setting the environment variable TORCHAUDIO_USE_BACKEND_DISPATCHER=1; the new logic will be enabled by default in Release 2.1.
# Fetch metadata using FFmpeg
metadata = torchaudio.info("test.wav", backend="ffmpeg")
# Load audio (with no backend parameter value provided, function prioritizes using FFmpeg if it is available)
waveform, rate = torchaudio.load("test.wav")
# Write audio using SoX
torchaudio.save("out.wav", waveform, rate, backend="sox")
Please see the documentation for torchaudio for more details.
Dropped Python 3.7 support (#3020) Following the upstream PyTorch (https://github.com/pytorch/pytorch/pull/93155), the support for Python 3.7 has been dropped.
Default to "precise" seek in torchaudio.io.StreamReader.seek (#2737, #2841, #2915, #2916, #2970)
Previously, the StreamReader.seek method seeked into a key frame closest to the given time stamp. A new option mode has been added which can switch the behavior to seeking into any type of frame, including non-key frames, that is closest to the given timestamp, and this behavior is now default.
Removed deprecated/unused/undocumented functions from datasets.utils (#2926, #2927)
The following functions are removed from datasets.utils
stream_urldownload_urlvalidate_fileextract_archive.Deprecated 'onesided' init param for MelSpectrogram (#2797, #2799)
torchaudio.transforms.MelSpectrogram assumes the onesided argument to be always True. The forward path fails if its value is False. Therefore this argument is deprecated. Users specifying this argument should stop specifying it.
Deprecated "sinc_interpolation" and "kaiser_window" option value in favor of "sinc_interp_hann" and "sinc_interp_kaiser" (#2922)
The valid values of resampling_method argument of resampling operations (torchaudio.transforms.Resample and torchaudio.functional.resample) are changed. "kaiser_window" is now "sinc_interp_kaiser" and "sinc_interpolation" is "sinc_interp_hann". The old values will continue to work, but users are encouraged to update their code.
For the reason behind of this change, please refer #2891.
Deprecated sox initialization/shutdown public API functions (#3010)
torchaudio.sox_effects.init_sox_effects and torchaudio.sox_effects.shutdown_sox_effects are deprecated. They were required to use libsox-related features, but are called automatically since v0.6, and the initialization/shutdown mechanism have been moved elsewhere. These functions are now no-op. Users can simply remove the call to these functions.
torchaudio.load, torchaudio.info and torchaudio.save.torchaudio.sox_effects.apply_effects_file and torchaudio.functional.apply_codec.
For I/O, to continue using file-like objects, please use the new dispatcher mechanism.
For effects, replacement functions will be added in the next release.torchaudio.io.StreamReader supports decoding media from byte strings contained in 1D tensors of torch.uint8 type. Using torch.Tensor type as a container for byte string is now deprecated. To pass byte strings, please wrap the string with io.BytesIO.
<table class="tg">
<thead>
<tr>
<th class="tg-0pky">Deprecated</th>
<th class="tg-0pky">Migration</th>
</tr>
</thead>
<tbody>
<tr>
<td class="tg-dvpl"><code>data = b"..."</code></br><code>src = torch.frombuffer(data, dtype=torch.uint8)</code></br><code>StreamReader(src)</code></td>
<td class="tg-dvpl"><code>data = b"..."</code></br><code>src = io.BytesIO(data)</code></br><code>StreamReader(src)</code></td>
</tr>
</tbody>
</table>torchaudio.functional.lfilter (#3080)Without the change in #2873, the WER results are:
| Model | dev-clean | dev-other | test-clean | test-other |
|---|---|---|---|---|
| WAV2VEC2_ASR_LARGE_LV60K_10M | 10.59 | 15.62 | 9.58 | 16.33 |
| WAV2VEC2_ASR_LARGE_LV60K_100H | 2.80 | 6.01 | 2.82 | 6.34 |
| WAV2VEC2_ASR_LARGE_LV60K_960H | 2.36 | 4.43 | 2.41 | 4.96 |
| HUBERT_ASR_LARGE | 1.85 | 3.46 | 2.09 | 3.89 |
| HUBERT_ASR_XLARGE | 2.21 | 3.40 | 2.26 | 4.05 |
After applying layer normalization, the updated WER results are:
| Model | dev-clean | dev-other | test-clean | test-other |
|---|---|---|---|---|
| WAV2VEC2_ASR_LARGE_LV60K_10M | 6.77 | 10.03 | 6.87 | 10.51 |
| WAV2VEC2_ASR_LARGE_LV60K_100H | 2.19 | 4.55 | 2.32 | 4.64 |
| WAV2VEC2_ASR_LARGE_LV60K_960H | 1.78 | 3.51 | 2.03 | 3.68 |
| HUBERT_ASR_LARGE | 1.77 | 3.32 | 2.03 | 3.68 |
| HUBERT_ASR_XLARGE | 1.73 | 2.72 | 1.90 | 3.16 |
shuffle is set True in BucketizeBatchSampler, the seed is only the same for the first epoch. In later epochs, each BucketizeBatchSampler object will generate a different shuffled iteration list, which may cause DPP training to hang forever if the lengths of iteration lists are different across nodes. In the 2.0.0 release, the issue is fixed by using the same seed for RNG in all nodes._fail_info_fileobj (#3032)torchaudio.io.StreamReader.torchaudio.functional.lfilter (#3018)AddNoise, Convolve, FFTConvolve, Speed, SpeedPerturbation, Deemphasis, and Preemphasis in torchaudio.transforms, and add_noise, fftconvolve, convolve, speed, preemphasis, and deemphasis in torchaudio.functional.fill_buffer method to torchaudio.io.StreamReader (#2954, #2971)buffer_chunk_size=-1 option to torchaudio.io.StreamReader (#2969)
When buffer_chunk_size=-1, StreamReader does not drop any buffered frame. Together with the fill_buffer method, this is a recommended way to load the entire media.reader = StreamReader("video.mp4")
reader.add_basic_audio_stream(buffer_chunk_size=-1)
reader.add_basic_video_stream(buffer_chunk_size=-1)
reader.fill_buffer()
audio, video = reader.pop_chunks()
torchaudio.io.StreamReader (#2975)
torchaudio.io.SteramReader now gives PTS (presentation time stamp) of the media chunk it is returning. To maintain backward compatibility, the timestamp information is attached to the returned media chunk.reader = StreamReader(...)
reader.add_basic_audio_stream(...)
reader.add_basic_video_stream(...)
for audio_chunk, video_chunk in reader.stream():
# Fetch timestamp
print(audio_chunk.pts)
print(video_chunk.pts)
# Chunks behave the same as torch.Tensor.
audio_chunk.mean(dim=1)
torchaudio.io.play_audio (#3026, #3051)
You can play audio with the torchaudio.io.play_audio function. (macOS only)torchaudio.utils.ffmpeg_utils, which can be used to query into the dynamically linked FFmpeg libraries.
get_demuxers()get_muxers()get_audio_decoders()get_audio_encoders()get_video_decoders()get_video_encoders()get_input_devices()get_output_devices()get_input_protocols()get_output_protocols()get_build_config()Refactor StreamReader/Writer implementation
torchaudio::ffmpeg namespace with torchaudio::io (#3013)pop_chunks implementations (#3002)Added logging to torchaudio.io.StreamReader/Writer (#2878)
Fixed the #threads used by FilterGraph to 1 (#2985)
Fixed the default #threads used by decoder to 1 in torchaudio.io.StreamReader (#2949)
Moved libsox integration from libtorchaudio to libtorchaudio_sox (#2929)
Added query methods to FilterGraph (#2976)
cuda_version (#2952)USE_CUDA detection (#3005)USE_ROCM detection (#3008)Nothing published for this version
This is a minor release, which is compatible with PyTorch 1.13.1 and includes bug fixes, improvements and documentation updates. There is no new featu
This is a minor release, which is compatible with PyTorch 1.13.1 and includes bug fixes, improvements and documentation updates. There is no new feature added.
TorchAudio 0.13.0 release includes:
TorchAudio 0.13.0 release includes:
Hybrid Demucs is a music source separation model that uses both spectrogram and time domain features. It has demonstrated state-of-the-art performance in the Sony Music DeMixing Challenge. (citation: https://arxiv.org/abs/2111.03600)
The TorchAudio v0.13 release includes the following features
SDR Results of pre-trained pipelines on MUSDB-HQ test set
| Pipeline | All | Drums | Bass | Other | Vocals |
|---|---|---|---|---|---|
| HDEMUCS_HIGH_MUSDB* | 6.42 | 7.76 | 6.51 | 4.47 | 6.93 |
| HDEMUCS_HIGH_MUSDB_PLUS** | 9.37 | 11.38 | 10.53 | 7.24 | 8.32 |
* Trained on the training data of MUSDB-HQ dataset. ** Trained on both training and test sets of MUSDB-HQ and 150 extra songs from an internal database that were specifically produced for Meta.
Special thanks to @adefossez for the guidance.
ConvTasNet model architecture was added in TorchAudio 0.7.0. It is the first source separation model that outperforms the oracle ideal ratio mask. In this release, TorchAudio adds the pre-trained pipeline that is trained within TorchAudio on the Libri2Mix dataset. The pipeline achieves 15.6dB SDR improvement and 15.3dB Si-SNR improvement on the Libri2Mix test set.
With the addition of four new audio-related datasets, there is now support for all downstream tasks in version 1 of the SUPERB benchmark. Furthermore, these datasets support metadata mode through a get_metadata function, which enables faster dataset iteration or preprocessing without the need to load or store waveforms.
Datasets with metadata functionality:
In release 0.12, TorchAudio released a CTC beam search decoder with KenLM language model support. This release, there is added functionality for creating custom Python language models that are compatible with the decoder, using the torchaudio.models.decoder.CTCDecoderLM wrapper.
torchaudio.io.StreamWriter is a class for encoding media including audio and video. This can handle a wide variety of codecs, chunk-by-chunk encoding and GPU encoding.
GriffinLim implementations in transforms and functional used the momentum parameter differently, resulting in inconsistent results between the two implementations. The transforms.GriffinLim usage of momentum is updated to resolve this discrepancy.torchaudio.info decode audio to compute num_frames if it is not found in metadata (#2740).
In such cases, torchaudio.info may now return non-zero values for num_frames.torchaudio.compliance.kaldi.fbank with dither option produced a different output from kaldi because it used a skewed, rather than gaussian, distribution for dither. This is updated in this release to correctly use a random gaussian instead.runtime_error exception with TORCH_CHECK (#2550, #2551, #2592)torchaudio.functional.resample function using the sinc resampling method, on float32 tensor with two channels and one second duration.CPU
| torchaudio version | 8k → 16k [Hz] | 16k → 8k | 16k → 44.1k | 44.1k → 16k |
|---|---|---|---|---|
| 0.13 | 0.256 | 0.549 | 0.769 | 0.820 |
| 0.12 | 0.386 | 0.534 | 31.8 | 12.1 |
CUDA
| torchaudio version | 8k → 16k [Hz] | 16k → 8k | 16k → 44.1k | 44.1k → 16k |
|---|---|---|---|---|
| 0.13 | 0.332 | 0.336 | 0.345 | 0.381 |
| 0.12 | 0.524 | 0.334 | 64.4 | 22.8 |
WER improvement on LibriSpeech dev and test sets
| Viterbi (v0.12) | Viterbi (v0.13) | KenLM (v0.12) | KenLM (v0.13) | |
|---|---|---|---|---|
| dev-clean | 10.7 | 10.9 | 4.4 | 4.2 |
| dev-other | 18.3 | 17.5 | 9.7 | 9.4 |
| test-clean | 10.8 | 10.9 | 4.4 | 4.4 |
| test-other | 18.5 | 17.8 | 10.1 | 9.5 |
:autosummary: in torchaudio docs (#2664, #2681, #2683, #2684, #2693, #2689, #2690, #2692)This is a minor release, which is compatible with PyTorch 1.12.1 and include small bug fixes, improvements and documentation update. There is no new f
This is a minor release, which is compatible with PyTorch 1.12.1 and include small bug fixes, improvements and documentation update. There is no new feature added.
For the full feature of v0.12, please refer to the v0.12.0 release note.
TorchAudio 0.12.0 includes the following:
TorchAudio 0.12.0 includes the following:
To support inference-time decoding, the release adds the wav2letter CTC beam search decoder, ported over from Flashlight (GitHub). Both lexicon and lexicon-free decoding are supported, and decoding can be done without a language model or with a KenLM n-gram language model. Compatible token, lexicon, and certain pretrained KenLM files for the LibriSpeech dataset are also available for download.
For usage details, please check out the documentation and ASR inference tutorial.
To improve flexibility in usage, the release adds two new beamforming modules under torchaudio.transforms: SoudenMVDR and RTFMVDR. They differ from MVDR mainly in that they:
reference_channel as an input argument in the forward method to allow users to select the reference channel in model training or dynamically change the reference channel in inference.Besides the two modules, the release adds new function-level beamforming methods under torchaudio.functional. These include
For usage details, please check out the documentation at torchaudio.transforms and torchaudio.functional and the Speech Enhancement with MVDR Beamforming tutorial.
StreamReader is TorchAudio’s new I/O API. It is backed by FFmpeg† and allows users to
For usage details, please check out the documentation and tutorials:
† To use StreamReader, FFmpeg libraries are required. Please install FFmpeg. The coverage of codecs depends on how these libraries are configured. TorchAudio official binaries are compiled to work with FFmpeg 4 libraries; FFmpeg 5 can be used if TorchAudio is built from source.
torchaudio.load, please install a compatible version of FFmpeg (Version 4 when using an official binary distribution).torchaudio.info now returns num_frames=0 for MP3.Hypothesis subclassed namedtuple. Containers of namedtuple instances, however, are incompatible with the PyTorch Lite Interpreter. To achieve compatibility, Hypothesis has been modified in release 0.12 to instead alias tuple. This affects RNNTBeamSearch as it accepts and returns a list of Hypothesis instances.complex128 to improve the precision and robustness of downstream matrix computations. The output dtype, however, is not correctly converted back to the original dtype. In release 0.12, we fix the output dtype to be consistent with the original input dtype.torchaudio.transforms.PitchShift, after its first call, to perform the operation on float32 Tensor with two channels and 8000 frames, resampled to 44.1 kHz across various shifted steps.| TorchAudio Version | 2 | 3 | 4 | 5 |
|---|---|---|---|---|
| 0.12 | 2.76 | 5 | 1860 | 223 |
| 0.11 | 6.71 | 161 | 8680 | 1450 |
__getattr__ to implement delayed initialization (#2377)Removed deprecated F.magphase, F.angle, F.complex_norm, and T.ComplexNorm. (#1934, #1935, #1942)
TorchAudio 0.11.0 release includes:
To support streaming ASR use cases, the release adds implementations of Emformer (docs), an RNN-T model that uses Emformer (emformer_rnnt_base), and an RNN-T beam search decoder (RNNTBeamSearch). It also includes a pipeline bundle (EMFORMER_RNNT_BASE_LIBRISPEECH) that wraps pre- and post-processing components, the beam search decoder, and the RNN-T Emformer model with weights pre-trained on LibriSpeech, which in whole allow for performing streaming ASR inference out of the box. For reference and reproducibility, the release provides the training recipe used to produce the pre-trained weights in the examples directory.
The masked prediction training of HuBERT model requires the masked logits, unmasked logits, and feature norm as the outputs. The logits are for cross-entropy losses and the feature norm is for penalty loss. The release adds HuBERTPretrainModel and corresponding factory functions (hubert_pretrain_base, hubert_pretrain_large, and hubert_pretrain_xlarge) to enable training from scratch.
The release adds an implementation of Conformer (docs), a convolution-augmented transformer architecture that has achieved state-of-the-art results on speech recognition benchmarks.
F.magphase, F.angle, F.complex_norm, and T.ComplexNorm. (#1934, #1935, #1942)
F.spectrogram, T.Spectrogram, F.phase_vocoder, and T.TimeStretch (#1957, #1958)
create_fb_matrix (#1998)
create_fb_matrix was replaced by melscale_fbanks in release 0.10. It is removed in 0.11. Please use melscale_fbanks.VCTK_092 class for the latest version of the dataset.diskcache_iterator and bg_iterator were deprecated in 0.10. They are removed in 0.11. Please cease the usage of them.<s>, <pad>, </s>, <unk>) that were not related to ASR tasks and not used. These dimensions were removed.This is a minor release compatible with PyTorch 1.10.2.
This is a minor release compatible with PyTorch 1.10.2.
There is no feature change in torchaudio from 0.10.1. For the full feature of v0.10, please refer to the v0.10.0 release notes.
This is a minor release, which is compatible with PyTorch 1.10.1 and include small bug fix, improvements and documentation update. There is no new fea
This is a minor release, which is compatible with PyTorch 1.10.1 and include small bug fix, improvements and documentation update. There is no new feature added.
TORCH_CUDA_ARCH_LIST delimiterFor the full feature of v0.10, please refer to the v0.10.0 release note.
Remove deprecated kaldi.resample_waveform
torchaudio 0.10.0 release includes:
HuBERT model architectures (“base”, “large” and “extra large” configurations) are added. In addition to that, support for pretrained weights from wav2vec 2.0, Unsupervised Cross-lingual Representation Learning and HuBERT are added.
These pretrained weights can be used for feature extractions and downstream task adaptation.
>>> import torchaudio
>>>
>>> # Build the model and load pretrained weight.
>>> model = torchaudio.pipelines.HUBERT_BASE.get_model()
>>> # Perform feature extraction.
>>> features, lengths = model.extract_features(waveforms)
>>> # Pass the features to downstream task
>>> ...
Some of the pretrained weights are fine-tuned for ASR tasks. The following example illustrates how to use weights and access to associated information, such as labels, which can be used in subsequent CTC decoding steps. (Note: torchaudio does not provide a CTC decoding mechanism.)
>>> import torchaudio
>>>
>>> bundle = torchaudio.pipelines.HUBERT_ASR_LARGE
>>>
>>> # Build the model and load pretrained weight.
>>> model = bundle.get_model()
Downloading:
100%|███████████████████████████████| 1.18G/1.18G [00:17<00:00, 73.8MB/s]
>>> # Check the corresponding labels of the output.
>>> labels = bundle.get_labels()
>>> print(labels)
('<s>', '<pad>', '</s>', '<unk>', '|', 'E', 'T', 'A', 'O', 'N', 'I', 'H', 'S', 'R', 'D', 'L', 'U', 'M', 'W', 'C', 'F', 'G', 'Y', 'P', 'B', 'V', 'K', "'", 'X', 'J', 'Q', 'Z')
>>>
>>> # Infer the label probability distribution
>>> waveform, sample_rate = torchaudio.load(hello-world.wav')
>>>
>>> emissions, _ = model(waveform)
>>>
>>> # Pass emission to (hypothetical) decoder
>>> transcripts = ctc_decode(emissions, labels)
>>> print(transcripts[0])
HELLO WORLD
A new model architecture, Tacotron2 is added, alongside several pretrained weights for TTS (text-to-speech). Since these TTS pipelines are composed of multiple models and specific data processing, so as to make it easy to use associated objects, a notion of bundle is introduced. Bundles provide a common access point to create a pipeline with a set of pretrained weights. They are available under torchaudio.pipelines module.
The following example illustrates a TTS pipeline where two models (Tacotron2 and WaveRNN) are used together.
>>> import torchaudio
>>>
>>> bundle = torchaudio.pipelines.TACOTRON2_WAVERNN_CHAR_LJSPEECH
>>>
>>> # Build text processor, Tacotron2 and vocoder (WaveRNN) model
>>> processor = bundle.get_text_preprocessor()
>>> tacotron2 = bundle.get_tacotron2()
Downloading:
100%|███████████████████████████████| 107M/107M [00:01<00:00, 87.9MB/s]
>>> vocoder = bundle.get_vocoder()
Downloading:
100%|███████████████████████████████| 16.7M/16.7M [00:00<00:00, 78.1MB/s]
>>>
>>> text = "Hello World!"
>>>
>>> # Encode text
>>> input, lengths = processor(text)
>>>
>>> # Generate (mel-scale) spectrogram
>>> specgram, lengths, _ = tacotron2.infer(input, lengths)
>>>
>>> # Convert spectrogram to waveform
>>> waveforms, lengths = vocoder(specgram, lengths)
>>>
>>> # Save audio
>>> torchaudio.save('hello-world.wav', waveforms, vocoder.sample_rate)
The loss function used in the RNN transducer architecture, which is widely used for speech recognition tasks, is added. The loss function (torchaudio.functional.rnnt_loss or torchaudio.transforms.RNNTLoss) supports float16 and float32 logits, has autograd and torchscript support, and can be run on both CPU and GPU, which has a custom CUDA kernel implementation for improved performance.
This release adds support for MVDR beamforming on multi-channel audio using Time-Frequency masks. There are three solutions (ref_channel, stv_evd, stv_power) and it supports single-channel and multi-channel (perform average in the method) masks. It provides an online option that recursively updates the parameters for streaming audio. Please refer to the MVDR tutorial.
This release adds GPU builds that support custom CUDA kernels in torchaudio, like the one being used for RNN transducer loss. Following this change, torchaudio’s binary distribution now includes CPU-only versions and CUDA-enabled versions. To use CUDA-enabled binaries, PyTorch also needs to be compatible with CUDA.
torchaudio.functional.lfilter now supports batch processing and multiple filters. Additional operations, including pitch shift, LFCC, and inverse spectrogram, are now supported in this release. The datasets CMUDict and LibriMix are added as well.
PCM_24 (the previous default) could cause warping. The default has been changed to PCM_16, which does not suffer this.power=None, torchaudio.functional.spectrogram and torchaudio.transforms.Spectrogram now defaults to return_complex=True, which returns Tensor of native complex type (such as torch.cfloat and torch.cdouble). To use a pseudo complex type, pass the resulting tensor to torch.view_as_real.torchaudio.functional.resample.specgram.extract_features of Wav2Vec2Model (#1776)
Wav2Vec2Model.feature_extractor().Wav2Vec2Model was updated. Wav2Vec2Model.encoder.read_out module is moved to Wav2Vec2Model.aux. If you have serialized state dict, please replace the key encoder.read_out with aux.num_out parameter has been changed to aux_num_out and other parameters are added before it. Please update the code from wav2vec2_base(num_out) to wav2vec2_base(aux_num_out=num_out).melscale_fbanks and deprecate create_fb_matrix (#1653)
linear_fbanks is introduced, create_fb_matrix is renamed to melscale_fbanks. The original create_fb_matrix is now deprecated. Please use melscale_fbanks.VCTK dataset (#1810)
VCTK_092 dataset.bg_iterator and diskcache_iterator are known to not improve the throughput of data loaders. Please cease their usage.Tacotron2
HuBERT
CUDA
| Tensor shape | [1,4,8000] | [1,4,16000] | [1,4,32000] |
|---|---|---|---|
| 0.10 | 119 | 120 | 123 |
| 0.9 | 160 | 184 | 240 |
Unit: msec
This release depends on pytorch 1.9.1 No functional changes other than minor updates to CI rules.
This release depends on pytorch 1.9.1 No functional changes other than minor updates to CI rules.
However, please note that the use of the pseudo complex type is now deprecated. These functions are tested to support TorchScript and autograd. For th…
torchaudio 0.9.0 release includes:
This release includes model architectures from wav2vec2.0 paper with utility functions that allow importing pretrained model parameters published on <code>fairseq</code> and Hugging Face Hub. Now you can easily run speech recognition with torchaudio. These model architectures also support TorchScript, and you can deploy them with ONNX or in non-Python environments, such as C++, Android and iOS. Please checkout our C++, Android and iOS examples. The following snippets illustrate how to create a deployable model.
# Import fine-tuned model from Hugging Face Hub
import transformers
from torchaudio.models.wav2vec2.utils import import_huggingface_model
original = Wav2Vec2ForCTC.from_pretrained("facebook/wav2vec2-base-960h")
imported = import_huggingface_model(original)
# Import fine-tuned model from fairseq
import fairseq
from torchaudio.models.wav2vec2.utils import import_fairseq_model
Original, _, _ = fairseq.checkpoint_utils.load_model_ensemble_and_task(
["wav2vec_small_960h.pt"], arg_overrides={'data': "<data_dir>"})
imported = import_fairseq_model(original[0].w2v_encoder)
# Build uninitialized model and load state dict
from torchaudio.models import wav2vec2_base
model = wav2vec2_base(num_out=32)
model.load_state_dict(imported.state_dict())
# Quantize / script / optimize for mobile
quantized_model = torch.quantization.quantize_dynamic(
model, qconfig_spec={torch.nn.Linear}, dtype=torch.qint8)
scripted_model = torch.jit.script(quantized_model)
optimized_model = optimize_for_mobile(scripted_model)
optimized_model.save("model_for_deployment.pt")
The internal implementation of lfilter has been updated to support autograd on both CPU and CUDA. Additionally, the performance on CPU is significantly improved. These improvements also apply to biquad variants.
The following table illustrates the performance improvements compared against the previous releases. lfilter was applied on float32 tensors with one channel and different number of frames.
<table> <tr> <td>torchaudio version </td> <td><p style="text-align: right">256</p> </td> <td><p style="text-align: right">512</p> </td> <td><p style="text-align: right">1024</p> </td> </tr> <tr> <td>0.9 </td> <td><p style="text-align: right"> 0.282</p>
</td> <td><p style="text-align: right"> 0.381</p>
</td> <td><p style="text-align: right"> 0.564</p>
</td> </tr> <tr> <td>0.8 </td> <td><p style="text-align: right"> 0.493</p>
</td> <td><p style="text-align: right"> 0.780</p>
</td> <td><p style="text-align: right"> 1.37</p>
</td> </tr> <tr> <td>0.7 </td> <td><p style="text-align: right"> 5.42</p>
</td> <td><p style="text-align: right"> 10.8</p>
</td> <td><p style="text-align: right"> 22.3</p>
</td> </tr> </table>
Unit: msec
torchaudio has functions that handle complex-valued tensors. In early days when PyTorch did not have a complex dtype, torchaudio adopted the convention to use an extra dimension to represent real and imaginary parts. In PyTorch 1.6, new dtyps, such as torch.cfloat and torch.cdouble were introduced to represent complex values natively. (In the following, we refer to torchaudio’s original convention as pseudo complex types, and PyTorch’s native dtype as native complex types.)
As the native complex types have become mature and stable, torchaudio has started to migrate complex functions to use the native complex type. In this release, the internal implementation was updated to use the native complex types, and interfaces were updated to allow passing/receiving native complex type directly. Users can choose to keep using the pseudo complex type or opt in to use native complex type. However, please note that the use of the pseudo complex type is now deprecated. These functions are tested to support TorchScript and autograd. For the detail of this migration plan, please refer to #1337.
Additionally, switching the internal implementation to the native complex types improved the performance. Since the internal implementation uses native complex type regardless of which complex type is passed/returned, users will automatically benefit from this performance improvement.
The following table illustrates the performance improvements from the previous release by comparing the time it takes for complex transforms to perform operation on float32 Tensor with two channels and 256 frames.
<table> <tr> <td>torchaudio version </td> <td><code>Spectrogram</code> </td> <td><code>TimeStretch</code> </td> <td><code>GriffinLim</code> </td> </tr> <tr> <td>0.9 </td> <td><p style="text-align: right"> 0.229</p>
</td> <td><p style="text-align: right"> 12.6</p>
</td> <td><p style="text-align: right"> 3320</p>
</td> </tr> <tr> <td>0.8 </td> <td><p style="text-align: right"> 0.283</p>
</td> <td><p style="text-align: right"> 126</p>
</td> <td><p style="text-align: right"> 5320</p>
</td> </tr> </table>
Unit: msec
<table> <tr> <td>torchaudio version </td> <td><code>Spectrogram</code> </td> <td><code>TimeStretch</code> </td> <td><code>GriffinLim</code> </td> </tr> <tr> <td>0.9 </td> <td><p style="text-align: right"> 0.195</p>
</td> <td><p style="text-align: right"> 0.599</p>
</td> <td><p style="text-align: right"> 36</p>
</td> </tr> <tr> <td>0.8 </td> <td><p style="text-align: right"> 0.219</p>
</td> <td><p style="text-align: right"> 0.687</p>
</td> <td><p style="text-align: right"> 60.2</p>
</td> </tr> </table>
Unit: msec
Along with the work of Complex Tensor Migration and Filtering Improvement mentioned above, more tests were added to ensure the autograd support. Now the following operations are guaranteed to support autograd up to second order.
lfilterallpass_biquadbiquadband_biquadbandpass_biquadbandrefect_biquadbass_biquadequalizer_biquadtreble_biquadhighpass_biquadlowpass_biquadAmplitudeToDBComputeDeltasFadeGriffinLimTimeMaskingFrequencyMaskingMFCCMelScaleMelSpectrogramResampleSpectralCentroidSpectrogramSlidingWindowCmnTimeStretch*VolNOTE:
amplitude_to_DBspectrogramgriffinlimresamplephase_vocoder*mask_along_axis_iidmask_along_axisgainspectral_centroidtorchaudio.transforms.TimeStretch and torchaudio.functional.phase_vocoder call atan2, which is not differentiable around zero. Therefore these functions are differentiable only when the input spectrogram does not contain values around zero.In release 0.8, the resampling operation was vectorized and its performance improved. In this release, the implementation of the resampling algorithm has been further revised.
rolloff parameter has been added for anti-aliasing control.torchaudio.transforms.Resample precomputes the kernel using float64 precision and caches it for even faster operation.torchaudio.functional.resample has been added and the original entry point, torchaudio.compliance.kaldi.resample_waveform is deprecated.The following table illustrates the performance improvements from the previous release by comparing the time it takes for torchaudio.transforms.Resample to complete the operation on float32 tensor with two channels and one-second duration.
<table> <tr> <td>torchaudio version </td> <td>8k → 16k [Hz] </td> <td>16k → 8k </td> <td>16k → 44.1k </td> <td>44.1k → 16k </td> </tr> <tr> <td>0.9 </td> <td><p style="text-align: right"> 0.192</p>
</td> <td><p style="text-align: right"> 0.559</p>
</td> <td><p style="text-align: right"> 0.478</p>
</td> <td><p style="text-align: right"> 0.467</p>
</td> </tr> <tr> <td>0.8 </td> <td><p style="text-align: right"> 0.537</p>
</td> <td><p style="text-align: right"> 0.753</p>
</td> <td><p style="text-align: right"> 43.9</p>
</td> <td><p style="text-align: right"> 17.6</p>
</td> </tr> </table>
Unit: msec
<table> <tr> <td>torchaudio version </td> <td>8k → 16k </td> <td>16k → 8k </td> <td>16k → 44.1k </td> <td>44.1k → 16k </td> </tr> <tr> <td>0.9 </td> <td><p style="text-align: right"> 0.203</p>
</td> <td><p style="text-align: right"> 0.172</p>
</td> <td><p style="text-align: right"> 0.213</p>
</td> <td><p style="text-align: right"> 0.212</p>
</td> </tr> <tr> <td>0.8 </td> <td><p style="text-align: right"> 0.860</p>
</td> <td><p style="text-align: right"> 0.559</p>
</td> <td><p style="text-align: right"> 116</p>
</td> <td><p style="text-align: right"> 46.7</p>
</td> </tr> </table>
Unit: msec
torchaudio implements some operations in C++ for reasons such as performance and integration with third-party libraries. This C++ module was only available on Linux and macOS. In this release, Windows packages also come with C++ module.
This C++ module in Windows package includes the efficient filtering implementation mentioned above, however, “sox_io” backend and torchaudio.functional.compute_kaldi_pitch are not included.
Since the 0.6 release, we have continuously improved I/O functionality. Specifically, in 0.8 the default backend has been changed from “sox” to “sox_io”, and the similar API change has been applied to “soundfile” backend. The 0.9 release concludes this migration by removing the deprecated backends. For the detail please refer to #903.
normalized argument from torchaudio.functional.griffinlim (#1369)
torchaudio.functional.sliding_window_cmn arg for correctness (#1347)
waveform=..., please change it to specgram=...torchaudio.transforms.Resample to precompute and cache the resampling kernel. (#1499, #1514)
resampler = torchaudio.transforms.Resample(orig_freq=8000, new_freq=44100)
resampler.to(torch.device("cuda"))
torchaudio no longer supports programmatic download of Common Voice dataset. Please remove the arguments from your code.torchaudio is adopting native complex type and the use of pseudo complex type and the related utility functions are now deprecated. Please refer to #1337 for the migration process.torchaudio.compliance.kaldi.resample_waveform (#1533)
torchaudio.functional.resample.torchaudio.transforms.MelScale now expects valid n_stft value (#1515)
n_stft.torchaudio.functional.lfilter (#1319)torchaudio.functional.lfilter (#1310, #1441)torchaudio.functional.resample (#1402)rolloff parameter (#1488)torchaudio.transforms.Resample (#1499, #1514, #1556)torchaudio.functional.phase_vocoder and torchaudio.transforms.TimeStretch (#1410)return_complex to torchaudio.functional.spectrogram and torchaudio.transforms.Spectrogram (#1366, #1551)__str__ override to AudioMetaData for easy print (#1339)sox/utils.cpp (#1306)check_length from validate_input_file (#1312)torchaudio.functional.griffinlim (#1368)torchaudio.transforms.MelScale when n_stft is invalid (#1505)__all__ (#1458)reference_cast in make_boxed_from_unboxed_functor (#1300)torchaudio.transforms.GriffinLim (#1433)librosa's Mel scale conversion with torchaudio’s in WaveRNN example (#1444)config.guess to support source build in recent architectures (#1484)torchaudio.functional.lfilter and biquad variants (#1400, #1438)torchaudio.transforms.FrequencyMasking (#1498)torchaudio.transforms.SlidingWindowCmn (#1482)torchaudio.transforms.MelScale (#1467)torchaudio.transforms.Vol (#1460)torchaudio.transforms.TimeStretch (#1420)torchaudio.transforms.AmplitudeToDB (#1447)torchaudio.transforms.GriffinLim (#1421)torchaudio.transforms.SpectralCentroid (#1425)torchaudio.transforms.ComputeDeltas (#1422)torchaudio.transforms.Fade (#1424)torchaudio.transforms.Resample (#1416)torchaudio.transforms.MFCC (#1415)torchaudio.transforms.Spectrogram / MelSpectrogram (#1340)torchaudio.functional.lfilter shape (#1360)torchaudio.functional.resample (#1516)torchaudio.functional.phase_vocoder (#1379)floor_divide with div (#1455)torch.assert_allclose with assertEqual (#1387)torchaudio.functional.lfilter autograd tests input size (#1443)torchaudio.transforms.InverseMelScale comparison test (#1437)torchaudio.transforms.TimeMasking and torchaudio.transforms.FrequencyMasking to perform out-of-place masking (#1481)power of torchaudio.transforms.MelSpectrogram as float only (#1572)torch.nn.functional.conv1d in torchaudio.functional.lfilter (#1318)torchaudio.functional.overdrive (#1299)sox_effects.apply_effects_tensor is CPU-only (#1459)sliding_window_cmn (#1383)This release depends on pytorch 1.8.1.
This release depends on pytorch 1.8.1.
Following the change of default backends, the legacy backend/interface have been marked as deprecated. The legacy backend/interface are still accessib…
This release supports Python 3.9.
Continuing from the previous release, torchaudio improves the audio I/O mechanism. In this release, we have four major updates.
Backend migration. We have migrated the default backend for audio I/O. The new default backend is “sox_io” (for Linux/macOS). The interface for “soundfile” backend has been also changed to align that of “sox_io”. Following the change of default backends, the legacy backend/interface have been marked as deprecated. The legacy backend/interface are still accessible, though it is strongly discouraged to use them. For the detail on the migration, please refer to #903.
File-like object support.
We have added file-like object support to I/O functions and sox_effects. You can perform the info, load, save and apply_effects_file operation on file-like objects.
# Query audio metadata over HTTP
# Will only fetch the first few kB
with requests.get(URL, stream=True) as response:
metadata = torchaudio.info(response.raw)
# Load audio from tar file
# No need to extract TAR file.
with tarfile.open(TAR_PATH, mode='r') as tarfile_:
fileobj = tarfile_.extractfile(SAMPLE_TAR_ITEM)
waveform, sample_rate = torchaudio.load(fileobj)
# Saving to Bytes buffer
# Using BytesIO, you can perform in-memory encoding/decoding.
buffer_ = io.BytesIO()
torchaudio.save(buffer_, waveform, sample_rate, format="wav")
# Apply effects (lowpass filter / resampling) while loading audio from S3
client = boto3.client('s3')
response = client.get_object(Bucket=S3_BUCKET, Key=S3_KEY)
waveform, sample_rate = torchaudio.sox_effects.apply_effect_file(
response['Body'], [["lowpass", "-1", "300"], ["rate", "8000"]])
[Beta] Codec Application.
Built upon the file-like object support, we added functional.apply_codec function, which can degrades audio data by applying audio codecs supported by “sox_io” backend, in in-memory fashion.
# Apply MP3 codec
degraded = F.apply_codec(
waveform, sample_rate, format="mp3", compression=-9)
# Apply GSM codec
degraded = F.apply_codec(waveform, sample_rate, format="gsm")
Encoding options.
We have added encoding options to save function of new backends. Now you can change the format and encodings with format, encoding and bits_per_sample options
# Save without any encoding option.
# The function will pick the encoding which the provided data fit
# For Tensor of float32 type, that is 32-bit floating-point PCM.
torchaudio.save("data.wav", waveform, sample_rate)
# Save as 16-bit signed integer Linear PCM
# The resulting file occupies half the storage but loses precision
torchaudio.save(
"data.wav", waveform, sample_rate, encoding="PCM_S", bits_per_sample=16)
More format support to "sox_io"’s save function. We have added support for GSM, HTK, AMB, and AMR-NB formats to "sox_io"’s save function.
torchaudio was utilizing CMake to build third party dependencies. Now torchaudio uses CMake to build its C++ extension. This will open the door to integrate torchaudio in non-Python environments (such as C++ applications and mobile). We will work on adding example applications and mobile integrations in upcoming releases.
This release introduces support for python 3.9. There is no 0.7.1 release, and the following changes are compared to 0.7.0.
This release introduces support for python 3.9. There is no 0.7.1 release, and the following changes are compared to 0.7.0.
download=True in CommonVoice (#1076)Deprecated SoxEffect and SoxEffectsChain
torchaudio is expanding its support for models and end-to-end applications. Please file an issue on github to provide feedback on them.
As you are likely already aware from the last release we’re currently in the process of making sox_io, which ships with new features such as TorchScript support and performance improvements, the new default. If you want to benefit from these features now, we encourage you to migrate. For more information see issue #903.
str.format to adopt changes in PyTorch, leading to improved error messages for TorchScript (#850)sox_utils.list_formats() for read and write (#811)VCTK_092 dataset (#812)sox_io backend (#871)soundfile backend to the one identical to sox_io backend. (#922)soundfile compatibility backend. (#922)torchaudio.compliance.kaldi.fbank (#947)pathlib.Path support to sox_io backend (#907)sox_io C++ implementation (#779)sox_io and sox_effects (#806)noise_shaping = True (#865)zip_safe = False to disable egg installation (#842)istft wrapper in favor of torch.istft. (#841)SoxEffect and SoxEffectsChain (#787)sox backend. (#904)soundfile. (#922)load_wav functions. (#905)Since sox_effects is now automatically initialized and shutdown (#572, #693), we are deprecating these functions (#709).
torchaudio now includes a new model module (with wav2letter included), new functionals (contrast, cvm, dcshift, overdrive, vad, phaser, flanger, biquad), datasets (GTZAN, CMU), and a new optional sox backend with support for torchscript. torchaudio now also supports Windows, with the soundfile backend.
torchaudio requires python 3.6 or more recent.
Updated pinned version of PyTorch to `v1.5.1`
v1.5.1Upgrading dataset DeprecationWarning to UserWarning so that the user gets the warning.
torchaudio includes new transforms (e.g. Griffin-Lim and inverse Mel scale), new filters (e.g. all pass, fade, band pass/reject, band, treble, deemph, riaa), and datasets (LJ Speech and SpeechCommands).
In particular, the construction parameters downsample, transform, target_transform, and return_dict are being deprecated.
torchaudio 0.4 improves on current transformations, datasets, and backend support.
We would like to thank again our contributors and the wider community for their significant contributions to this release. In particular we'd like to thank @keunwoochoi, @ksanjeevan, and all the other maintainers and contributors of torchaudio-contrib for their significant and valuable additions around augmentations (#285) and batching (#327).
downsample, transform, target_transform, and return_dict are being deprecated.torchaudio.functional.detect_pitch_frequency. (#313, #322)torchaudio.transforms: TimeStretch, FrequencyMasking, TimeMasking. (#285, #333, #348)torchaudio.transform.ComplexNorm. (#285, #333)torchaudio.functional.compute_deltas. (#268, #326)torchaudio.functional.gain and torchaudio.functional.dither (#319, #360). We welcome work to continue the effort to implement features available in SoX, see #260.equalizer_biquad (#315, #340), lowpass_biquad, highpass_biquad (#275), lfilter, and biquad (#275, #291, #326) in torchaudio.functional.torchaudio.functional.mfcc. (#228)MelScale and librosa. (#294)torchaudio.compliance.kaldi.resample_waveform where internal variables where not moved to the GPU when used. (#277)istft where the dtype and device of parameters were not created on the same device as the tensor provided by the user. (#264)load_state_dict). (#246)torchaudio.load to [-1, 1]. (#283)This release is to update the dependency to PyTorch 1.3.1.
This release is to update the dependency to PyTorch 1.3.1.
This release is to update the dependency to PyTorch 1.3.0.
This release is to update the dependency to PyTorch 1.3.0.
…and conventions. See the section on backwards breaking changes for a migration guide.
torchaudio has been redesigned to be an extension of PyTorch and part of the domain APIs (DAPI) ecosystem. Domain specific libraries such as this one are kept separated in order to maintain a coherent environment for each of them. As such, torchaudio is an ML library that provides relevant signal processing functionality, but it is not a general signal processing library. The full rationale of this new standardization can be found in the README.md.
In light of these changes some transforms have been removed or have different argument names and conventions. See the section on backwards breaking changes for a migration guide.
We provide binaries via pip and conda. They require PyTorch 1.2.0 and newer. See https://pytorch.org/ for installation instructions.
We would like to thank our contributors and the wider community for their significant contributions to this release. We are happy to see an active community around torchaudio and are eager to further grow and support it.
In particular we'd like to thank @keunwoochoi, @ksanjeevan, and all the other maintainers and contributors of torchaudio-contrib for their significant and valuable additions around standardization and the support of complex numbers (https://github.com/pytorch/audio/pull/131, https://github.com/pytorch/audio/issues/110, https://github.com/keunwoochoi/torchaudio-contrib/issues/61, https://github.com/keunwoochoi/torchaudio-contrib/issues/36).
An implementation of basic transforms with a Kaldi-like interface.
We added the functions spectrogram, fbank, and resample_waveform (https://github.com/pytorch/audio/pull/119, https://github.com/pytorch/audio/pull/127, and https://github.com/pytorch/audio/pull/134). For more details see the documentation on torchaudio.compliance.kaldi which mirrors the arguments and outputs of Kaldi features.
As an example we can look at the sinc interpolation resampling similar to Kaldi’s implementation. In the figure below, the blue dots are the original signal and red dots are the downsampled signal with half the original frequency. The red dot elements are approximately every other original element.
specgram = torchaudio.compliance.kaldi.spectrogram(waveform, frame_length=...)
fbank = torchaudio.compliance.kaldi.fbank(waveform, num_mel_bins=...)
resampled_waveform = torchaudio.compliance.kaldi.resample_waveform(waveform, orig_freq=...)
Constructing a signal from a spectrogram can be used in applications like source separation or to generate audio signals to listen to. More specifically torchaudio.functional.istft is the inverse of torch.stft. It has the same parameters (+ additional optional parameter of length) and returns the least squares estimation of an original signal.
torch.manual_seed(0)
n_fft = 5
waveform = torch.rand(2, 5)
stft = torch.stft(waveform, n_fft=n_fft)
approx_waveform = torchaudio.functional.istft(stft, n_fft=n_fft, length=waveform.size(1))
>>> waveform
tensor([[0.4963, 0.7682, 0.0885, 0.1320, 0.3074],
[0.6341, 0.4901, 0.8964, 0.4556, 0.6323]])
>>> approx_waveform
tensor([[0.4963, 0.7682, 0.0885, 0.1320, 0.3074],
[0.6341, 0.4901, 0.8964, 0.4556, 0.6323]])
Compose:
Please use core abstractions such as nn.Sequential() or a for-loop over a list of transforms.SPECTROGRAM, F2M, and MEL have been removed. Please use Spectrogram, MelScale, and MelSpectrogramLC2CL and BLC2CBL): While the LC layout might be common in signal processing, support for it is out of scope of this library and transforms such as LC2CL only aid their proliferation. Please use transpose if you need this behavior.Scale, PadTrim, DownmixMono: Please use division in place of Scale torch.nn.functional.pad/trim in place of PadTrim , torch.mean on the channel dimension in place of DownmixMono.torchaudio.legacy has been removed. Please use torchaudio.load and torchaudio.saveSpectrogram used to be of dimension (channel, time, freq) and is now (channel, freq, time). Similarly for MelScale, MelSpectrogram, and MFCC, time is the last dimension. Please see our README for an explanation of the rationale behind these changes. Please use transpose to get the previous behavior.MuLawExpanding was renamed to MuLawDecoding as the inverse of MuLawEncoding ( https://github.com/pytorch/audio/pull/159)SpectrogramToDB was renamed to AmplitudeToDB ( https://github.com/pytorch/audio/pull/170). The input does not necessarily have to be a spectrogram and as such can be used in many more cases as the name should reflect.Spectrogram, AmplitudeToDB, MelScale, MelSpectrogram, MFCC, MuLawEncoding, and MuLawDecoding. (https://github.com/pytorch/audio/pull/118)Spectrogram, AmplitudeToDB, MelScale, MelSpectrogram, MFCC, MuLawEncoding, and MuLawDecoding (https://github.com/pytorch/audio/pull/118)test_transforms.py where double tensors were compared with floats (https://github.com/pytorch/audio/pull/132)vctk.read_audio (issue https://github.com/pytorch/audio/issues/143) as there were issues with downsampling using SoxEffectsChain (https://github.com/pytorch/audio/pull/145)sox_close (https://github.com/pytorch/audio/pull/174)Your coding agent can read these notes before it upgrades. Set up the MCP server →