NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
NuGet · #3058 most downloaded on NuGet
This NuGet package provides native and managed libraries for the Microsoft Cognitive Services Speech SDK. Documentation is available at https://aka.ms/csspeech.
Last release today
28 Sep 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 56 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
68 releases · first in 2018
The Beta version of Python support (3.5 and above) is available with this release. For more information, see here](../../quickstart-python.md).
Python
The Beta version of Python support (3.5 and above) is available with this release. For more information, see here](../../quickstart-python.md).
JavaScript
The Speech SDK for JavaScript has been open-sourced. The source code is available on GitHub .
We now support Node.js, more info can be found here .
The length restriction for audio sessions has been removed, reconnection will happen automatically under the cover.
Connection object
From the Recognizer , you can access a Connection object. This object allows you to explicitly initiate the service connection and subscribe to connect and disconnect events. (This feature isn't yet available from JavaScript and Python.)
Support for Ubuntu 18.04.
Android
Enabled ProGuard support during APK generation.
Improvements in the internal thread usage, reducing the number of threads, locks, mutexes.
Improved error reporting / information. In several cases, error messages haven't been propagated out all the way out.
Updated development dependencies in JavaScript to use up-to-date modules.
Fixed memory leaks due to a type mismatch in RecognizeAsync .
In some cases exceptions were being leaked.
Fixing memory leak in translation event arguments.
Fixed a locking issue on reconnect in long running sessions.
Fixed an issue that could lead to missing final result for failed translations.
C#: If an async operation wasn't awaited in the main thread, it was possible the recognizer could be disposed before the async task was completed.
Java: Fixed a problem resulting in a crash of the Java VM.
Objective-C: Fixed enum mapping; RecognizedIntent was returned instead of RecognizingIntent .
JavaScript: Set default output format to 'simple' in SpeechConfig .
JavaScript: Removing inconsistency between properties on the config object in JavaScript and other languages.
Updated and fixed several samples (for example output voices for translation, etc.).
Added Node.js samples in the sample repository .
One column per quarter.
Proxy Support: In the SpeechConfig object, you can now call a function to set the proxy information (hostname, port, username, and password). This fea
Support for Android x86/x64.
Proxy Support: In the SpeechConfig object, you can now call a function to set the proxy information (hostname, port, username, and password). This feature isn't yet available on iOS.
Improved error code and messages. If a recognition returned an error, this did already set Reason (in canceled event) or CancellationDetails (in recognition result) to Error . The canceled event now contains two additional members, ErrorCode and ErrorDetails . If the server returned additional error information with the reported error, it will now be available in the new members.
Added additional verification in the recognizer configuration, and added additional error message.
Improved handling of long-time silence in middle of an audio file.
NuGet package: for .NET Framework projects, it prevents building with AnyCPU configuration.
Fixed several exceptions found in recognizers. In addition, exceptions are caught and converted into Canceled event.
Fix a memory leak in property management.
Fixed bug in which an audio input file could crash the recognizer.
Fixed a bug where events could be received after a session stop event.
Fixed some race conditions in threading.
Fixed an iOS compatibility issue that could result in a crash.
Stability improvements for Android microphone support.
Fixed a bug where a recognizer in JavaScript would ignore the recognition language.
Fixed a bug preventing setting the EndpointId (in some cases) in JavaScript.
Changed parameter order in AddIntent in JavaScript, and added missing AddIntent JavaScript signature.
Reliability improvements and bug fixes:
Reliability improvements and bug fixes:
Fixed potential fatal error due to race condition in disposing recognizer
Fixed potential fatal error when unset properties occur.
Added additional error and parameter checking.
Objective-C: Fixed possible fatal error caused by name overriding in NSString.
Objective-C: Adjusted visibility of API
JavaScript: Fixed regarding events and their payloads.
Documentation improvements.
In our sample repository , a new sample for JavaScript was added.
With this release, a number of breaking changes are introduced. Check this page for details.
Support for Objective-C on iOS. Check out our Objective-C quickstart for iOS .
Support for JavaScript in browser. Check out our JavaScript quickstart .
Breaking changes
UWP apps built with the Speech SDK now can pass the Windows App Certification Kit (WACK). Check out the UWP quickstart .
UWP apps built with the Speech SDK now can pass the Windows App Certification Kit (WACK). Check out the UWP quickstart .
Support for .NET Standard 2.0 on Linux (Ubuntu 16.04 x64).
Experimental: Support Java 8 on Windows (64-bit) and Linux (Ubuntu 16.04 x64). Check out the Java Runtime Environment quickstart .
Functional change
Breaking changes
On Java (Android), the SpeechFactory.configureNativePlatformBindingWithDefaultCertificate function no longer requires a path parameter. Now the path is automatically detected on all supported platforms.
The get-accessor of the property EndpointUrl in Java and C# was removed.
In Java, the audio synthesis result on the translation recognizer is implemented now.
Fixed a bug that could cause inactive threads and an increased number of open and unused sockets.
Fixed a problem, where a long-running recognition could terminate in the middle of the transmission.
Fixed a race condition in recognizer shutdown.
Support Android platform (API 23: Android 6.0 Marshmallow or higher). Check out the Android quickstart .
Support Android platform (API 23: Android 6.0 Marshmallow or higher). Check out the Android quickstart .
Support .NET Standard 2.0 on Windows. Check out the .NET Core quickstart .
Experimental: Support UWP on Windows (version 1709 or later).
Check out the UWP quickstart .
Note that UWP apps built with the Speech SDK don't yet pass the Windows App Certification Kit (WACK).
Support long-running recognition with automatic reconnection.
Functional changes
StartContinuousRecognitionAsync() supports long-running recognition.
The recognition result contains more fields. They're offset from the audio beginning and duration (both in ticks) of the recognized text and additional values that represent recognition status, for example, InitialSilenceTimeout and InitialBabbleTimeout .
Support AuthorizationToken for creating factory instances.
Breaking changes
Recognition events: NoMatch event type was merged into the Error event.
SpeechOutputFormat in C# was renamed to OutputFormat to stay aligned with C++.
The return type of some methods of the AudioInputStream interface changed slightly:
In Java, the read method now returns long instead of int .
In C#, the Read method now returns uint instead of int .
In C++, the Read and GetFormat methods now return size_t instead of int .
C++: Instances of audio input streams now can be passed only as a shared_ptr .
Fixed incorrect return values in the result when RecognizeAsync() times out.
The dependency on media foundation libraries on Windows was removed. The SDK now uses Core Audio APIs.
Documentation fix: Added a regions page to describe the supported regions.
A recognizer now can consume a stream as the audio source. For more information, see the related how-to guide .
Functional changes
AudioInputStream
A recognizer now can consume a stream as the audio source. For more information, see the related how-to guide .
Detailed output format
When you create a SpeechRecognizer , you can request Detailed or Simple output format. The DetailedSpeechRecognitionResult contains a confidence score, recognized text, raw lexical form, normalized form, and normalized form with masked profanity.
Fixed a possible callback issue in the USP layer during shutdown.
If a recognizer consumed an audio input file, it was holding on to the file handle longer than necessary.
Removed several deadlocks between the message pump and the recognizer.
Fire a NoMatch result when the response from service is timed out.
The media foundation libraries on Windows are delay loaded. This library is required for microphone input only.
The upload speed for audio data is limited to about twice the original audio speed.
On Windows, C# .NET assemblies now are strong named.
Documentation fix: Region is required information to create a recognizer.
More samples have been added and are constantly being updated. For the latest set of samples, see the Speech SDK samples GitHub repository .
This release is the first public preview release of the Azure Speech SDK.
This release is the first public preview release of the Azure Speech SDK.
Your coding agent can read these notes before it upgrades. Set up the MCP server →