NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #3572 most downloaded on pub.dev
OpenAI Whisper ASR (Automatic Speech Recognition) for Flutter
Last release 6 days ago
02 Oct 2026
Ships unpredictably
gaps range from 2 weeks to 11 months
Nearly every release is documented
notes for 22 of 22 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
22 releases · first in 2025
One column per month.
Fixed Android builds failing at checkDebugAarMetadata on newer toolchains (#30): the plugin compiled against compileSdk 34 while its own dependency ff
checkDebugAarMetadata on newer toolchains (#30): the plugin compiled against compileSdk 34 while its own dependency ffmpeg_kit_flutter_new_min 2.1.0 requires consumers to compile against 35 or later, and an app cannot override a plugin's compileSdk. The plugin now uses compileSdk 36, Flutter's own default. Thanks @d4kr for the reportminSdkVersion 21 → 24, README updated). ffmpeg-kit already required API 24, so apps below it never built; the plugin's declaration now says so instead of advertising API 21Opt-in Silero voice-activity detection for transcribe (#29): pass vadModelPath (controller and low-level API) to run whisper.cpp's built-in VAD before
transcribe (#29): pass vadModelPath (controller and low-level API) to run whisper.cpp's built-in VAD before decoding, so only detected speech is transcribed — the standard defence against whisper hallucinating or looping over the leading/trailing silence of push-to-talk recordings. vadSpeechPadMs optionally overrides the padding kept around speech. Off by default; implemented in all three native shims. Contributed by @Ranjan-Bhagatfailed to open VAD model … instead of a generic errorkeepModelLoaded: whisper.cpp keeps VAD state on the model context and only resets it when VAD runs again, so a parked model that has run VAD is only reused by requests with the same VAD model. Without that guard, a non-VAD request on such a context got its timestamps remapped through the previous request's speech map, and a request with a different VAD model silently used the first onekeepModelLoaded / releaseModel being silently ignored on macOS: the 2.6.0 resident-model support never reached the macOS shim, so macOS callers still reloaded the model on every request and live sessions never borrowed a parked model. The macOS shim now matches the iOS oneleaks --atExit, plus an end-to-end run of the macOS example app through the public Dart APIOpt-in resident model between transcriptions (#26): pass keepModelLoaded: true to transcribe (controller and low-level API) and the loaded whisper_con
keepModelLoaded: true to transcribe (controller and low-level API) and the loaded whisper_context is parked in the native layer instead of freed, so the next request with the same model file skips the multi-second load — the difference between usable and not for push-to-talk dictation. Free it with the new releaseModel() (controller and Whisper), or by transcribing once more with the flag off. Implemented on all five platforms (the Android/Windows/Linux shim and the iOS/macOS shim)no_context) before decoding, so a warm transcription is byte-identical to a cold one and one request's output can never condition (or leak a hallucination into) the next. initialPrompt still applies per requesttranscribeLive / startWhisperLiveSession accept keepModelLoaded, which parks the session's model on stop() so session-per-utterance callers skip the per-session load too. A session also borrows an already-parked model automatically when the path matches — and always returns it on stop, so a one-shot caller that parked a model never silently loses it to a live session. Cross-session independence needs nothing extra: every stream decode already runs with no_contextreleaseModel frees it (and is a safe no-op when nothing is parked), and 3 rounds of concurrent transcriptions pass leaks --atExit with zero leaked bytesLive transcription now works with locally provided model files (#25): WhisperController.transcribeLive accepts modelPath as an alternative to model, a
WhisperController.transcribeLive accepts modelPath as an alternative to model, and startWhisperLiveSession is now part of the public API — it takes a modelPath and leaves audio delivery to the caller via WhisperLiveSession.feedsdk: '>=3.7.0 <4.0.0' and flutter: '>=3.29.0'. Since 2.4.0, ffi ^2.1.4 has needed Dart 3.7, but the stale >=3.1.0 constraint made pub get fail with a version-solving error on older SDKs instead of a clear SDK-requirement message. Verified on the floor (Flutter 3.29.3 / Dart 3.7.2) and on current stable (3.44.8)Added transcription progress reporting (#6): transcribe (controller and low-level API) takes an optional onProgress callback invoked with whisper.cpp'
transcribe (controller and low-level API) takes an optional onProgress callback invoked with whisper.cpp's inference progress as a 0–100 percentage (coarse steps). Implemented via a NativeCallable.listener handed to whisper_full_params.progress_callback, so events arrive on the calling isolate without blocking inferenceWhisperController.transcribe now exposes withSegments and splitOnWord (#14): pass withSegments: true to get per-segment timestamps in result.transcrip
WhisperController.transcribe now exposes withSegments and splitOnWord (#14): pass withSegments: true to get per-segment timestamps in result.transcription.segments (fromTs/toTs as Duration), and add splitOnWord: true for one segment per word. Previously segments were only reachable through the low-level Whisper.transcribe APIdiarize parameter now actually does something (#13): it enables whisper.cpp's tinydiarize speaker-turn detection. Use the new WhisperModel.smallEnTdrz model (English only) together with diarize: true and withSegments: true — segments after which the speaker changes carry speakerTurnNext == true. With regular (non-tdrz) models the flag remains a no-op, as it silently was since 1.7.0Added Linux support (x64): the vendored whisper.cpp v1.9.1 builds into libwhisper_ggml.so through the standard Flutter Linux CMake toolchain — both on
libwhisper_ggml.so through the standard Flutter Linux CMake toolchain — both one-shot (transcribe) and live (transcribeLive) transcription worklibggml*.so binaries from linux/ — leftovers that never constituted a working implementation (no whisper code, wrong bundling variable) and only inflated the package-DWHISPER_GGML_AVX2=OFF for baseline SSE2) and uses an ffmpeg executable from PATH for non-WAV inputtranscribe called whisper_full with a null context instead of returning an error response (found by the new Linux test suite)transcribe (unreadable/wrong-format WAV, failed inference) returned without freeing the loaded model — ~150 MB leaked per failed call with the base modelAdded Windows support (x64): the vendored whisper.cpp v1.9.1 now also builds into whisper_ggml.dll through the standard Flutter Windows CMake toolchai
whisper_ggml.dll through the standard Flutter Windows CMake toolchain — both one-shot (transcribe) and live (transcribeLive) transcription work-DWHISPER_GGML_AVX2=OFF to the plugin CMake for a baseline SSE2 buildffmpeg executable from PATH when available (ffmpeg_kit has no Windows implementation); without it, input must already be 16 kHz mono WAV/O2) in Windows debug builds, matching the iOS/macOS behaviourUpgraded the vendored whisper.cpp from a 2023-era snapshot to v1.9.1 — roughly 15× faster transcription (an 11-second clip takes ~0.4 s with the base
base model on Apple Silicon), with Accelerate enabled on iOS/macOS-march=armv8.2-a+fp16 Android compile flags that caused SIGILL crashes on armv8.0 devices (#15)-O3 in debug builds on iOS/macOS, so development-time transcription is no longer unusably slowRemoved the windows and linux platform declarations — neither has a native whisper implementation, and pub.dev advertised support that failed at runti
windows and linux platform declarations — neither has a native whisper implementation, and pub.dev advertised support that failed at runtime. Unsupported platforms now throw a clear UnsupportedError (#20)channel-error on unrelated plugins like path_provider (#16)Fixed a native crash (SIGABRT) when whisper emits text containing invalid UTF-8 — e.g. a token boundary splitting a multi-byte character, common with
Added live (streaming) transcription: WhisperController.transcribeLive takes a stream of 16 kHz mono PCM16 audio and returns a WhisperLiveSession that
WhisperController.transcribeLive takes a stream of 16 kHz mono PCM16 audio and returns a WhisperLiveSession that emits progressively refined partial transcripts while the user speaks; stop() returns the final textstream_start / stream_feed / stream_stop API on Android, iOS, and macOS); inference runs on a dedicated isolate and never blocks the UI[BLANK_AUDIO] markers and hallucinated repetition; tunable per session via gateRmsMin, gateVoiceRatio, and gateNoiseFloorCapsuppressNonSpeechTokens parameter (whisper_full_params.suppress_non_speech_tokens) to TranscribeRequest, transcribe, and transcribeLiveAdded initialPrompt parameter to TranscribeRequest and WhisperController.transcribe
initialPrompt parameter to TranscribeRequest and WhisperController.transcribeinitial_prompt through to whisper_full_params.initial_prompt on Android, iOS, and macOS to bias decoding toward domain-specific vocabulary, names, and punctuationnoContext parameter (whisper_full_params.no_context, equivalent to Python whisper's condition_on_previous_text=False) on Android, iOS, and macOS to disable cross-segment text conditioning — helps against hallucinated repetition on short utterancesnoContext: false leave whisper.cpp defaults, so existing callers see no behaviour changeflutter_riverpod dependency, which was constraining consumers to riverpod 2.x even though the package never imported itConnected diarize transcribe parameter to the underlying whisper C++ code
diarize transcribe parameter to the underlying whisper C++ codediarize parameter to the transcribe methodAdded auto language support for iOS
auto language support for iOSexample projectSwitched main FFmpeg from heavy ffmpeg_kit_flutter_new: ^1.6.1 to lightweight ffmpeg_kit_flutter_new_min: ^2.1.0
ffmpeg_kit_flutter_new: ^1.6.1 to lightweight ffmpeg_kit_flutter_new_min: ^2.1.0recorder dependency for example project from v5.2.1 to v6.0.0Added ability to use "auto" language detection
pubspec.yaml dependenciesUpgraded Android bindings to work with Flutter 3.29
Fixed Android v1 embedding issue by adding override for ffmpeg_kit_flutter_full_gpl
* Cleaned up code
* Added support for MacOS
Added support for Android and iOS
Your coding agent can read these notes before it upgrades. Set up the MCP server →