NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2116 most downloaded on PyPI
A powerful framework for building realtime voice AI agents
Last release 3 days ago
01 Oct 2026
Ships fairly regularly
a new release about every 8 days
Nearly every release is documented
notes for 55 of the last 60 stable releases
4 versions withdrawn
withdrawn after publishing
3 years old
222 releases · first in 2023
feat(smallestai): add TTS continuations protocol support by @harshitajain165 in #7111
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.8.3...livekit-agents@1.8.4
One column per month.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.8.2...livekit-agents@1.8.3
Run loop blocking detection. The agent now detects code that blocks the asyncio event loop and surfaces the offending call in Agent Insights, so you c
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.8.1...livekit-agents@1.8.2
Speech models that are able to speak and listen simultaneously are now supported via the new DuplexModel class, with OpenAI's GPT-Live model as the fi
Speech models that are able to speak and listen simultaneously are now supported via the new DuplexModel class, with OpenAI's GPT-Live model as the first to be implemented.
session = AgentSession(
llm=GPTLiveModel(
voice="marin",
# backend Responses model that handles reasoning and tools
responses_options={
"model": "gpt-5.6-luna",
"instructions": "Use tools when current information is required.",
},
),
)Read more about duplex models in our docs.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.8.0...livekit-agents@1.8.1
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Spans, metrics, and logs now follow the OTel GenAI semantic conventions, so
Langfuse, Datadog Agent Observability, and other GenAI-aware backends read
LiveKit traces natively. PII filtering also moved in-process.
Migration Notes
| Before | After |
|---|---|
event gen_ai.system.message |
attribute gen_ai.system_instructions |
events gen_ai.user.message, gen_ai.assistant.message, gen_ai.tool.message |
attribute gen_ai.input.messages |
event gen_ai.choice (role/content/tool_calls) |
attributes gen_ai.output.messages + gen_ai.response.finish_reasons |
realtime gen_ai.* on the agent_turn span |
on the new realtime_inference child span |
realtime gen_ai.operation.name = "chat" |
= "generate_content" |
gen_ai.provider.name = "api.openai.com", "AWS Bedrock", "google" |
"openai", "aws.bedrock", "gcp.gen_ai" |
custom llm_node always tagged with the configured model/provider |
tagged only when that LLM served the call; otherwise the node span records the convention itself |
| PII stripped at LiveKit Cloud's collector | stripped in-process; under project redaction content never reaches any exporter, Cloud included |
telemetry._chat_ctx_to_otel_events(chat_ctx) |
telemetry.gen_ai.to_system_instructions() / to_input_messages() / to_output_messages() |
telemetry.utils._redaction_enabled() |
telemetry.utils.redaction_enabled(span_attributes=None) |
Breaking
gen_ai.system.message,gen_ai.user.message, gen_ai.assistant.message, gen_ai.tool.message andgen_ai.choice are gone, replaced by the attributes in the table above (JSON,agent_turn onto a new realtime_inference child span,agent_turn return nothing. Itsgen_ai.operation.name also changed from chat to generate_content.gen_ai.provider.name is normalized to the registry spelling. Providers outside thellm_node is no longer credited with the configured model and providergen_ai.usage.input_cached_tokens is emitted again on the pipeline path (reversinggen_ai.usage.reasoning.output_tokensgen_ai.usage.reasoning_tokens. A backend summing both spellings double-counts.telemetry._chat_ctx_to_otel_events() removed; telemetry.utils._redaction_enabled()redaction_enabled(span_attributes=None).Added
gen_ai.request.stream, gen_ai.output.type,gen_ai.conversation.id, gen_ai.response.{id,model,finish_reasons,time_to_first_chunk}.gen_ai.usage.cache_write.input_tokens, .reasoning.output_tokens.{text,audio,image}.* counts.execute_tool spans carry gen_ai.tool.{name,type,call.id,call.arguments,call.result,description};gen_ai.agent.name, sessions gen_ai.workflow.name.error.type on failures, set even when redaction is on.gen_ai.client.token.usage, gen_ai.client.operation.duration,gen_ai.client.operation.time_to_first_chunk, gen_ai.invoke_agent.duration,gen_ai.execute_tool.duration, alongside the existing lk.agents.*.OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=0 (ortelemetry.gen_ai.set_capture_content(False)) omits message payloads, toolset_tracer_provider(..., allow_pii=False) / LIVEKIT_TELEMETRY_ALLOW_PII=0 stripsFull Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.7.1...livekit-agents@1.8.0
fix(voice): agent and user state while a tool runs by @longcw in #6937
livekit to 1.1.15 by @1egoman in #6970Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.7.0...livekit-agents@1.7.1
Introducing PII Redaction Support
This release adds PII redaction support for Agent Observability. It semantically redacts detected entities from chat history and audio recordings. It also filters sensitive data from logs and traces while retaining useful diagnostic context.
To support this feature, we renamed sensitive trace attributes and log fields emitted by the framework. If your third-party observability queries depend on these names, update them when you upgrade.
Expressive mode allows voice agents to speak with natural prosody and emotion. Speech delivery is determined by emotion tags generated from the context of the conversation. Enable expressive mode in your AgentSession:
from livekit.agents import AgentSession, inference
session = AgentSession(
stt=inference.STT(model="deepgram/nova-3", language="multi"),
llm=inference.LLM(model="google/gemma-4-31b-it"),
tts=inference.TTS(model="fishaudio/s2.1-pro", voice="b347db033a6549378b48d00acb0d06cd"),
expressive=True,
# ... vad, turn_detection
)See the documentation and the blog post for more information.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.10...livekit-agents@1.7.0
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
fix(cartesia): redact API keys from websocket handshake errors by @LHMQ878 in #6740
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.9...livekit-agents@1.6.10
Add support for Deepgram's Flux TTS API (v2/speak) integration by @dg-edcharbeneau in https://github.com/livekit/agents/pull/6511
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.8...livekit-agents@1.6.9
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.8...livekit-agents@1.6.9
> `console` and `dev` modes are deprecated in favor of lk agent (#6656). Existing workflows that shell out to python agent.py console or python agent.…
[!WARNING]
consoleanddevmodes are deprecated in favor oflk agent(#6656). Existing workflows that shell out topython agent.py consoleorpython agent.py devshould migrate to thelk agentCLI.
[!IMPORTANT] Inference Ink-2 speech onset fix (#6630) — if you rely on the STT model for interruption detection without a local VAD, upgrade for this fix. Previously the provider's speech-onset signal was delayed until first non-empty interim transcript, so STT interruptions could be delayed on Ink-2 via Inference. Setups with a local VAD in the pipeline are unaffected.
lk agent by @u9g in https://github.com/livekit/agents/pull/6656Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.7...livekit-agents@1.6.8
Warning
console and dev modes are deprecated in favor of lk agent (#6656). Existing workflows that shell out to python agent.py console or python agent.py dev should migrate to the lk agent CLI.
Important
Inference Ink-2 speech onset fix (#6630) — if you rely on the STT model for interruption detection without a local VAD, upgrade for this fix. Previously the provider's speech-onset signal was delayed until first non-empty interim transcript, so STT interruptions could be delayed on Ink-2 via Inference. Setups with a local VAD in the pipeline are unaffected.
lk agent by @u9g in #6656Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.7...livekit-agents@1.6.8
fix(core): reset user away timer on final STT transcript AGT-3149 by @chenghao-mou in https://github.com/livekit/agents/pull/6478
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.6...livekit-agents@1.6.7
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.6...livekit-agents@1.6.7
elevenlabs: deprecate STT model_id in favor of model by @u9g in #6378
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.5...livekit-agents@1.6.6
fix(assemblyai): respect mode preset for turn-silence; remap deprecated u3-pro by @dlange-aai in https://github.com/livekit/agents/pull/6220
mode preset for turn-silence; remap deprecated u3-pro by @dlange-aai in https://github.com/livekit/agents/pull/6220remove() method to ChatContext for flexible item removal by @3eid in https://github.com/livekit/agents/pull/5131Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.4...livekit-agents@1.6.5
mode preset for turn-silence; remap deprecated u3-pro by @dlange-aai in #6220remove() method to ChatContext for flexible item removal by @3eid in #5131Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.4...livekit-agents@1.6.5
disable retry for eot errors by @chenghao-mou in #6195
Turn Detector
Turn handling/committing
Warning
users are advised to upgrade if you use agent handoff with STT and are on 1.5.14 ~ 1.6.3
Core
Plugins
language_code doc to remove error case by @inickt in #6197Version update
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.3...livekit-agents@1.6.4
fix(slng): expose speed in update_options by @adityajha2005 in #6175
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.2...livekit-agents@1.6.3
feat(assemblyai): add universal-3-5-pro (now default) and Voice Focus streaming params by @dlange-aai in https://github.com/livekit/agents/pull/6119
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.1...livekit-agents@1.6.2
Determining when the agent should respond is a delicate balance: responding too early can interrupt the user, while responding too late introduces unn
Determining when the agent should respond is a delicate balance: responding too early can interrupt the user, while responding too late introduces unnecessary latency and awkward silence. We solve this with our state-of-the-art turn detector, which leverages both audio and text semantics to identify the optimal moment to respond.
Read more about our turn detector in our blog.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.6.0...livekit-agents@1.6.1
When a long-running tool is in progress, the user will be met with silence until it completes. Asynchronous tools can hand control back to the LLM bef
When a long-running tool is in progress, the user will be met with silence until it completes. Asynchronous tools can hand control back to the LLM before it finishes, streaming updates into the conversation as it progresses.
class TravelAgent(Agent):
@llm.function_tool(flags=llm.ToolFlag.CANCELLABLE, on_duplicate="confirm")
async def book_flight(self, ctx: RunContext, origin: str, destination: str, date: str) -> str:
"""Book a flight."""
# First update is delivered immediately and releases control back to the LLM,
# so the agent can say e.g. "Sure, searching flights to Tokyo — this'll take a minute."
await ctx.update(
f"Searching flights from {origin} to {destination} on {date}. "
"This will take a couple of minutes."
)
await asyncio.sleep(30) # simulate long-running work
# Later updates are coalesced into a deferred reply, delivered when the agent
# is idle: "Good news — best price is $289 on Delta, confirming now."
await ctx.update("Found 3 options. Best price $289 on Delta. Confirming now.")
await asyncio.sleep(40) # simulate more long-running work
# The final return is also delivered when the agent is idle.
return f"Booked! Confirmation FL-{random.randint(100000, 999999)}."
Use filler words to break long silences via ctx.with_filler(). You can set delay to the amount of quiet seconds per interval for when the filler words are played. You can also rotate between phrases, set interval to configure the cooldown between filler fires.
followups = [
"Almost there, just confirming.",
"Still working on it, won't be long.",
"Hang tight — almost done.",
]
async with ctx.with_filler(
lambda step: followups[step], delay=5, interval=10, max_steps=len(followups)
):
await asyncio.sleep(40) # simulate a long API call
Read more about async tools in our docs.
pytest --unit/--plugin openai/ … flags by @Bobronium in https://github.com/livekit/agents/pull/5945Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.17...livekit-agents@1.6.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
feat(gradium STT): add language option to gradium.STT by @FLoppix in https://github.com/livekit/agents/pull/5878
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.15...livekit-agents@1.5.17
Nothing published for this version
(openai realtime): add status_details to incomplete response logs by @tinalenguyen in https://github.com/livekit/agents/pull/5873
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.14...livekit-agents@1.5.15
Update download-files deprecation message by @bcherry in https://github.com/livekit/agents/pull/5781
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.12...livekit-agents@1.5.14
Nothing published for this version
deprecate mcp_servers param on Agent and AgentSession by @longcw in https://github.com/livekit/agents/pull/5667
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.10...livekit-agents@1.5.12
Nothing published for this version
fix: surface Deepgram TTS websocket errors by @nightcityblade in https://github.com/livekit/agents/pull/5728
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.9...livekit-agents@1.5.10
An outbound call can reach a person, voicemail, an IVR menu, or a number that can't accept messages. Answering machine detection (AMD) listens to the
An outbound call can reach a person, voicemail, an IVR menu, or a number that can't accept messages. Answering machine detection (AMD) listens to the start of the call, classifies it with an LLM, and returns a result so your agent can respond appropriately.
Read more about using AMD in our blog.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.8...livekit-agents@1.5.9
feat(interruption): barge-in cooldown window for corrections by @chenghao-mou in https://github.com/livekit/agents/pull/5269
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.7...livekit-agents@1.5.8
fix(openai): forward session.update on RealtimeModel.update_options by @longcw in https://github.com/livekit/agents/pull/5531
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.6...livekit-agents@1.5.7
Add Qwen 3 TTS support for Simplismart-livekit plugin by @simplipratik in https://github.com/livekit/agents/pull/5474
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.5...livekit-agents@1.5.6
(hedra): note deprecation in readme by @tinalenguyen in https://github.com/livekit/agents/pull/5475
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.4...livekit-agents@1.5.5
Refines default behavior for preemptive generation to better handle long or intermittent user speech, reducing unnecessary downstream inference and as
Refines default behavior for preemptive generation to better handle long or intermittent user speech, reducing unnecessary downstream inference and associated cost increases.
Also introduces PreemptiveGenerationOptions for developers who need fine-grained control over this behavior.
https://github.com/livekit/agents/blob/78a66bcf79c5cea82989401c408f1dff4b961a5b/livekit-agents/livekit/agents/voice/turn.py#L115
class PreemptiveGenerationOptions(TypedDict, total=False):
"""Configuration for preemptive generation."""
enabled: bool
"""Whether preemptive generation is enabled. Defaults to ``True``."""
preemptive_tts: bool
"""Whether to also run TTS preemptively before the turn is confirmed.
When ``False`` (default), only LLM runs preemptively; TTS starts once the
turn is confirmed and the speech is scheduled."""
max_speech_duration: float
"""Maximum user speech duration (s) for which preemptive generation
is attempted. Beyond this threshold, preemptive generation is skipped
since long utterances are more likely to change and users may expect
slower responses. Defaults to ``10.0``."""
max_retries: int
"""Maximum number of preemptive generation attempts per user turn.
The counter resets when the turn completes. Defaults to ``3``."""
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.3...livekit-agents@1.5.4
> livekit-agents 1.5 introduced many new features. You can check out the changelog [here](https://github.com/livekit/agents/releases/tag/livekit-agent
[!NOTE]
livekit-agents 1.5 introduced many new features. You can check out the changelog here.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.2...livekit-agents@1.5.3
fix(utils): preserve type annotations in deprecate_params by @longcw in https://github.com/livekit/agents/pull/5200
[!NOTE]
livekit-agents 1.5 introduced many new features. You can check out the changelog here.
generate_reply timeout to 10 seconds by @qionghuang6 in https://github.com/livekit/agents/pull/5205min_words_to_interrupt to Phonic plugin options by @qionghuang6 in https://github.com/livekit/agents/pull/5304Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.1...livekit-agents@1.5.2
> livekit-agents 1.5 introduced many new features. You can check out the changelog [here](https://github.com/livekit/agents/releases/tag/livekit-agent
[!NOTE]
livekit-agents 1.5 introduced many new features. You can check out the changelog here.
Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.5.0...livekit-agents@1.5.1
Old keyword arguments (min_endpointing_delay, allow_interruptions, etc.) still work but are deprecated and will emit warnings.
Preemptive generation starts LLM and TTS inference before the end of a user’s turn is detected, reducing overall latency.
To disable it:
session = AgentSession(preemptive_generation=False)
The headline feature of v1.5.0: an audio-based ML model that distinguishes genuine user interruptions from incidental sounds like backchannels ("mm-hmm"), coughs, sighs, or background noise. Enabled by default — no configuration needed.
Key stats:
When a false interruption is detected, the agent automatically resumes playback from where it left off — no re-generation needed.
To opt out and use VAD-only interruption:
session = AgentSession(
...
turn_handling=TurnHandlingOptions(
interruption={
"mode": "vad",
},
),
)
Blog post: https://livekit.com/blog/adaptive-interruption-handling
Endpointing delays now adapt to each conversation's natural rhythm. Instead of a fixed silence threshold, the agent uses an exponential moving average of pause durations to dynamically adjust when it considers the user's turn complete.
session = AgentSession(
...
turn_handling=TurnHandlingOptions(
endpointing={
"mode": "dynamic",
"min_delay": 0.3,
"max_delay": 3.0,
},
),
)
TurnHandlingOptions APIEndpointing and interruption settings are now consolidated into a single TurnHandlingOptions dict passed to AgentSession. Old keyword arguments (min_endpointing_delay, allow_interruptions, etc.) still work but are deprecated and will emit warnings.
session = AgentSession(
turn_handling={
"turn_detection": "vad",
"endpointing": {"min_delay": 0.5, "max_delay": 3.0},
"interruption": {"enabled": True, "mode": "adaptive"},
},
)
New SessionUsageUpdatedEvent provides structured, per-model usage data — token counts, character counts, and audio durations — broken down by provider and model:
@session.on("session_usage_updated")
def on_usage(ev: SessionUsageUpdatedEvent):
for usage in ev.usage.model_usage:
print(f"{usage.provider}/{usage.model}: {usage}")
Usage types: LLMModelUsage, TTSModelUsage, STTModelUsage, InterruptionModelUsage.
You can also access aggregated usage at any time via the session.usage property:
usage = session.usage
for model_usage in usage.model_usage:
print(model_usage)
Usage data is also included in SessionReport (via model_usage), so it's available in post-session telemetry and reporting out of the box.
ChatMessage.metricsEach ChatMessage now carries a metrics field (MetricsReport) with per-turn latency data:
transcription_delay — time to obtain transcript after end of speechend_of_turn_delay — time between end of speech and turn decisionon_user_turn_completed_delay — time in the developer callbackContext summarization now includes function calls and their outputs when building summaries, preserving tool-use context across the conversation window.
Set the agent log level via LIVEKIT_LOG_LEVEL environment variable or through ServerOptions, without touching your code.
| Deprecated | Replacement | Notes |
|---|---|---|
metrics_collected event |
session_usage_updated event + ChatMessage.metrics |
Usage/cost data moves to session_usage_updated; per-turn latency moves to ChatMessage.metrics. Old listeners still work with a deprecation warning. |
UsageCollector |
ModelUsageCollector |
New collector supports per-model/provider breakdown |
UsageSummary |
LLMModelUsage, TTSModelUsage, STTModelUsage |
Typed per-service usage classes |
RealtimeModelBeta |
RealtimeModel |
Beta API removed |
AgentFalseInterruptionEvent.message / .extra_instructions |
Automatic resume via adaptive interruption | Accessing these fields logs a deprecation warning |
AgentSession kwargs: min_endpointing_delay, max_endpointing_delay, allow_interruptions, discard_audio_if_uninterruptible, min_interruption_duration, min_interruption_words, turn_detection, false_interruption_timeout, resume_false_interruption |
turn_handling=TurnHandlingOptions(...) |
Old kwargs still work but emit deprecation warnings. Will be removed in v2.0. |
Agent / AgentTask kwargs: turn_detection, min_endpointing_delay, max_endpointing_delay, allow_interruptions |
turn_handling=TurnHandlingOptions(...) |
Same migration path as AgentSession. Will be removed in future versions. |
generate_reply to resolve with the current GenerationCreatedEvent by @qionghuang6 in https://github.com/livekit/agents/pull/5147Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.4.6...livekit-agents@1.5.0
Nothing published for this version
Nothing published for this version
(google realtime): replace deprecated mediaChunks by @tinalenguyen in https://github.com/livekit/agents/pull/5089
null in enum array for nullable enum schemas by @MSameerAbbas in https://github.com/livekit/agents/pull/5080required field in tool schema when function has no parameters by @longcw in https://github.com/livekit/agents/pull/5082Full Changelog: https://github.com/livekit/agents/compare/livekit-agents@1.4.5...livekit-agents@1.4.6
Your coding agent can read these notes before it upgrades. Set up the MCP server →