NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #2954 by repository stars
Last release 16 days ago
22 Sep 2026
Ships fairly regularly
a new release about every 5 weeks
Rarely documented
notes for 5 of 24 stable releases
Nothing withdrawn
no release was ever pulled
1 years old
34 releases · first in 2025
One column per month.
Nothing published for this version
Nothing published for this version
Olla is a high-performance proxy and load balancer for LLM infrastructure.
Olla is a high-performance proxy and load balancer for LLM infrastructure.
# Docker
docker pull ghcr.io/thushan/olla:v0.0.29
# Binary (see assets below)
./olla --config config.yaml
/internal/ui/ - Overview (fleet status, success rate, latency, live requests-per-second sparkline), Endpoints (per-endpoint health, priority, latency, model count) and Models (inventory grouped by family, with hosting endpoints). Built in Svelte 5 + TailwindCSS, polled (not pushed) against the existing /internal/status* JSON, and served from the same listener as the proxy - no second port (#205).dashboard.access_policy config block (allowed_cidrs + allowed_hosts), loopback-only by default. The published Docker image ships with allowed_cidrs pre-widened to the RFC1918 ranges so docker run -p 40114:40114 ghcr.io/thushan/olla:latest gives you a working dashboard with no config mount - which also means the container is reachable from anyone else on the same LAN through that published port (#213). See the Admin Dashboard docs for the full security model.make build-web (e.g. plain go build or go install) now logs a clear startup warning and serves 503 at /internal/ui/ instead of a silent placeholder (#213).GET /internal/metrics in Prometheus text format, built from the same data as /internal/status and /internal/stats/models - no external exporter needed for core proxy monitoring. Thanks to @Puupuls for the contribution (#188)./internal/status, /internal/status/endpoints and /internal/status/models now emit a weak ETag, so polling clients (including the new dashboard) can use If-None-Match and get 304s instead of re-fetching the full payload (#205, #213).sticky_outcome, routing strategy/action/reason, provider_model, translator fallback reason), not just response headers. Addresses issue #178 (#182).Info again in the access log; a prior regression had silently demoted them to Debug (#211).model_aliases entry silently proxy to the wrong backend: a request for an alias whose target model exists on no endpoint now fails fast with 404/503 and routing_action: rejected, matching how unknown models are already handled. Reported by @skaravos (#197, issue #191).config/models.yaml now actually parses. It used \d inside double-quoted YAML scalars, which is invalid YAML, so the file has never loaded on any install - the failure was swallowed and Olla silently fell back to embedded defaults, discarding any customisation with no diagnostic. Reported by @billford (#206, issue #204).--validate-config flag checks configuration and provider profiles without starting the server, with a pass/warn/fail report and exit codes (#210).logging.level in config is now actually applied to the runtime logger, instead of being parsed and ignored (#209).model_extraction.family_aliases and special_rules.preserve_family into the family-extraction pipeline - both were parsed but previously had no effect. Fixes Kimi-K2 being misclassified under deepseek (shared GGUF architecture lineage) (#208).user:pass@host credentials now fails startup instead of loading silently and leaking the credentials into every status/dashboard JSON response. The boot error rewrites the URL into a ready-to-paste auth: block with placeholders, so the fix doesn't require re-typing the real credentials (#213).id and url fields on /internal/status and /internal/status/endpoints are now derived from a sanitised URL (userinfo, query and fragment stripped) rather than echoing the raw configured URL (#213)./internal/status no longer reports "status": "critical" on a healthy fresh boot with no traffic yet. A new system.has_traffic boolean lets clients branch on the no-traffic state directly instead of parsing the success_rate string (#213)..dockerignore for a smaller, more predictable build context (#213).- url: "http://user:pass@host:8080" now fails startup. Move credentials into an auth: block:
- url: "http://host:8080"
auth:
type: basic
username: user
password: passid/url fields: now derived from a sanitised URL rather than the raw config value. Bookmarked dashboard deep-links to a specific endpoint will change once after upgrading. Clients comparing url for identity should switch to id."critical". Check system.has_traffic if you branch on traffic state./olla/proxy/ with zero healthy endpoints now returns 503 instead of 502. Update any monitor keyed on that status code.owned_by in converted model listings: the shared organisation-extraction logic used by the vLLM, vLLM-MLX, SGLang, llama.cpp, LMDeploy and Lemonade converters now also matches a hyphenated leading segment (e.g. Qwen2.5-7B → owned_by: "Qwen2.5") rather than falling through to a generic default. Treat owned_by as a best-effort label, not a canonical identifier.Thanks to everyone who contributed to this release:
GET /internal/metrics Prometheus endpoint (#188).config/models.yaml parse failure (issue #204).Full Changelog: v0.0.28...v0.0.29
Documentation: thushan.github.io/olla | Issues: github.com/thushan/olla/issues
An endpoint URL with embedded user:pass@host is rejected at config load. Previously
such a URL loaded silently and the credentials flowed into every status/dashboard JSON
surface as the literal URL string. The endorsed credential path is the auth: block,
which is held as json:"-" fields on the internal endpoint record and never reaches the
status layer.
Before:
- url: "http://user:pass@host:8080"
name: "protected-backend"
type: "openai-compatible"
After:
- url: "http://host:8080"
name: "protected-backend"
type: "openai-compatible"
auth:
type: basic
username: user
password: pass
For secrets kept out of config, use the username_file / password_file variants (see
Endpoint Authentication).
The boot error rewrites the rejected URL into an auth: block with the credentials
replaced by placeholders, so the fix is a copy-paste away without the error ever echoing
your real credentials. For a full user:pass@host URL this is an exactly equivalent
basic block; for a username-only user@host URL there is no equivalent basic
configuration (basic auth requires a password), so the error offers migration
alternatives instead - a basic block with a password placeholder, or a bearer block if
the username was really a token.
/internal/status no longer reports "status": "critical" on a fresh boot with all
endpoints healthy. With no proxy traffic, the system status derives from endpoint health
alone, success_rate reports "N/A", and a new always-present boolean
system.has_traffic lets clients branch on the no-traffic state without parsing the
success-rate string. Previously a healthy fresh boot fell through a < 90.0 success-rate
threshold and reported critical, which coupled with the dashboard produced a misleading
red status on first start.
id derivation changedThe id values surfaced on /internal/status, /internal/status/endpoints, and the
per-model endpoint_ids on /internal/status/models are now derived from the sanitised
URL (scheme+host+port+path) with positional disambiguation for siblings that share a
sanitised form. Credential rotation no longer changes the ID, because userinfo, query,
and fragment are stripped before hashing. Bookmarked dashboard deep-links to a specific
endpoint row will change once after upgrading, then stay stable for endpoints with a
distinct name or a distinct sanitised URL. IDs may change again if a sibling sharing the
same sanitised URL is later added or removed, because the shared -N suffix is positional
and gets renumbered (and the degenerate case of two endpoints sharing both name and
sanitised URL is not immune either). The IDs are base36 FNV-1a, identical across all three
status payloads for the same endpoint from the same repository snapshot.
url in status responses is now sanitisedThe url field on /internal/status and /internal/status/endpoints now has userinfo,
query string, and fragment stripped before it is surfaced. (/internal/status/models
exposes no URL at all - only endpoint display names and endpoint_ids.) Previously this
field echoed the raw configured URL, which could include embedded credentials. Clients
that compared this field against the raw config value for identity should switch to the
id field, which is designed for that purpose.
The three status JSON routes now emit a weak ETag of the form W/"<base36>" over
their stable fields. If-None-Match uses weak comparison, so clients that echo the
ETag verbatim are unaffected. Clients doing strong-only comparison should allow weak
matches. See Conditional requests and compression
for the full contract.
Successful proxy requests log at Info again in the access log. A regression had
demoted them to Debug, which silently broke the "log all requests" expectation the
security practices document promises. Only /internal/ GET/HEAD polling that returns
2xx or 304 stays at Debug so an open dashboard tab does not flood the log; 4xx,
5xx, and any non-GET/HEAD method under /internal/ continue to log at Info.
A binary built without make build-web (e.g. via go install or a plain go build)
now logs a clear startup line and serves 503 at /internal/ui/ with a body naming the
fix. Previously such a binary served a silent placeholder. See
Development Setup: building the dashboard.
/internal/ui/A read-only, single-page dashboard is now embedded in the binary and served from the
same listener as the proxy: Overview (fleet status, success rate, latency, a live
requests-per-second sparkline), Endpoints (per-endpoint health, priority, latency,
model count), and Models (inventory grouped by family, with hosting endpoints). It
polls the existing /internal/status* JSON every 5-15 seconds with ETag/304
caching - no WebSocket, no push, and nothing in it changes state.
There is no authentication. Access is controlled entirely by the new dashboard.access_policy
config block (allowed_cidrs + allowed_hosts), which defaults to loopback-only. The
published Docker image ships with allowed_cidrs pre-widened to the RFC1918 ranges and
the listener bound to 0.0.0.0, so docker run -p 40114:40114 ghcr.io/thushan/olla:latest
followed by opening /internal/ui/ works with no config mount - and also means the
dashboard, and the container, are reachable from anyone else on the same LAN through that
published port. The access policy gates /internal/ui/ only: the /internal/status*,
/internal/health, /internal/metrics and /version JSON it reads stay as reachable as
they already were, allowed_cidrs or not. Unifying that gap under one policy is tracked
in issue #214.
See Admin Dashboard for the full security model and configuration reference.
GET /internal/metricsPrometheus-format metrics are now exposed directly, built from the same data as
/internal/status and /internal/stats/models - system status, per-endpoint health,
security counters, model usage and routing. No external exporter is required for core
proxy monitoring. Thanks to @Puupuls for the contribution. See
System API for the metric series.
A request that lands on an endpoint whose circuit breaker is open now fails over to the next available endpoint instead of the request failing outright. The endpoint is removed from that request's candidate list only; its persisted health is left untouched, because an open breaker already reflects accumulated failure state and demoting health on top of it is the proxy engine's job for genuine connectivity failures, not the retry path's. Exhaustion errors now distinguish circuit-breaker-open counts from connection-failure counts instead of conflating the two under one message.
A request for a model_aliases entry whose target model exists on no endpoint used to be
logged as rejected but still proxied to a compatible backend anyway, returning 200 from
the wrong model. It now returns 404/503 with routing_action: rejected, matching how
unknown models are already handled. One side effect: /olla/proxy/ with zero healthy
endpoints now returns 503 instead of 502 - update any monitor keyed on that specific
status code.
owned_by in converted model listings may return the raw matched segmentThe organisation-extraction logic used by the vLLM, vLLM-MLX, SGLang, llama.cpp, LMDeploy
and Lemonade converters is now shared. Alongside the existing org/model slash split, it
also recognises a hyphen-separated leading segment against a known-organisation list and
returns that segment verbatim - so a model ID like Qwen2.5-7B now yields owned_by: "Qwen2.5" rather than falling through to the converter's generic default. Anything
parsing owned_by for exact organisation matching should treat it as a best-effort label,
not a canonical identifier.
config/models.yaml customisations are no longer silently discardedThe shipped config/models.yaml used \d inside double-quoted YAML scalars, which is not
a valid YAML escape - the file has never parsed, on any install. The failure was swallowed
and Olla silently fell back to embedded defaults, so any customisation (e.g. capability
tagging via name_patterns) had no effect and no diagnostic. The regex patterns are fixed
to single-quoted scalars, config loading now happens eagerly at boot rather than lazily on
first request, and a found-but-unparseable candidate logs a WARN with the path and YAML
error instead of failing silently.
--validate-config flagRun olla --validate-config to check the configuration and provider profiles without
starting the server - a clear pass/warn/fail report with exit codes, useful in CI or
before a restart.
logging.level in config is now honouredThe logging.level config field was parsed but never applied to the runtime logger.
Precedence is OLLA_LOGGING_LEVEL env var, then config file, then default; an invalid
value now warns and falls back instead of being silently ignored.
Olla is a high-performance proxy and load balancer for LLM infrastructure.
Olla is a high-performance proxy and load balancer for LLM infrastructure.
# Docker
docker pull ghcr.io/thushan/olla:v0.0.28
# Binary (see assets below)
./olla --config config.yamlThis is a huge release and contains a lot of new exciting changes and bugfixes!
anthropic_support.messages_path), fixing 404s against DMR's /anthropic/v1/messages (#171).downloaded flag is mapped to available state, so downloaded models route correctly instead of being treated as unhealthy. Thanks to @matthewjhunter for the report and fix (#161, issue #160).type: "openai" now resolves model listings correctly via /olla/openai/v1/models; removed the duplicate openai.yaml profile that caused non-deterministic route registration. openai-compatible is the canonical value,openai remains an accepted alias. Thanks to @petersimmons1972 for the report (#151, issue #148).Highly requested feature finally implemented!
bearer, api_key and basic auth, with credentials from inline strings, ${ENV_VAR}, or _file siblings for Docker/k8s. Works with vLLM/llama.cpp/LiteLLM --api-key or any bearer-auth reverseconfig_error instead of dead; 429 honours Retry-After without tripping the circuit breaker; POST retries are skipped once response bytes have flushed (no double billing on mid-stream resets).max_body_size on non-proxy and translator routes (including chunked bodies) (#173).X-Olla-* headers from upstream responses (anti-spoofing), sanitise X-Request-ID to prevent log injection, and stop leaking the auth scheme in status JSON (#173).rs/cors) for browser-based clients like OpenWebUI and dashboards. Off by default; auto-exposes the X-Olla-* header set. Thanks to @ccsmart for the request (#159, issue #156).reasoning/reasoning_content now maps to Anthropic thinking blocks in both streaming and non-streaming responses (#176).usage chunks, synthesises fallback output_tokens, keeps same-chunk reasoning/content/tool-call deltas intact, and fixes repeated/interleaved tool-call handling (#176).Flush/Unwrap on wrapped writers (#176).context_management) are now ignored during translation and forwarded unchanged in passthrough mode. Thanks to @RenxuLogan for the report (#157, issue #154).OLLA_-overridable (#164, #165):server.read_header_timeout (default 10s) - guards against Slowloris-style slow-header attacks.proxy.response_header_timeout (default 30s) - raise for on-demand model loaders (e.g. Lemonade) that would otherwise abort cold starts at 30s. Now honoured by both Olla and Sherpa. Thanks to @matthewjhunter (#164, issue #163).proxy.connection_keep_alive (default 30s) and proxy.tls_handshake_timeout (default 10s).model_registry.unification.cleanup_interval / stale_threshold are now honoured (previously silently ignored, hard-coded to 5m) (#171).atomic.Pointer, snapshotted once per request/stream, so a hot-reload can't hand an in-flight request a torn config (#173, #176).LoadOrCompute, eliminating racing transport rebuilds on first use (#173).litepool.Get returns errors instead of panicking (handles nil and typed-nil factory results) (#176).runtime.GC(), an unread request counter) (#173)./internal/status/models now returns endpoint names for every entry (was inconsistently URL-then-names). Consumers reading endpoints[0] as a URL must switch to names (#173).X-Olla-* response headers are dropped, so a chained Olla-in-front-of-Olla setup loses the inner instance's headers (#173).Olla and the default load balancer is least-connections (#172)./olla-validate --quick / --nightly) with a local multi-protocol mock backend and fault injection - no Docker, CI-safe (#174)./version self-path reporting (previously reported a 404'ing /internal/version) (#171).Thanks to everyone who reported issues and contributed to this release:
response_header_timeout (#164, issue #163).type: "openai" endpoints (#148).Full Changelog: v0.0.27...v0.0.28
Documentation: thushan.github.io/olla | Issues: github.com/thushan/olla/issues
Nothing published for this version
We've added native support for LMDeploy after so long as well as bugfixes for sticky sessions thanks to @lbatalha .
We've added native support for LMDeploy after so long as well as bugfixes for sticky sessions thanks to @lbatalha.
# Docker
docker pull ghcr.io/thushan/olla:v0.0.27
# Binary (see assets below)
./olla --config config.yamlDocumentation: thushan.github.io/olla | Issues: github.com/thushan/olla/issues
This is a bugfix release that addresses SSL connection issues due to mis-handling of the Host header.
This is a bugfix release that addresses SSL connection issues due to mis-handling of the Host header.
# Docker
docker pull ghcr.io/thushan/olla:v0.0.26
# Binary (see assets below)
./olla --config config.yamlDocumentation: thushan.github.io/olla | Issues: github.com/thushan/olla/issues
Nothing published for this version
Olla is a high-performance proxy and load balancer for LLM infrastructure.
Olla is a high-performance proxy and load balancer for LLM infrastructure.
# Docker
docker pull ghcr.io/thushan/olla:v0.0.25
# Binary (see assets below)
./olla --config config.yamlThanks to @dnnspaul for contributing the Model Aliasing feature to Olla to alias models easily via the configuration.
We've now got a way of having sticky sessions in Olla to help keep requests aligned to KV Caches across multiple endpoints, taken from the working implementation in TensorFoundry's FoundryOS.
Lots of bugfixes and chores from March & April.
Documentation: thushan.github.io/olla | Issues: github.com/thushan/olla/issues
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →