NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3612 most downloaded on PyPI
Containers for machine learning
Last release 12 days ago
22 Sep 2026
Release timing varies
gaps range from 8 days to 2 months
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
184 releases · first in 2021
One column per quarter.
GPU compatibility warnings in cog doctor . cog doctor now warns when the resolved PyTorch and CUDA wheel ships no kernels for the local GPU, catching
cog doctor. cog doctor now warns when the resolved PyTorch and CUDA wheel ships no kernels for the local GPU, catching images that build successfully but fail when they run CUDA operations. (#3104)cog push. cog push <target> --json emits one versioned JSON document with digest-pinned image, model, and managed-weight references for image and bundle projects. Progress and diagnostics remain on stderr. (#3178).dockerignore guidance. The getting-started documentation now explains Cog's Docker build context and how to exclude development-only files to reduce build time and image size. (#3165)--break-system-packages and creates the expected python alias for both prebuilt and direct CUDA 13 base images. (#3160)Local servers now bind to 127.0.0.1 instead of 0.0.0.0 . cog serve , cog run -p , cog predict , and cog train no longer expose the prediction endpoint
127.0.0.1 instead of 0.0.0.0. cog serve, cog run -p, cog predict, and cog train no longer expose the prediction endpoint to your whole network by default. To get the old behavior back, use cog serve --host 0.0.0.0 or cog run -p 0.0.0.0:8888 — -p now accepts host:port. (#2812)Browser playground. cog playground starts a local, schema-driven web UI for talking to a running model — a Postman-like tool that reads the model's OpenAPI schema and builds the input form for you, with a request inspector and webhook relay. Point it at anything speaking the Cog HTTP API with --target. (#3123)
cog serve --playground. Serve the model and the playground in one command. The playground gets its own port (--playground-port, default 9000, 0 picks a free one) and both URLs are printed on startup. (#3154)
@cog.concurrent decorator. Configure async runner concurrency on the method that does the work instead of in cog.yaml:
class Runner(cog.BaseRunner):
@cog.concurrent(max=4)
async def run(self, prompt: str = cog.Input()) -> str:
return await generate(prompt)The value is extracted statically at build time. concurrency.max in cog.yaml still works and still takes precedence, but it now warns on build — remove it when you migrate. (#3089)
Input validation before build and run. cog predict, cog run, and cog train now validate -i inputs against the model's OpenAPI schema before building the image or starting the container, so a typo fails in a second rather than after a full build. Covers unknown input names, missing required inputs, type mismatches, enum values, numeric bounds, string length, and regex patterns. The schema is generated statically from cog.yaml; for existing images it's read from the openapi_schema label. (#3097)
accept on Path and File inputs. Declare the MIME types or extensions an input takes — Input(accept="image/*"), Input(accept="audio/wav,audio/mp3"), Input(accept=".safetensors,.bin") — and Cog emits it as x-cog-accept in the OpenAPI schema so UIs and clients know what to offer. Using accept on a non-file type is a build error. (#2863)
Local wheels and source archives in requirements. build.python_requirements now accepts bare local paths like ./dist/mylib-0.1.0-py3-none-any.whl or ./packages/localpkg.tar.gz. Paths resolve relative to the requirements file, and only the referenced artifacts get staged into the build. .whl, .zip, .tar.gz, .tgz, .tar.bz2, and .tar.xz are supported. (#3095)
cog serve, cog run, cog predict, and cog train now run the model with console-formatted logs instead of production JSON. Deployed models are unaffected. (#2806)cog weights. cog weights import and cog weights pull show live Docker-style progress while they populate the local weight store. (#3087)--use-cog-base-image works, and how it interacts with --use-cuda-base-image. (#3088)nvidia/cuda images failed with externally-managed-environment because uv-managed Python is marked PEP 668 externally managed and uv 0.9.x enforces it. Cog now passes --break-system-packages wherever it installs into a uv-managed Python, not just in cog base-image. (#3073)cudnn with a CUDA minor version now validates. A config pairing cuda: "11.6" with cudnn: "8" was rejected up front even though 11.6 resolves fine to a patch-level base image like nvidia/cuda:11.6.2-cudnn8-devel-ubuntu20.04. Validation now matches versions the same way image selection does. (#3070)uname -m reports aarch64 there, but the release asset is cog_Linux_arm64, so the documented curl command and tools/install.sh both 404'd. Both now normalize the architecture. (#3066)info.version was hard-coded to 0.1.0; it now carries the version that built the schema. (#3075)cog weights pull is quieter. Directory entries in weight tarballs no longer produce skipping unexpected tar entry lines. (#3076)Server-Sent Event prediction streams. HTTP prediction requests can now ask for Accept: text/event-stream to receive start , output , log , metric , an
Accept: text/event-stream to receive start, output, log, metric, and terminal completed events for predictors that explicitly opt in with @streaming / @cog.streaming. Reconnecting clients can replay retained in-flight prediction events with PUT /predictions/{id}. (#3019)replicate/cog-examples have moved into this repository under examples/, with updated run.py / BaseRunner examples for common workflows including images, training, notebooks, streaming, concurrency, context, and Replicate API usage. (#3055)cog weights. Every cog weights subcommand now prints a warning that the weights workflow is experimental and should not be relied on in production workflows yet. (#3025)run() as the primary prediction entry point while preserving legacy predict() fallback behavior, and resolves inherited and imported targets more consistently. (#3027)cog.Secret inputs now work with coglet-backed predictions. Predictors that annotate inputs as Secret, Optional[Secret], or Secret | None now receive cog.types.Secret values at runtime again, restoring behavior that regressed in the Rust/coglet runtime rewrite. (#3057)cog doctor --fix now shows available remediation text. Findings without an auto-fix now display their remediation message instead of incorrectly saying no auto-fix is available. (#3031)from __future__ import annotations no longer reject valid setup() -> None methods or accept invalid run() -> None methods because None annotations were stored as strings. (#3034)cog serve now mounts weights like cog run. Models served locally now receive configured weights consistently with local prediction runs. (#3044)None at the HTTP edge. Optional predictor inputs omitted from HTTP requests are now passed as None instead of being treated inconsistently by request handling. (#3051)d6c2b70 Bump version to 0.21.0-rc.3
Server-Sent Event prediction streams. HTTP prediction requests can now ask for Accept: text/event-stream to receive start , output , log , metric , an
Accept: text/event-stream to receive start, output, log, metric, and terminal completed events for predictors that explicitly opt in with @streaming / @cog.streaming. Reconnecting clients can replay retained in-flight prediction events with PUT /predictions/{id}. (#3019)cog weights. Every cog weights subcommand now prints a warning that the weights workflow is experimental and should not be relied on in production workflows yet. (#3025)run() as the primary prediction entry point while preserving legacy predict() fallback behavior, and resolves inherited and imported targets more consistently. (#3027)cog doctor --fix now shows available remediation text. Findings without an auto-fix now display their remediation message instead of incorrectly saying no auto-fix is available. (#3031)from __future__ import annotations no longer reject valid setup() -> None methods or accept invalid run() -> None methods because None annotations were stored as strings. (#3034)cog run command. The cog predict command has been renamed to cog run with full backward compatibility. cog predict still works as an alias.
cog run command. The cog predict command has been renamed to cog run with full backward compatibility. cog predict still works as an alias. (#3015)cog push and weights commands. You can now reference models by name (e.g., r8.im/user/model) instead of full image URLs when pushing or managing weights. (#3018)Opaque annotation to exclude fields from the generated schema. (#3001).cog/ directory that is automatically filtered from the Docker build context. (#3000)uv. (#2999)cog push no longer includes the image tag (e.g., :latest), preventing 404 errors when users click the link. (#3020)cac3e91 feat: experimental managed weights
Support for TypedDict in schema generation. Fixed an issue where TypedDict type annotations would cause schema generation to fail.
cog doctor command. Diagnose common Cog setup issues, check configuration, and verify that everything is working correctly. Run cog doctor to validate
cog doctor command. Diagnose common Cog setup issues, check configuration, and verify that everything is working correctly. Run cog doctor to validate your environment. (#2923)COG_LEGACY_SCHEMA=1 to opt out if you encounter issues. (#2950)r8.im/... image names. (#2954)Secret = Input(default=None) is treated as optional. Secret inputs with None defaults are now correctly identified as optional in the generated schema. (#2949)`cog run` is now `cog exec`. cog run still works as a hidden alias with a deprecation warning -- existing scripts won't break yet, but update them.
cog run is now cog exec. cog run still works as a hidden alias with a deprecation warning -- existing scripts won't break yet, but update them. (#2916)async def setup() actually runs now. In 0.17.x, async setup coroutines were silently dropped -- setup appeared to succeed but none of the code executed, causing AttributeError on every prediction. (#2921)setup() (httpx clients, aiohttp sessions, asyncio queues) no longer crash because setup and predict run on different loops. (#2927)dict and list[dict] work as input types. These were supported as outputs but rejected as inputs, breaking chat-style message inputs. (#2928)list[X] | None works as an input type. The type system only had Required, Optional, and Repeated -- not optional-and-repeated. Both the Python SDK and Go schema generator now handle this correctly. (#2882)cog push now shows status during the docker save phase instead of sitting silent while large images export to disk. (#2797)record_metric() enforces naming rules -- must start with a letter, no consecutive underscores, max 128 chars, max 4 segments. predict_time and the cog. prefix are reserved. (#2911)b43abeae4c94f52b8f4bbf037ee9d48e98dcbf3d fix: replace deprecated library usage patterns
635fffb8b981fa09ad4dabb4fc6c40344ad3657b fix: don't coerce URL strings in str-typed inputs (regression #2868)
Under the hood, though, this is the foundation for what's coming next. A few things did change... see the breaking changes section at the bottom.
This is a big release. The prediction server has been rewritten in Rust, Pydantic dependency conflicts are a thing of the past, and several long-requested QoL features are here. We've tested extensively against models on Replicate and the vast majority work without any changes. Under the hood, though, this is the foundation for what's coming next. A few things did change... see the breaking changes section at the bottom.
The Python HTTP server that ran inside Cog containers has been replaced with a Rust-based server called coglet. It uses a two-process architecture -- a Rust parent handling HTTP and orchestration, and a Python worker subprocess running your predict function. You don't need to change anything in your code. Predictions are faster to start, the server handles concurrency better, and worker crashes no longer take down the whole container. More importantly, this is the runtime we'll be building on -- it unlocks things like native streaming, smarter scheduling, and tighter hardware integration that weren't possible with the old Python server.
You can now emit custom metrics from your predict function:
from cog import current_scope
def predict(self, prompt: str) -> str:
scope = current_scope()
scope.record_metric("tokens_generated", 128)
scope.record_metric("timing.inference", 0.42)
...
Metrics appear in the prediction response alongside predict_time. Supports "replace" (default), "incr", and "append" accumulation modes. Dot-path keys create nested objects.
Add a healthcheck() method to your Predictor to inject custom validation into the /health-check endpoint:
def healthcheck(self) -> bool:
# check GPU is responsive, weights are loaded, etc.
return True
Runs with a 5-second timeout, even when the model is busy processing predictions.
Callers can now pass a context dict with prediction requests, accessible in your predict function via current_scope().context. Useful for forwarding metadata like API tokens, region hints, or prediction IDs without changing your function signature.
Cog's type system was built on Pydantic, which meant the SDK pinned a specific Pydantic version in every container. If a package you depended on needed a different version, you were stuck. That's gone now -- cog.BaseModel is a standard Python dataclass, and Pydantic isn't installed at all unless you add it yourself. The API is unchanged (class Output(BaseModel): text: str still works), and if your predict function returns a Pydantic model it'll still be serialized correctly.
cog push and cog login now work with any OCI-compliant registry -- GHCR, GCR, ECR, Docker Hub, self-hosted. The provider is selected automatically based on the registry host in your image name.
New build.sdk_version field in cog.yaml lets you pin the Python SDK version independently from the CLI:
build:
sdk_version: "0.16.6"
Omit it to get the latest stable release. Set "prerelease" to opt into pre-release builds. Minimum supported version is 0.16.0.
cog push can now upload layers directly to the registry in parallel with automatic chunking (96 MB chunks, 5 concurrent uploads), blob deduplication, and retry with exponential backoff. Set COG_PUSH_OCI=1 to enable. Falls back to docker push if anything goes wrong.
uv for package managementDockerfiles now use uv instead of pip for faster, more reliable dependency installation inside containers.
The old hardcoded 5-minute setup timeout has been removed. You can now configure your own timeout with the COG_SETUP_TIMEOUT environment variable (in seconds). If unset, there's no internal timeout.
CLI output now uses color-coded prefixes (✔ green for success, ⚠ yellow for warnings, ✗ red for errors) and auto-detects color support from TTY/environment. Use --no-color or NO_COLOR=1 to disable.
python_version is now required in the build: section of cog.yaml. Builds fail with a clear error if it's missing. Add python_version: "3.13" (or 3.10/3.11/3.12).python_packages.cog train is deprecated. It still works but prints a warning. Will be removed in a future release so we can replace it with something better.emit_metric() is deprecated in favor of current_scope().record_metric(). The old function still works as a compat shim.a55abe518188ba53c3da7b7be5e9bde9adfb04a4 Bump version to 0.17.0-rc.4
49442909c5fce8bb9fa2a1ef98c23a33cf519f62 fix(sdk): restore emit_metric as deprecated compat shim
5ed0cee0f6ac6941e6e6a9d1cb460fc59f772fda chore: bump version to 0.17.0-rc.2
415d87c2c8b82d3b3c0cdc7958ec345061bd7a8f Remove deprecated fast/monobase build system
415d87c2c8b82d3b3c0cdc7958ec345061bd7a8f Remove deprecated fast/monobase build system
415d87c2c8b82d3b3c0cdc7958ec345061bd7a8f Remove deprecated fast/monobase build system
415d87c2c8b82d3b3c0cdc7958ec345061bd7a8f Remove deprecated fast/monobase build system
415d87c2c8b82d3b3c0cdc7958ec345061bd7a8f Remove deprecated fast/monobase build system
Nothing published for this version
ebdf3ebf53b7434ed206b898a55f2db6ca398dcb Also run release-specific CI for v0.16-maint branch
da2de8fcc0d97025a3fa5f1ef13631cfcc2165e3 Fix: use tmpImageId for schema validation when pushing to r8.im
1743e400ef689fd5c4f350595c16314d8d368f7d Fix integration tests
Nothing published for this version
7ef0e6439d0ebfea3454e427242d1cead4031d82 Bump actions/setup-python from 5 to 6
-buildvcs=false on go build (#2516)02102a94d101e05f7ab3ae82afbfbaf6c3d06ba6 Swap out deprecated ast.Str for ast.Constant
ast.Str for ast.Constant (#2514)e862f4ee download llms.txt from docs when running cog init
cog init -x--pipelines include AGENTS.md (#2499)0a6d85677912b8364bd6959e9827a8782ae36c1c Do not compare build metadata if null
Nothing published for this version
b9e72144787b8118495324aebf1556109119908b Add integration test for float with cog-runtime
23b93cc6f6db5cd64d39991b49b312be2dfc30e0 Add integration test for 3.13 base images
9d5593105fe45874fc3bb02fb1376ad7c9f1fe69 Deduplicate includes
c40d9534542b9912b929dd06efe088caa76f961f Add warning for cog_runtime flag
52a64d21ad566db53ab8f1b6bca9bac582d17bf3 Add cog_runtime flag
39583000c73f98911ae12aeeecb1e60411d4591e Add Torch 2.7.1 base images
b737a7bb1fd72ac31a19840ac8769f10fc79d29f Add specific handling for pushing procedure with versions
a57a47d63cd8349d59f1f3a3b2fcf71d7f5d513f Skip analysing packages with @
a433dddf60a561f24be0699fa1cee6ffb64bba83 Document the cog CLI
5df702ae08b4a804c373b020df1c578658f8e212 Fix FastAPI deprecation warnings by migrating to lifespan handlers
cog predict to support --json flag (#2404)2c4ab6ce1d9db6bbe257f07bc7a4809b98a42d0c Support pulling to specific folders
Check if token is empty by @8W9aG in https://github.com/replicate/cog/pull/2392
Full Changelog: https://github.com/replicate/cog/compare/v0.15.3...v0.15.4
70f567fb727dbe15a5051d1d15cb65380c026c63 Add integration test for python 3.13
os.PathLike output from models (#2388)59dc987b1c85b1d9d667b0a51e5b6e42130753d4 Add checking of pipelines runtime requirements before pushing
e5c83e711270674f1a439147d3982ed00fb768e9 Add deprecated field to Input
63b6246296c0a0b0dbb92d8b7f9f642271f0621e Add integration test for setup run in trainer
24f693ef9394ee0edd01941c49a2952959e77430 Add call_graph tool
4c8b448264e3d269b286eb810ec77e831544c8ab Add OCI registry client
command.Command (#2307)7cdd08daa38e33800091fa217163c8e163e77afa Add environment variable for defining coglet version
cog_binary fixture (#2304)7b0d5f1982d531a14cc576faacca786ccc291ea7 Add new version log
docker.Build to command.Command (#2293)e01c1dcfcb1f529843d35ab05496a461afa00da2 Add .python-version to .dockerignore
docker.Xxx helpers with functions on command.Command (#2288)738ad3d909f3110d6662cf16ca08ebfd1bdc0569 Disable provenance attestations
b07ab054e6bfdc4c8153e1e16cb4ffd1c4842d71 Add complex types to cog predict
beba0424aba813033c1d03eeb3bae9f3d17a5be7 Add empty namespace package for cog.ext
cog.ext (#2246)a4bac6e88c2f29158c5145358cb9550d241e34b4 Add a config option to enable fast builds
pathlib.Path and descendants as uploadables (#2235)uv + misc fixes while setting up a new dev environment (#2233)994b2f3fc83df8be49037e5f347a912648b443f5 Allow cancelation on async models
edd038500be8fdcaaee18691f3a65e620e7320f9 Warn the user when they use deprecated fields
Your coding agent can read these notes before it upgrades. Set up the MCP server →