NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3786 most downloaded on PyPI
SkyPilot: Manage all your AI compute.
Last release 13 days ago
22 Sep 2026
Release timing varies
gaps range from 8 days to 2 months
Nearly every release is documented
notes for 38 of 41 stable releases
2 versions withdrawn
withdrawn after publishing
4 years old
62 releases · first in 2022
One column per quarter.
Nothing published for this version
uv pip install " skypilot>=0.13.2rc1 "
Get it now with:
uv pip install "skypilot>=0.13.2rc1"replace_skypilot_config when the body raises by @kevinmingtarja in #10236resolved_config by @DanielZhangQD in #10270sky api logs traceback on 4xx from /api/stream by @DanielZhangQD in #10341Note truncated.
`bash uv pip install "skypilot>=0.13.1rc1" `
Get it now with:
uv pip install "skypilot>=0.13.1rc1"
disk_size instead of ephemeral_storage field by @kyuds in https://github.com/skypilot-org/skypilot/pull/9273Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.13.0rc1...v0.13.1rc1
Deprecations & Breaking Changes:
SkyPilot v0.13.0 introduces Sky Batch for large-scale batch inference and data processing across worker pools, a native Hugging Face storage backend (hf://), a generalized lifecycle-hooks framework for running custom logic around cluster events, and the sky debug-dump command for one-shot troubleshooting. This release also adds per-instance cost caps, NVIDIA Blackwell / CUDA 13 images, and GKE Autopilot support.
Get it now with:
uv pip install "skypilot[all]==0.13.0"
Or, upgrade your team SkyPilot API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.13.0
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Deprecations & Breaking Changes:
api_access renamed to api_server_access, now defaulting to True (#9169). Tasks now automatically receive API server credentials. The legacy api_access key is still accepted for backward compatibility.sky.jobs.queue(version=1) remains deprecated and continues to emit a warning; use sky.jobs.queue(version=2), which returns richer job metadata as dictionaries (#9118).all_includes_in_cluster from the SkyPilot config schema (#9671).**sky batch** brings large-scale batch inference and data processing to SkyPilot with a simple Python API (#8850, #9438, #9456). Given a dataset and a processing function, Sky Batch splits the data into batches, distributes them across a pool of GPU/CPU workers, and collects the results, with no manual sharding, provisioning, or failure handling. Workers are reused across jobs, so expensive setup (installing packages, downloading weights, loading models onto GPUs) happens only once.
import sky
# Point at your input data in cloud storage
ds = sky.batch.Dataset(source='s3://my-bucket/prompts/')
# Map a processing function over the dataset across a worker pool
ds.map(my_inference_fn, pool='my-gpu-pool', output='s3://my-bucket/results/')
Inputs and outputs can be read from and written to Amazon S3 (s3://) or Google Cloud Storage (gs://).
SkyPilot now supports Hugging Face as a first-class storage backend via the hf:// scheme (#9698, #9763). It wraps read-write Hugging Face Buckets (Xet-backed, S3-like object storage) and, through the same handle, read-only mounting of model, dataset, and space repos. Bucket lifecycle operations go through huggingface_hub, and MOUNT/MOUNT_CACHED modes use the FUSE-based hf-mount, consistent with SkyPilot's existing mount infrastructure.
file_mounts:
# Read-write Hugging Face Bucket
/data:
source: hf://buckets/my-org/my-bucket
# Read-only model / dataset / space repo
/model:
source: hf://meta-llama/Llama-3.1-8B
mode: MOUNT
URL formats: hf://buckets/<ns>/<bucket> (bucket, read-write), hf://<ns>/<model>[@<rev>], hf://datasets/<ns>/<ds>, and hf://spaces/<ns>/<sp> (repos, read-only). Authentication reads HF_TOKEN from the environment or the Hugging Face CLI token file.
The autostop hook introduced in v0.12 has grown into a generalized lifecycle-hooks framework via resources.hooks (#9064). Each hook declares a script and the events it fires on: autostop, preemption, and down. Use it to checkpoint state, sync experiment trackers, or send notifications around cluster lifecycle transitions.
resources:
hooks:
- run: |
wandb sync
curl -X POST $SLACK_WEBHOOK -d '{"text": "Cluster shutting down"}'
events: [autostop, down]
timeout: 300
The legacy autostop.hook YAML continues to work with a deprecation warning.
sky debug-dump: One-Shot TroubleshootingThe new **sky debug-dump** command collects diagnostic information (logs, state, and configuration) for clusters, requests, and managed jobs into a single zip for troubleshooting and bug reports (#8868, #9752, #9753). Dumps are stored under ~/.sky/debug_dumps/ and automatically cross-link related resources (e.g. a cluster and its associated requests) so you capture complete context in one command.
sky debug-dump --cluster my-cluster
You can now set a per-instance hourly budget cap with resources.max_hourly_cost (#9132, #9170). Only instances at or below the limit are considered during optimization. When use_spot is set, the cap applies against spot prices; otherwise against on-demand prices.
resources:
accelerators: A100
max_hourly_cost: 10.0
SkyPilot's default GPU image now ships NVIDIA driver 580 (open) + CUDA 13 for Blackwell support (#9639), with a legacy CUDA 12 image retained for pre-Turing GPUs and an arm64 build fix (#9789). RunPod adds RTX PRO Blackwell GPUs to its catalog (#9187).
SkyPilot now detects GKE Autopilot clusters and trusts Node Auto-Provisioning (NAP) instead of walking existing node pools, so GPU requests provision correctly on Autopilot instead of being incorrectly rejected (#9565).
MIN_COMPATIBLE_API_VERSION bump for 0.12 clients (#9121)._next/ assets instead of index.html (#9722); don't cache error responses on /dashboard/_next paths (#9498).async_creator (#9483), put_async for the queue backend (#9346).enabled_clouds and sky check (#9284), cached permission checks (#9125).devMode flag for in-cluster iteration (#9609), persist the whole sky_logs dir (#9454), optional server log mirror (#9754), configurable probes (#9477), disable session affinity on ingress (#9347), option to disable serving the server log (#9140).autopilot.enabled detection and NAP trust (#9565).context_configs (#9640).**ephemeral-storage requests/limits** on pods (#9053), allowed_nodes config for node-level filtering (#9151), pod spread constraints (#9296).pod_config to override runtimeClassName for GPU pods (#9241); probe Ray ports when hostNetwork: true (#9644).network_tier: best for OCI OKE RoCE (#9622); allow excluding the in-cluster context from allowed_contexts: all (#9637).volumeMounts patch-merge key from name to mountPath (#9358), write permissions for all volume mounts at startup (#9221), owner-identity check for scoped kubeconfig names (#9725, #9190). Updated default/base k8s images (#9618, #9590).k8s-hostpath volume type for node-local cache (#9165).access_mode from an existing PVC for use_existing volumes (#9359); check the existing backend when creating volumes (#9104); add volume mount check (#9343).Ti units for volume sizes ≥ 1024Gi (#9369); don't rewrite region to in-cluster when in-cluster auth is unavailable (#9743).hf://) for buckets and repos (#9698, #9763).MOUNT (#9225).S3CompatibleStore (#9080); fixed git "dubious ownership" during workdir sync (#9352).__pycache__ from the built wheel (#9413).file_mounts destinations again (#9537); fixed cluster re-provisioning and atomic SSH-key writes (#9476); prioritize user SSH keys over cluster keys in Host * config (#9099)./dev/fuse fd after passing it to the client (#9463); fixed a flock leak in the uv_spawn child process (#9431); fixed browser opening on WSL for sky api login (#9178).SKY_RUNTIME_DIR and removed import-time filesystem side effects (#9605); let packages expose client SDKs under the sky namespace (#9784).**--tail for sky jobs logs** (#9449), preload_content for sky.jobs.tail_logs (#9662), and a wait helper in the programmatic SDK (#9138).sky jobs launch (#9761, #9135, #9191); optionally merge cluster launch-progress events into job events (#9804).ServiceStatusRunner extension point (#9801), new user-configurable options (#9116), and improved max-services error messages (#8897).AWS
aws login (2026 AWS CLI command) (#9526) and subnet selection (#9127).subnet_names span multiple AZs (#9791); added AZ-fetch timeouts to prevent sky check hangs (#9244, #9240).GCP
provisioningModel=RESERVATION_BOUND for DENSE reservations (#9779); added missing required instance types to the catalog (#9472); pinned/bumped the gcloud SDK (#9587, #9535); sped up sky check (#9095).Nebius
open_ports replacing UFW (#9563); use managed disks to create instances (#9432); parallelize multi-node VM creation (#9570); custom image_id (#9133); fixed autodown (#9573).Slurm
cpu_partition config for CPU-only tasks (#9281) and gpu_partition_map for GRES without a GPU type (#9237).--time and fall back to DefaultTime when MaxTime is UNLIMITED (#9633); preserve SLURM_CONF* env vars when launching srun (#9425).sbatch (#9174), numeric partition names treated as integers (#9168), and --mem when memory is not consumable (#9171).RunPod / Vast / PrimeIntellect
create_template (#9550).vastai-sdk v6+ (#9314); made account-level SSH key registration best-effort (#9631).query_instances TypeError on retry (#9259) and instance update after _get_instance_info (#9327).sky ssh up (#9402).Dashboard
allowed_users (#9696); show counts above the users and service-account tables (#9655)./status when the controller is up (#9645).CLI
**sky api login --no-browser** for headless environments, continuing to poll when a browser can't be opened (#9677, #9676).**sky debug-dump** troubleshooting command (#8868).cost-report (#9504).kubernetes < 36 due to an RBAC regression, while accepting dict[K, V] openapi_types for kubernetes >= 36 (#9685, #9682).azure-cli < 2.87.0 (azure-mgmt-storage 25.0.0 breaks Azure storage) (#9774).pyopenssl upper limit (#8070); pinned click < 8.3.0 in relevant runtime installs (#9459, #9763).MOUNT_CACHED config documentation (#9130); api_server_access documentation (#9115); refreshed managed jobs, quickstart, and agent-skills docs (#9564, #9250).Thank you to all contributors who made this release possible!
@aylei, @Michaelvll, @DanielZhangQD, @zpoint, @kevinmingtarja, @lloyd-brown, @concretevitamin, @cg505, @SeungjinYang, @rohansonecha, @kevinzwang, @kyuds, @romilbhardwaj, @alex000kim, @andylizf, @cblmemo, @hentt30, @williamsnell, @Jobarion, and many others.
Note: this is a draft. Use GitHub's "Generate release notes" on the tagged release to finalize the exact contributor @handles and the "New Contributors" section.
For a complete list of changes, see the commit history.
Deprecations & Breaking Changes:
SkyPilot v0.13.0rc1 introduces Sky Batch for large-scale batch inference and data processing across worker pools, a native Hugging Face storage backend (hf://), a generalized lifecycle-hooks framework for running custom logic around cluster events, and the **sky debug-dump** command for one-shot troubleshooting. This release also adds per-instance cost caps, NVIDIA Blackwell / CUDA 13 images, and GKE Autopilot support.
Get it now with:
uv pip install "skypilot[all]==0.13.0rc1"
Or, upgrade your team SkyPilot API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.13.0rc1
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Deprecations & Breaking Changes:
**api_access renamed to api_server_access, now defaulting to True** (#9169). Tasks now automatically receive API server credentials. The legacy api_access key is still accepted for backward compatibility.**sky.jobs.queue(version=1) remains deprecated** and continues to emit a warning; use sky.jobs.queue(version=2), which returns richer job metadata as dictionaries (#9118).all_includes_in_cluster from the SkyPilot config schema (#9671).**sky batch** brings large-scale batch inference and data processing to SkyPilot with a simple Python API (#8850, #9438, #9456). Given a dataset and a processing function, Sky Batch splits the data into batches, distributes them across a pool of GPU/CPU workers, and collects the results, with no manual sharding, provisioning, or failure handling. Workers are reused across jobs, so expensive setup (installing packages, downloading weights, loading models onto GPUs) happens only once.
import sky
# Point at your input data in cloud storage
ds = sky.batch.Dataset(source='s3://my-bucket/prompts/')
# Map a processing function over the dataset across a worker pool
ds.map(my_inference_fn, pool='my-gpu-pool', output='s3://my-bucket/results/')
Inputs and outputs can be read from and written to Amazon S3 (s3://) or Google Cloud Storage (gs://).
SkyPilot now supports Hugging Face as a first-class storage backend via the hf:// scheme (#9698, #9763). It wraps read-write Hugging Face Buckets (Xet-backed, S3-like object storage) and, through the same handle, read-only mounting of model, dataset, and space repos. Bucket lifecycle operations go through huggingface_hub, and MOUNT/MOUNT_CACHED modes use the FUSE-based hf-mount, consistent with SkyPilot's existing mount infrastructure.
file_mounts:
# Read-write Hugging Face Bucket
/data:
source: hf://buckets/my-org/my-bucket
# Read-only model / dataset / space repo
/model:
source: hf://meta-llama/Llama-3.1-8B
mode: MOUNT
URL formats: hf://buckets/<ns>/<bucket> (bucket, read-write), hf://<ns>/<model>[@<rev>], hf://datasets/<ns>/<ds>, and hf://spaces/<ns>/<sp> (repos, read-only). Authentication reads HF_TOKEN from the environment or the Hugging Face CLI token file.
The autostop hook introduced in v0.12 has grown into a generalized lifecycle-hooks framework via resources.hooks (#9064). Each hook declares a script and the events it fires on: autostop, preemption, and down. Use it to checkpoint state, sync experiment trackers, or send notifications around cluster lifecycle transitions.
resources:
hooks:
- run: |
wandb sync
curl -X POST $SLACK_WEBHOOK -d '{"text": "Cluster shutting down"}'
events: [autostop, down]
timeout: 300
The legacy autostop.hook YAML continues to work with a deprecation warning.
sky debug-dump: One-Shot TroubleshootingThe new **sky debug-dump** command collects diagnostic information (logs, state, and configuration) for clusters, requests, and managed jobs into a single zip for troubleshooting and bug reports (#8868, #9752, #9753). Dumps are stored under ~/.sky/debug_dumps/ and automatically cross-link related resources (e.g. a cluster and its associated requests) so you capture complete context in one command.
sky debug-dump --cluster my-cluster
You can now set a per-instance hourly budget cap with resources.max_hourly_cost (#9132, #9170). Only instances at or below the limit are considered during optimization. When use_spot is set, the cap applies against spot prices; otherwise against on-demand prices.
resources:
accelerators: A100
max_hourly_cost: 10.0
SkyPilot's default GPU image now ships NVIDIA driver 580 (open) + CUDA 13 for Blackwell support (#9639), with a legacy CUDA 12 image retained for pre-Turing GPUs and an arm64 build fix (#9789). RunPod adds RTX PRO Blackwell GPUs to its catalog (#9187).
SkyPilot now detects GKE Autopilot clusters and trusts Node Auto-Provisioning (NAP) instead of walking existing node pools, so GPU requests provision correctly on Autopilot instead of being incorrectly rejected (#9565).
MIN_COMPATIBLE_API_VERSION bump for 0.12 clients (#9121)._next/ assets instead of index.html (#9722); don't cache error responses on /dashboard/_next paths (#9498).async_creator (#9483), put_async for the queue backend (#9346).enabled_clouds and sky check (#9284), cached permission checks (#9125).devMode flag for in-cluster iteration (#9609), persist the whole sky_logs dir (#9454), optional server log mirror (#9754), configurable probes (#9477), disable session affinity on ingress (#9347), option to disable serving the server log (#9140).autopilot.enabled detection and NAP trust (#9565).context_configs (#9640).**ephemeral-storage requests/limits** on pods (#9053), allowed_nodes config for node-level filtering (#9151), pod spread constraints (#9296).pod_config to override runtimeClassName for GPU pods (#9241); probe Ray ports when hostNetwork: true (#9644).network_tier: best for OCI OKE RoCE (#9622); allow excluding the in-cluster context from allowed_contexts: all (#9637).volumeMounts patch-merge key from name to mountPath (#9358), write permissions for all volume mounts at startup (#9221), owner-identity check for scoped kubeconfig names (#9725, #9190). Updated default/base k8s images (#9618, #9590).k8s-hostpath volume type for node-local cache (#9165).access_mode from an existing PVC for use_existing volumes (#9359); check the existing backend when creating volumes (#9104); add volume mount check (#9343).Ti units for volume sizes ≥ 1024Gi (#9369); don't rewrite region to in-cluster when in-cluster auth is unavailable (#9743).hf://) for buckets and repos (#9698, #9763).MOUNT (#9225).S3CompatibleStore (#9080); fixed git "dubious ownership" during workdir sync (#9352).__pycache__ from the built wheel (#9413).file_mounts destinations again (#9537); fixed cluster re-provisioning and atomic SSH-key writes (#9476); prioritize user SSH keys over cluster keys in Host * config (#9099)./dev/fuse fd after passing it to the client (#9463); fixed a flock leak in the uv_spawn child process (#9431); fixed browser opening on WSL for sky api login (#9178).SKY_RUNTIME_DIR and removed import-time filesystem side effects (#9605); let packages expose client SDKs under the sky namespace (#9784).**--tail for sky jobs logs** (#9449), preload_content for sky.jobs.tail_logs (#9662), and a wait helper in the programmatic SDK (#9138).sky jobs launch (#9761, #9135, #9191); optionally merge cluster launch-progress events into job events (#9804).ServiceStatusRunner extension point (#9801), new user-configurable options (#9116), and improved max-services error messages (#8897).AWS
aws login (2026 AWS CLI command) (#9526) and subnet selection (#9127).subnet_names span multiple AZs (#9791); added AZ-fetch timeouts to prevent sky check hangs (#9244, #9240).GCP
provisioningModel=RESERVATION_BOUND for DENSE reservations (#9779); added missing required instance types to the catalog (#9472); pinned/bumped the gcloud SDK (#9587, #9535); sped up sky check (#9095).Nebius
open_ports replacing UFW (#9563); use managed disks to create instances (#9432); parallelize multi-node VM creation (#9570); custom image_id (#9133); fixed autodown (#9573).Slurm
cpu_partition config for CPU-only tasks (#9281) and gpu_partition_map for GRES without a GPU type (#9237).--time and fall back to DefaultTime when MaxTime is UNLIMITED (#9633); preserve SLURM_CONF* env vars when launching srun (#9425).sbatch (#9174), numeric partition names treated as integers (#9168), and --mem when memory is not consumable (#9171).RunPod / Vast / PrimeIntellect
create_template (#9550).vastai-sdk v6+ (#9314); made account-level SSH key registration best-effort (#9631).query_instances TypeError on retry (#9259) and instance update after _get_instance_info (#9327).sky ssh up (#9402).Dashboard
allowed_users (#9696); show counts above the users and service-account tables (#9655)./status when the controller is up (#9645).CLI
**sky api login --no-browser** for headless environments, continuing to poll when a browser can't be opened (#9677, #9676).**sky debug-dump** troubleshooting command (#8868).cost-report (#9504).kubernetes < 36 due to an RBAC regression, while accepting dict[K, V] openapi_types for kubernetes >= 36 (#9685, #9682).azure-cli < 2.87.0 (azure-mgmt-storage 25.0.0 breaks Azure storage) (#9774).pyopenssl upper limit (#8070); pinned click < 8.3.0 in relevant runtime installs (#9459, #9763).MOUNT_CACHED config documentation (#9130); api_server_access documentation (#9115); refreshed managed jobs, quickstart, and agent-skills docs (#9564, #9250).Thank you to all contributors who made this release possible!
@aylei, @Michaelvll, @DanielZhangQD, @zpoint, @kevinmingtarja, @lloyd-brown, @concretevitamin, @cg505, @SeungjinYang, @rohansonecha, @kevinzwang, @kyuds, @romilbhardwaj, @alex000kim, @andylizf, @cblmemo, @hentt30, @williamsnell, @Jobarion, and many others.
Note: this is a draft. Use GitHub's "Generate release notes" on the tagged release to finalize the exact contributor @handles and the "New Contributors" section.
For a complete list of changes, see the commit history.
`bash uv pip install "skypilot>=0.12.3.post1" `
Get it now with:
uv pip install "skypilot>=0.12.3.post1"
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.12.3rc1...v0.12.3.post1
`bash uv pip install "skypilot>=0.12.3" `
Get it now with:
uv pip install "skypilot>=0.12.3"
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.12.3rc1...v0.12.3
Nothing published for this version
`bash uv pip install "skypilot>=0.12.2.post1" `
Get it now with:
uv pip install "skypilot>=0.12.2.post1"
What's changed?
`bash uv pip install "skypilot>=0.12.2" `
Get it now with:
uv pip install "skypilot>=0.12.2"
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.12.1...v0.12.2
Nothing published for this version
`bash uv pip install "skypilot>=0.12.1" `
Get it now with:
uv pip install "skypilot>=0.12.1"
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.12.0...v0.12.1
Nothing published for this version
Fix Zip Slip vulnerability in /upload endpoint (#8723).
SkyPilot v0.12.0 brings major new capabilities: Slurm integration for running SkyPilot on existing Slurm clusters, Job Groups for heterogeneous parallel workloads like RL training, an Agent Skill that teaches AI coding agents to use SkyPilot, Recipes for sharing reusable YAML templates across teams, and significant Pool enhancements including autoscaling. This release also drops Python 3.7/3.8, with Python 3.9+ now required.
Get it now with:
uv pip install "skypilot[all]>=0.12.0"
Or, upgrade your team SkyPilot API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.12.0
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Deprecations & Breaking Changes:
sky.jobs.queue(version=1) is deprecated and will be removed in v0.13. Use sky.jobs.queue(version=2) instead. The new version returns richer job metadata as dictionaries (#9118).SkyPilot now supports connecting your Slurm clusters, bringing its unified interface to one of the most widely used job schedulers in high-performance computing (#5491, #8198, #8219, #8268, #8291, #8470, #8604, #8729, and 25+ additional PRs). Users can launch SkyPilot clusters and managed jobs on Slurm clusters with the same CLI and YAML they use for cloud and Kubernetes, enabling seamless workload portability across all AI infra.
<p align="center"> <img width="80%" alt="SkyPilot dashboard showing Slurm GPU availability" src="https://github.com/user-attachments/assets/2790a276-1fd7-4edb-8643-e04200236dae" /> </p>
Key capabilities include:
sky show-gpus for Slurm partitionssbatch_options in task YAML# Launch on a Slurm cluster
resources:
accelerators: H100:8
infra: slurm
# View GPU availability across Slurm clusters
sky show-gpus --infra slurm
# Launch a training job on Slurm
sky launch --infra slurm/my-cluster train.yaml
SkyPilot now ships an official Agent Skill that teaches AI coding agents—Claude Code, Codex, and others—how to use SkyPilot (#8823, #9017, #9037). With the skill installed, your agent can launch clusters, run managed jobs, serve models, compare GPU pricing, and manage cloud resources—all through natural language.
Install the skill (docs) by telling your agent:
Fetch and follow the install guide https://github.com/skypilot-org/skypilot/blob/HEAD/agent/INSTALL.md
In our Scaling Autoresearch blog post, we gave Claude Code the SkyPilot agent skill and access to a 16-GPU Kubernetes cluster. Over 8 hours, the agent autonomously submitted ~910 experiments in parallel, achieving a 9x speedup over sequential search—and even discovered hardware-specific optimizations on its own.
<p align="center"> <img width="60%" alt="Scaling Autoresearch with SkyPilot Agent Skill" src="https://blog.skypilot.co/scaling-autoresearch/assets/banner.png" /> </p>
Example interactions your agent can now handle:
| Capability | Example Prompt |
|---|---|
| Launch dev clusters | "Launch a cluster with 4 A100 GPUs. Auto-stop after 30 min idle." |
| Fine-tune models | "Fine-tune Llama 3.1 8B on my dataset at s3://my-data. Use spot instances." |
| Distributed training | "Run PyTorch DDP training across 4 nodes with 8 H100s each." |
| Serve models | "Deploy Llama 3.1 70B with vLLM. Autoscale 1-3 replicas based on QPS." |
| Compare pricing | "What's the cheapest 8x H200 across AWS, GCP, Lambda, and CoreWeave?" |
| Multi-cloud failover | "Submit jobs that try our Slurm cluster first and fall back to AWS." |
SkyPilot Job Groups let you define multiple tasks with different resource requirements that run together as a single managed job (#8456, #8664, #8686, #8688, #8713, #8940). This is ideal for reinforcement learning workflows where training, inference, and data servers need different hardware—H100s for policy training, cheaper GPUs for rollout inference, and high-memory CPUs for replay buffers.
SkyPilot provisions all resources together, configures networking automatically with built-in service discovery ({task_name}-{node_index}.{job_group_name}), and manages the full lifecycle as one unit.
# Multi-document YAML for a Job Group
---
name: rl-training
execution: parallel
primary_tasks: [ppo-trainer]
---
name: data-server
resources:
cpus: 4+
run: |
python data_server.py --port 8000
---
name: ppo-trainer
resources:
accelerators: H100:1
run: |
python ppo_trainer.py --data-server data-server-0.rl-training:8000
<p align="center"> <img width="60%" alt="Job Groups architecture" src="https://blog.skypilot.co/job-groups/assets/banner.png" /> </p>
Recipes allow teams to store and share SkyPilot YAMLs in a centralized, team-accessible registry (#8755, #8825, #8851, #8873, #8876, #8882). Launch standardized workloads directly from the CLI or dashboard without local YAML files, reducing DevOps overhead and ensuring consistent configurations across the team.
Recipes support clusters, managed jobs, pools, and SkyServe, with built-in validation that blocks local file dependencies to ensure portability.
# Launch a recipe directly
sky launch recipes:dev-cluster
# Launch with custom overrides
sky launch recipes:gpu-cluster --cpus 16 --gpus H100:4 --env DATA_PATH=s3://my-data
<p align="center"> <img width="80%" alt="Recipes dashboard" src="https://blog.skypilot.co/skypilot-recipes/images/create-recipe.png" /> </p>
SkyPilot Pools receive major upgrades with autoscaling, multiple jobs per worker, heterogeneous pools, and memory-aware scheduling (#8483, #8192, #8315, #8279, #8509, #7891).
# autoscaling-pool.yaml
pool:
min_workers: 0
max_workers: 10
resources:
accelerators: H100
setup: |
echo "Setup complete!"
The SkyPilot dashboard receives significant performance improvements and new observability features (#8523, #8718, #8651, #8534, #8539).
Deploy-mode API servers now auto-enable consolidation mode (#9090), which runs managed job controllers on the API server itself instead of launching separate controller VMs. This eliminates controller overhead costs and simplifies managed job operations for team deployments.
Parallel uploads are now the default for MOUNT_CACHED file mounts, delivering a 7x speedup — flush time dropped from 151s to 21s for a ~14.6 GB test workload (#8455). Advanced tuning options are available via data.mount_cached config (#8810, #8831, #8994).
<p align="center"> <img width="50%" alt="mount-cached-speedup" src="https://github.com/user-attachments/assets/03dab45c-3694-44ba-948e-166ea8a6450b" /> </p>
SkyPilot now automatically configures Elastic Fabric Adapter (EFA) on EKS, delivering ~78.8 GB/s inter-node bandwidth (vs ~4.1 GB/s without EFA) — critical for distributed training performance (#8557, #8771):
resources:
network_tier: best
<p align="center"> <img width="50%" alt="efa-speedup" src="https://github.com/user-attachments/assets/e28b3d47-5378-41b6-a2eb-b8faf496f21a" /> </p>
A new autostop hook mechanism allows running custom scripts before a cluster is automatically stopped — for example, to save checkpoints, sync W&B runs, or send notifications (#8412):
resources:
autostop:
idle_minutes: 10
hook: |
wandb sync
curl -X POST $SLACK_WEBHOOK -d '{"text": "Cluster shutting down"}'
hook_timeout: 300
The SkyPilot dashboard now automatically detects W&B links and other external URLs generated by your AI workloads — no more digging through job logs (#8405).
<p align="center"> <img width="60%" alt="External Links" src="https://github.com/user-attachments/assets/26c38623-60e1-4994-aab6-24d1f29249b6" /> </p>
Users can now specify exit codes that trigger automatic job recovery in managed jobs, useful for transient failures with known error codes (#8324):
resources:
job_recovery:
recover_on_exit_codes: [29]
SkyPilot now auto-detects WSL and seamlessly configures VSCode Remote-SSH for Windows users (#8669).
/upload endpoint (#8723).sky api login (#8590).sky api stop hang on zombie processes in containers (#8839)./users endpoint (#8973).AUTOSTOPPING status to UP for old clients (#9073).fullnameOverride support (#8528), RWX persistent storage with RollingUpdate (#8537), unified ingress resource (#8532), disabling basic auth middleware (#8694), CoreWeave credentials (#8200), DigitalOcean credentials (#7931), SSH node pool config (#8249), Slurm credentials (#8729), AWS config mounting (#8827).remote_identity override in task config (#8659).sky show-gpus and infra dashboard (#8222).skypilot-workspace) for filtering (#8893) and user annotations on k8s pods (#9065).allowed_contexts is set (#8821).sky launch (#8589) and revamped volume background refresh (#8524).MOUNT_CACHED with configurable sequential fallback (#8455).MOUNT_CACHED configurations for workload-specific tuning (#8810, #8831, #8994).vastdata://) (#9041, #9072).--graceful flag for cluster storage operations (#8753).provision.install_conda config to disable conda installation on provisioned nodes (#8662).SKYPILOT_USER environment variable available in jobs (#8747)..sky directory location—put .sky somewhere other than $HOME (#8153).rclone flush script (#8930).sky cancel reliability (#8203), uv >=0.10.5 stripping execute permissions on XFS (#8904), WebSocket SSH proxy timeout under concurrent connections (#9001), openssh version and SetEnv compatibility (#9075).sky down failure (#8460).WORKDIR set to site-packages causing setup failures (#8378).--no-owner --no-group for both uploads and downloads (#8556).RuntimeError: dictionary changed size during iteration in ContextualEnviron (#8962).internal_external_ips and internal_services to ManagedJobRecord for richer job metadata (#8735).AWS
LocalDisk feature: Access NVMe instance storage on supported instance types (#8762, #8807, #8661).GCP
Azure
use_internal_ips, vpc_name, ssh_proxy_command (#8986).Nebius
subnet_id (#8814) and disk tiers (#8905).Vast
RunPod
New Clouds
Other
lstrip('ssh-') bug (#8417), autoscaler config polluting SSH node pools (#9012).Admin Control
UserRequest (#8741).Dashboard
CLI
sky gpus command group with list and label subcommands (#8691).--secret-file option for passing secrets from files (#8646).sky show-gpus (#8576), fix pandas >=3.0.0 compatibility (#8643).ssh even if Ray cluster is in bad state (#8649).sky serve status fails (#8406).--all flag for sky jobs pool status (#8392).volume from sky -h (#8228).sky api info (#9076).sky serve (#8141).Thank you to all contributors who made this release possible!
@aflah02, @alex000kim, @andylizf, @atoniolo76, @aylei, @bilelomrani1, @Bokki-Ryu, @cblmemo, @cg505, @concretevitamin, @dan-blanchard, @DanielZhangQD, @haimmarko-lgtm, @hentt30, @huksley, @ibrahimnd2000, @Jayachander123, @Jobarion, @kevinmingtarja, @kevinzwang, @kyuds, @liuwb, @lloyd-brown, @lucamanolache, @m-braganca, @Michaelvll, @nakinnubis, @ngi, @oelachqar, @oliviert, @otutukingsley, @panf2333, @Philmod, @php-workx, @qicz, @rohansonecha, @romilbhardwaj, @SalikovAlex, @seahyinghang8, @SeungjinYang, @smwaqas89, @williamsnell, @wurambo, @zpoint
New Contributors:
Special thanks to the community for bug reports, feature requests, and pull requests that helped improve SkyPilot!
For a complete list of changes, see the commit history.
Nothing published for this version
Nothing published for this version
Zip Slip vulnerability fix — patched path traversal in /upload endpoint
SkyPilot v0.11.2 delivers Slurm support in Beta, JobGroups for heterogeneous parallel workloads, and significantly enhanced Pools with autoscaling, multi-job scheduling and heterogeneous GPU support. This release also brings Autostop Hooks, 7x MOUNT_CACHED mode uploads speed up, automatic EFA on EKS, and numerous admin, security, and performance improvements.
Get it now with:
uv pip install "skypilot>=0.11.2"
Or, upgrade your team SkyPilot API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.11.2
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Breaking Change: Python 3.9+ required — Python 3.7 and 3.8 are no longer supported (#8489). Please upgrade before installing this release.
SkyPilot now supports Slurm as a new infrastructure backend, enabling users to orchestrate workloads on HPC clusters alongside cloud VMs and Kubernetes — all through the same unified interface (docs, #5491, #8138).
This release brings comprehensive Slurm capabilities:
--image-id for reproducible, GPU-accelerated workloads (#8604, #8609)ssh <cluster> drops you inside the Slurm job allocation, so nvidia-smi correctly reflects only your allocated GPUs (#8268)See the Slurm documentation for setup instructions.
<p align="center"> <img width="3024" height="792" alt="image" src="https://github.com/user-attachments/assets/0510257b-c158-4ef1-8432-e7a619a679bc" /> </p>
JobGroups enable running multiple jobs with different resource requirements together as a managed group (blog, #8456). Define multi-task pipelines in a single multi-document YAML file:
# Header: job group metadata
name: rl-training
execution: parallel
primary_tasks: [trainer]
termination_delay: 30s
---
name: trainer
resources:
accelerators: A100:8
run: python train.py
---
name: reward-server
resources:
accelerators: A100:1
run: python reward_server.py
Key capabilities:
sky jobs logs --task-name <name> for viewing specific task logs<p align="center"> <img width="80%" alt="JobGroups banner" src="https://blog.skypilot.co/job-groups/assets/banner.png" /> </p>
Example applications included: RL post-training (RLHF) pipeline and parallel train-eval pipeline.
SkyPilot Pools receive significant upgrades in this release:
min_workers, max_workers, and target queue length; a QueueLengthAutoscaler handles the rest while protecting running jobs from cancellation (#8483)any_of resource configurations and the scheduler dynamically resolves to available hardware (#8315):resources:
any_of:
- accelerators: T4:1
- accelerators: A100:1
--num-jobs argument (#7891)SkyPilot dashboard now automatically detects your W&B links generated by your AI workloads. No need to dig into the job logs to figure out where your training panels are. (#8405)
<p align="center"> <img width="60%" alt="image 1" src="https://github.com/user-attachments/assets/26c38623-60e1-4994-aab6-24d1f29249b6" /> </p>
An autostop hook mechanism allow running custom scripts before a cluster is automatically stopped — for example, to save checkpoints, sync W&B runs, or send Slack notifications (#8412):
resources:
autostop:
idle_minutes:10
hook:|
wandb sync
curl -X POST $SLACK_WEBHOOK -d '{"text": "Cluster shutting down"}'
hook_timeout:300
sky logs --autostop command to view hook execution logssky exec is rejected on AUTOSTOPPING clusters; sky launch waits for autostop to complete before restartingSkyPilot now automatically configures Elastic Fabric Adapter (EFA) on EKS with a single flag (#8557):
resources:
network_tier: best
This automates what was previously a complex manual setup, delivering ~78.8 GB/s inter-node bandwidth (vs ~4.1 GB/s without EFA), critical for distributed training performance. EFA interfaces are allocated proportionally to the requested GPU count.
<p align="center"> <img width="50%" alt="efa-speedup" src="https://github.com/user-attachments/assets/e28b3d47-5378-41b6-a2eb-b8faf496f21a" /> </p>
Parallel uploads are now the default for MOUNT_CACHED file mounts, delivering a 7x speedup — flush time dropped from 151s to 21s for a ~14.6 GB test workload (#8455). A new data.mount_cached.sequential_upload config option allows reverting to sequential uploads if needed.
<p align="center"> <img width="50%" alt="mount-cached-speedup" src="https://github.com/user-attachments/assets/03dab45c-3694-44ba-948e-166ea8a6450b" /> </p>
Users can now specify exit codes that trigger automatic job recovery in managed jobs (#8324):
resources:
job_recovery:
recover_on_exit_codes:[29]
When a job exits with a specified code, SkyPilot automatically recovers it — useful for transient failures with known error codes.
Automatically detect that SkyPilot is running in WSL, and seamless set up VSCode Remote-SSH for Windows users (#8669)
ray-node container (#8353, #8444)kubernetes.set_pod_resource_limits — set pod CPU/memory limits relative to requests for pod resource limit enforcement (#8644)NOT_READY status — actionable error messages for PVC issues; background refresh daemon; --refresh flag (#8524)remote_identity override (#8659), not-ready node exclusion (#8172), GKE autoscaler compatibility (#8326)/upload endpoint (#8723)sky api login — works around Chrome Private Network Access restrictions (#8590)SERVICE_ACCOUNT (#8386)fullnameOverride (#8528)p5e.48xlarge H200 and Melbourne region support (#8465, #8055)create_instance_kwargs (#8536), Vast.ai SSH fix (#8614)--secret-file CLI flag — avoid shell history exposure (#8646)provision.install_conda config — skip Miniconda for faster launches (#8662).sky location (#8153), restore autostop on start (#8022), multiple skylets per host (#8156)greenlet import fix (#8653), Resources.copy() fix (#8648), pandas compatibility (#8643), Docker MOTD fix (#8632)lstrip('ssh-') bug fix (#8417), package distribution fix (#8508)Thank you to all contributors who made this release possible!
@aflah02, @alex000kim, @andylizf, @aylei, @Bokki-Ryu, @brianstrauch, @cblmemo, @cg505, @concretevitamin, @cwhitak3r, @dan-blanchard, @DanielZhangQD, @Elden123, @funkypenguin, @hentt30, @isagi-y22, @JiangJiaWei1103, @kevinmingtarja, @koaning, @kyuds, @laimis9133, @liuwb, @lloyd-brown, @lucamanolache, @m-braganca, @Maknee, @Michaelvll, @mk0walsk, @mmcclean-aws, @mt5225, @nakinnubis, @oelachqar, @otutukingsley, @Philmod, @php-workx, @qicz, @rohansonecha, @romilbhardwaj, @sachdva, @SalikovAlex, @seahyinghang8, @SeungjinYang, @YashIIT0909, @yurekami, @atoniolo76, @zpoint
New Contributors:
Special thanks to the community for bug reports, feature requests, and pull requests that helped improve SkyPilot!
For a complete list of changes, see the commit history.
Nothing published for this version
Nothing published for this version
Nothing published for this version
This patch release is a minor bump from v0.11.1 to to get you the latest fixes:
This patch release is a minor bump from v0.11.1 to to get you the latest fixes:
Install the release candidate:
# Select needed clouds.
uv pip install 'skypilot[kubernetes,aws,gcp]==0.11.2rc1'
Upgrade your remote API server:
NAMESPACE=skypilot # TODO: change to your installed namespace
RELEASE_NAME=skypilot # TODO: change to your installed release name
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--version 0.11.2-rc.1 \
--reset-then-reuse-values \
--set apiService.image=null # reset to default image
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.11.1...v0.11.2rc1
This patch release is a minor bump from v0.11.0 to to get you the latest fixes:
This patch release is a minor bump from v0.11.0 to to get you the latest fixes:
sqlite3.OperationalError: database is locked
See the full v0.11 release notes for everything new in SkyPilot v0.11!
Deprecation: rbac.rules → rbac.namespaceRules; ingress.nodePortEnabled → ingress-ngnix.controller.service.type=NodePort (see details) helm chart
SkyPilot v0.11.0 delivers major new features: Pools and Managed Jobs Consolidation Mode; significant improvement enterprise-readiness at large scale: supports hundreds of AI engineers with a single API server instance, avoid OOM, >10x performance improvement on many requests, additional observability, and more; and UX improvements: Templates, Python SDK, CI/CD, Git Support, and more.
Get it now with:
uv pip install "skypilot>=0.11.0"
Or, upgrade your team SkyPilot API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.11.0
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
SkyPilot supports spawning a pool that launches a set of workers across many clouds and Kubernetes clusters (docs, #6260, #6426, #6459, #6552, #6591, #6665, #6675, #7332, #7963, #8008, #7876, #8047, #7930,#8039, #7846, #7855, #7919, #7920). Jobs can be scheduled on this pool and distributed to workers as they become available.
Key benefits include:
Learn more in our blog post.
<p align="center"> <img width="80%" alt="batch_inference_architecture" src="https://github.com/user-attachments/assets/964a7c82-5c25-4848-a40f-30bb3ab65c45" /> </p>
Consolidation Mode is general available (#7122, #7127, #7396, #7459, #7498, #7560, #7601, #7619, #7717, #7720, #7847, #8082, #8021, #8106). This enables:
# config.yaml
jobs:
controller:
consolidation_mode: true # Currently defaults to False.
<p align="center"> <img width="80%" alt="Frame 5" src="https://github.com/user-attachments/assets/62873c90-7634-4e58-85e5-f3cf13b01d31" /> </p>
We have significantly optimized the Managed Jobs controller, allowing it to handle 2000+ parallel jobs on a single 8-CPU controller—an 18x improvement in job capacity with the same controller size (#7051, #7371, #7379, #7408, #7432, #7473, #7487, #7488, #7494, #7519, #7585, #7595, #7945, #7966, #7979,#8036, #8095).
<p align="center"> <img width="80%" alt="image 1" src="https://github.com/user-attachments/assets/50a0127c-e3d6-43dd-a898-830c58fd9610" /> </p>
<p align="center"> <img width="80%" alt="image_1" src="https://github.com/user-attachments/assets/56d4b8b9-c394-4b37-90f2-99e7586bb851" /> </p>
<p align="center"> <img width="80%" alt="image_2" src="https://github.com/user-attachments/assets/243d63a5-98eb-44df-a5db-ef6a187c9e92" /> </p>
<p align="center"> <img width="80%" alt="image_3" src="https://github.com/user-attachments/assets/81fc122c-ea77-49b9-ad62-09e0da58e809" /> </p>
<p align="center"> <img width="80%" alt="image_4" src="https://github.com/user-attachments/assets/f8b5c1ce-9ea9-4c36-8d3b-274987295961" /> </p>
UX, Robustness, Performance Improvement
Volume Support for Existing PVC and Ephemeral Volumes (#7915, #7971,#8179)
Exising PVC: Reference pre-existing Kubernetes PersistentVolumeClaims as a SkyPilot volume (#7915).
# volume.yaml
name: existing-pvc-name
type: k8s-pvc
infra: k8s/context1
use_existing: true
config:
namespace: namespace
Ephemeral Volumes: automatically create volumes when a cluster is launched and deleted when the cluster is torn down, making them ideal for temporary storage across multiple nodes, such as caches and intermediate results.
# task.sky.yaml
file_mounts:
/mnt/cache:
size: 100Gi
CoreWeave Integration Announcement
CoreWeave now officially support SkyPilot (#6386, #6519, #6895, #7756, #7838). This integration provides: Infiniband Support, Object Storage, Autoscaling.
See the announcement on Coreweave blog. <p align="center"> <img width="80%" alt="image_5" src="https://github.com/user-attachments/assets/f6679b23-4aeb-44be-95e4-79640a17d9a8" /> </p>
AMD GPU Support Announcement
SkyPilot now fully supports AMD GPUs on Kubernetes clusters (#6378, #6944). This includes: GPU Detection and Scheduling, Dashboard Metrics, ROCm Support.
See the announcement on AMD Rocm blog. <p align="center"> <img width="80%" alt="image_6" src="https://github.com/user-attachments/assets/b4f6c078-9d65-4a72-86b5-8a36e9e51e1a" /> </p>
More Clouds Support
SkyPilot Templates
SkyPilot now ships predefined YAML templates for launching clusters with popular frameworks and patterns. Templates are automatically available on all new SkyPilot clusters. (#7935, #7965)
You can now launch a multi-node Ray cluster by adding a single line to your YAML’s run block:
run: |
# One-line setup for a distributed Ray cluster
~/.sky/templates/ray/start_cluster
# Submit your job
python train.py
Programmability: SkyPilot Python SDK
SkyPilot Python SDK is significantly improved with:
<p align="center"> <img width="80%" alt="image_7" src="https://github.com/user-attachments/assets/2fda4419-ddd8-437d-a576-8b3114ca5a7e" /> </p>
logs = sky.tail_logs(cluster_name, job_id, follow=True, preload_content=False)
for line in logs:
if line is not None:
if 'needle in the haystack' in line:
print("found it!")
break
logs.close()
resource_config = user_request.task.get_resource_config()
resource_config['use_spot'] = True
user_request.task.set_resources(resource_config)
Integrating with CI/CD
With the improved SDK, we can integrate SkyPilot with GitHub Actions and other orchestrator to spin off your AI workloads automatically. (#7932) <p align="center"> <img width="80%" alt="image_8" src="https://github.com/user-attachments/assets/5371b954-466a-4613-a726-ce8f856fb860" /> </p>
Native Git Support
You can now use your private git repositories directly as your SkyPilot workdir (#6294, #6257, #6268). SkyPilot handles the cloning and syncing automatically. You can also use --git-url and --git-ref options with sky serve up (#8012).
# task.sky.yaml
workdir:
url: <https://github.com/my-org/my-repo.git>
ref: 1234ab # commit hash or branch name
You can also find the commit hash of your workdir in the Dashboard: <p align="center"> <img width="80%" alt="image_1 1" src="https://github.com/user-attachments/assets/c9632f54-0693-464c-9956-f775668559ed" /> </p>
Autostop/Autodown Based on SSH Sessions
In addition to running jobs on clusters, you can now configure autostop/autodown to wait for active SSH sessions (#6361, #6485).
# task.sky.yaml
resources:
autostop:
wait_for: jobs_and_ssh
We released high-performance distributed training examples for large models with checkpointing support (#6525, #6551, #6242, #6273, #6443).
<p align="center"> <img width="80%" alt="image_2 1" src="https://github.com/user-attachments/assets/3ace5534-d289-476a-b4b4-c6b73c6925f4" /> </p>
rbac.rules → rbac.namespaceRules; ingress.nodePortEnabled → ingress-ngnix.controller.service.type=NodePort (see details) helm chartpost_provision_runcmd (#7943).allowed_contexts configuration (#7878).--retry-until-up: Fixed the -retry-until-up flag to properly retry failed launches across all cloud zones (#8079).sky logs --provision during cluster launches (#7888).sky show-gpus, fix the default Neuron-based AMI (#7896, #7958).AWS_CONFIG_FILE environment variable for custom credential paths (#8050).SSH Node Pools
pod_config and custom_metadata specifications for multiple SSH node pools through context_configs (#7660, #7913).provision_timeout field for controlling SSH node provisioning wait times (#7660).SecretStr to prevent accidental exposure in logs (#8040).Task.get_resources_config() returns resources as a dictionary matching YAML format, and Task.set_resources() now accepts dictionaries—making programmatic resource manipulation as intuitive as editing YAML (#7857).SKYPILOT_SETUP_NUM_GPUS_PER_NODE environment variable available during setup phase for configuring software based on GPU count (#7092).sky serve logs (#8002)sky status --kubernetes for the team SkyPilot API server (#7989), removed sky local up --ips in favor of sky ssh up (#8065)Thank you to all contributors who made this release possible!
@adocherty, @alex000kim, @aylei, @brianstrauch, @cblmemo, @cg505, @coopslarhette, @cwhitak3r, @DanielZhangQD, @hyoxt121, @kevinmingtarja, @kyuds, @lloyd-brown, @massaindustries, @Michaelvll, @rohansonecha, @romilbhardwaj, @SalikovAlex, @seongsukwon-moreh, @SeungjinYang, @zpoint
New Contributors:
Special thanks to the community for bug reports, feature requests, and pull requests that helped improve SkyPilot!
For a complete list of changes, see the commit history.
Nothing published for this version
This is a preview of the upcoming 0.11.0 release!
This is a preview of the upcoming 0.11.0 release!
Install the release candidate:
# Select needed clouds.
uv pip install 'skypilot[kubernetes,aws,gcp]==0.11.0rc1'
Upgrade your remote API server:
NAMESPACE=skypilot # TODO: change to your installed namespace
RELEASE_NAME=skypilot # TODO: change to your installed release name
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--version 0.11.0-rc.1 \
--reset-then-reuse-values \
--set apiService.image=null # reset to default image
If you try out 0.11.0rc1, please let us know in Slack!
Improved CI reliability by fixing flaky tests, enhancing dependency testing, and tuning vulnerability scanning (#6991, #7047, #7049, #7064, #7103, #71…
This release focuses on production stability, performance optimization, and fixing critical bugs that affected reliability in multi-user and Kubernetes environments. Key highlights include major performance improvements for managed jobs and the dashboard, resolution of race conditions and resource leaks, and expanded cloud/accelerator support.
This release includes 400+ merged pull requests with high-priority critical fixes and additional improvements spanning bug fixes, performance enhancements, new features, and comprehensive documentation updates.
<div align="center"> <img width="89%" alt="jobs-controller" src="https://github.com/user-attachments/assets/fe7db06f-e236-4b5e-8fbc-b149b0cc2379" /> </div>
consolidation_mode (#7127, #7396, #7459, #7122, #7498, #7601, #7560, #7619, #7717, #7720).
sky jobs launch# config.yaml
jobs:
controller:
consolidation_mode: true # Currently defaults to False.
<div align="center"> <img width="80%" alt="image 1" src="https://github.com/user-attachments/assets/3d1ec48f-cbe6-4e5f-b7ad-3afc06d627e8" /> </div>
<div align="center"> <img width="80%" alt="Screenshot 2025-11-11 at 11 51 07 AM" src="https://github.com/user-attachments/assets/f592c5b4-1c4b-4368-bf6a-4a4f8ba11177" /> </div>
import sky
from sky import jobs as managed_jobs
for i in range(100):
resource = sky.Resources(accelerators='A100:8')
task = sky.Task(resources=resource,
workdir='.',
run='python batch_inference.py')
managed_jobs.launch(task, name=f'hello-{i}')
<div align="center"> <img width="80%" alt="image" src="https://github.com/user-attachments/assets/af66c3ef-3a7e-4ecc-9a2d-bb09ee7cec91" /> </div>
<div align="center"> <img width="80%" alt="image 4" src="https://github.com/user-attachments/assets/11a336fb-ccb9-4df9-8ba1-48245e95dcee" /> </div>
<div align="center"> <img width="80%" alt="image 5" src="https://github.com/user-attachments/assets/e9acc7c9-6292-417d-ae75-9d1b7043a7d8" /> </div>
sky status and API endpoints, improving mean response time by 10-35% (#7689, #7690, #7665, #7705, #7708).orjson for API serialization, improving tail latency for large payloads (#7734)./dev/shm exhaustion) (#7678) and fusermount-server leaks (#7398).BrokenProcessPool crashes when canceling sky logs (#7607).curl | sh pattern for fluent-bit installation with official package repositories (#7126).sky down output to show a clear summary when multiple clusters fail or succeed (#7225, #7635).sky status) to show activity when the server is under load (#7631).sky ui as a shorter alias for sky dashboard and sky volume as an alias for sky volumes (#7565, #7746).stream_and_get can now retrieve the latest request without an ID (#6965).sky logs --provision <cluster> command (#7682).SKYPILOT_NUM_JOBS environment variable for pools to complement SKYPILOT_JOB_RANK (#7542).sky jobs apply --workers <n> (#7236).sky local up --name (#7244, #7394).sky storage delete CLI command which failed to delete cloud buckets (#7279).yum (#7740).New Cloud Provider Support
Kubernetes Improvements
allowed_contexts: all to simplify setup in controlled environments (#7196)..failed file generation to improve setup debugging (#7207, #7208, #7218, #7219).AWS
use_ssm: false configuration (#7387).Other Clouds
/etc/hosts (#7773).AsyncFileLock that occurred when async operations were cancelled (#7627, #7664).server-log in Helm deployments for better observability (#7226).connection already closed error when using PostgreSQL (#7584).sky jobs logs would fail with ClusterNotUpError due to stale cache (#7585).skypilot[aws] (#7646).MOUNT_CACHED (#7421).Dict return types with strongly-typed Pydantic models for SDKs (#6833, #6847, #7404).CommandGen feature, simplifying the Task API (#7801).Install or upgrade to v0.10.5 using pip:
pip install -U "skypilot[all]==0.10.5"
Or with uv:
uv pip install -U "skypilot[all]==0.10.5"
For API server deployments using Helm:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.10.5
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Thank you to all contributors who made this release possible! 🎉
@bobokvsky, @coopslarhette, @massaindustries, @mluogh, @eric-czech, @pokgak, @Elden123, @jmalukaite, @SamuelMarks, @brianstrauch, @lynnliu030, @EricBryann, @Michaelvll, @romilbhardwaj, @concretevitamin, @andylizf, @alex000kim, @cg505, @kevinmingtarja, @zpoint, @aylei, @rohansonecha, @SeungjinYang, @DanielZhangQD, and the entire SkyPilot community.
Special thanks to the community for bug reports, feature requests, and pull requests that helped make this release more robust and feature-rich. This release includes 400+ merged PRs from dozens of contributors! Also, thanks @alex000kim for helping with this release note!
For a complete list of all changes: https://github.com/skypilot-org/skypilot/compare/v0.10.3...v0.10.5
Nothing published for this version
This patch release is a minor bump over v0.10.3.post1 to fix robustness against Coreweave clusters, i.e., adding additional retries and fallback for t
This patch release is a minor bump over v0.10.3.post1 to fix robustness against Coreweave clusters, i.e., adding additional retries and fallback for the Kubernetes API calls:
To upgrade:
pip install "skypilot==0.10.3.post2"
# Restart your local API server
sky api stop; sky api start
Or, upgrade the API server:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.10.3
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:0.10.3.post2 \
--version $VERSION --devel --reuse-values
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.10.3.post1...v0.10.3.post2
This patch release is a minor bump over v0.10.3 to fix a dependency issue caused by a breaking change in uvicorn==0.36.0:
This patch release is a minor bump over v0.10.3 to fix a dependency issue caused by a breaking change in uvicorn==0.36.0:
uvicorn dependency to mitigate AttributeError: 'Config' object has no attribute 'setup_event_loop' error. (#7287)If you see an issue like above, upgrade your SkyPilot with:
pip install "skypilot==0.10.3.post1"
# Restart your local API server
sky api stop; sky api start
If you are using a remote API server, it should still work with v0.10.3.
SkyPilot v0.10.3 delivers 2-10x performance improvement, enhanced cloud integration, improved Kubernetes integration, and strengthened production reli
SkyPilot v0.10.3 delivers 2-10x performance improvement, enhanced cloud integration, improved Kubernetes integration, and strengthened production reliability features for AI/ML workloads across clouds.
sky status (#6883, #6858, #6868, #6871, #6882, #6940, #6948, #6892, #6908)Upgrade your SkyPilot API server to 0.10.3 and get the improvement out of the box!
More SkyPilot API server and cluster metrics can now be set up for better observability, including:
CoreWeave autoscaler is now supported, with a single config field change (#6895):
kubernetes:
autoscaler: coreweave
SkyPilot now supports Microsoft Entra ID SSO login (#7045, #7028), besides Okta and Google Workspace. See more docs for setting up SSO login for API server: https://docs.skypilot.co/en/latest/reference/auth.html#sso-recommended <img width="60%" alt="image 1" src="https://github.com/user-attachments/assets/aeebe56c-b24a-4a70-ac37-a18597b81125" />
tail_logs (#6902)Kubernetes Improvements
Nebius Cloud
list_instances (#6867)AWS
use_ssm flag for secure connections (#6830)RunPod
GPU Support
sky show-gpus (#7006)To install or upgrade:
# Using pip
pip install -U "skypilot>=0.10.3"
# Using uv (recommended)
uv pip install "skypilot>=0.10.3"
To upgrade your SkyPilot API server to version 0.10.3, use the following Helm command:
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.10.3
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Note: The --set apiService.image=berkeleyskypilot/skypilot:$VERSION is needed in case your previous API server was using the nightly build.
Thank you to all contributors who made this release possible! 🎉
@SeungjinYang, @cg505, @SalikovAlex, @DanielZhangQD, @webconn, @m6xwzzz, @kevinmingtarja, @zpoint, @clayrosenthal, @lloyd-brown, @kyuds, @aylei, @romilbhardwaj, @Michaelvll, @andylizf, @Maknee, @alex000kim, @rohansonecha, @seongsukwon-moreh, @concretevitamin, and the entire SkyPilot community.
Special thanks to the community for bug reports, feature requests, and pull requests that helped make this release more robust and feature-rich. This release includes 100+ merged PRs from 30+ contributors! Also, thanks to @alex000kim for helping polishing this release note.
For the complete list of changes, see the full changelog.
SkyPilot v0.10.2 brings values in cluster management, improved Kubernetes support, programmatic SDK, and numerous stability and performance enhancemen
SkyPilot v0.10.2 brings values in cluster management, improved Kubernetes support, programmatic SDK, and numerous stability and performance enhancements for production use from teams with a large number of workloads.
Get it now with:
uv pip install "skypilot>=0.10.2"
Get provisioning logs with the new --provision option for sky logs (#6638):
sky logs --provision <cluster-name>
Find the detailed reason for your cluster failure, e.g., OOM, with the cluster events (#6590, #6593, #6615, #6620, #6621, #6667, #6658, #6609, #6617):
SkyPilot introduces a new preload_content option for tail_logs to enable processing logs while streaming.
logs = sky.tail_logs(cluster_name, job_id, follow=True, preload_content=False)
for line in logs:
if line is not None:
if 'needle in the haystack' in line:
print("found it!")
break
logs.close()
sky down for AWS clusters with exposed ports by 4x (#6629, #6663, #6720)--config option documentation (#6794)sky api info (#6748)sky.endpoints function (#6599)load_balancing_policy: instance_aware_least_load
replica_policy:
target_qps_per_replica:
"H100:1": 2.5 # H100 can handle 2.5 QPS
"A100:1": 1.25 # A100 can handle 1.25 QPS
"A10:1": 0.5 # A10 can handle 0.5 QPS
We thank all contributors who made this release possible!
New Contributors: @miltava, @tomzx, @webconn, @hongsu-moreh, @ibpark-moreh, @nathan-liner
All Contributors: @DanielZhangQD, @cblmemo, @kyuds, @SeungjinYang, @lloyd-brown, @romilbhardwaj, @rohansonecha, @zpoint, @aylei, @SalikovAlex, @Michaelvll, @kevinmingtarja, @miltava, @Maknee, @cg505, @tomzx, @concretevitamin, @webconn, @sethkimmel3, @andylizf, @hongsu-moreh, @ibpark-moreh, @lucamanolache, @nathan-liner
Special thanks to the community for bug reports, feature requests, and pull requests that helped improve SkyPilot!
For a complete list of changes, see the commit history.
Nothing published for this version
SkyPilot v0.10.1 improves enterprise production readiness with large-scale distributed training capabilities (Llama 4 400B, OpenAI GPT-OSS), introduce
SkyPilot v0.10.1 improves enterprise production readiness with large-scale distributed training capabilities (Llama 4 400B, OpenAI GPT-OSS), introduces/enhances integration with AMD GPUs and leading GPU clouds (CoreWeave, Nebius), and delivers enhanced reliability features for mission-critical AI workloads.
Get it now with:
uv pip install "skypilot>=0.10.1"
Use your (private) git repositories as your SkyPilot workdir:
# task.sky.yaml
workdir:
url: https://github.com/my-org/my-repo.git
ref: 1234ab # commit hash or branch name
Find your commit hash of your workdir in Dashboard:
<p align="center"> <img src="https://i.imgur.com/h9U838b.png" width="80%" /> </p>
Added in (#6294, #6257, #6268).
We released high-performance distributed training examples for large models with checkpointing support (#6525, #6551, #6242, #6273).
<p align="center"> <img src="https://i.imgur.com/dAov9ud.png" width="80%" /> </p>
Agentic training example with VeRL is now also available (#6443).
AMD ROCm is now supported on Kubernetes clusters with SkyPilot!
# task.sky.yaml
resources:
infra: k8s/my-amd-cluster
image_id: docker:rocm/pytorch-training:v25.6
accelerators: MI300:4
SkyPilot now supports CoreWeave clusters with native Infiniband and object store support (#6386, #6487, #6483).
Use your CoreWeave cluster with SkyPilot:
resources:
infra: k8s/my-coreweave-cluster
network: best # Enable infiniband
B200 GPUs and spot instances are now supported on Nebius cloud (#6474, #6478, #6267).
# task.sky.yaml
resources:
infra: nebius
accelerators: B200:8
use_spot: true
Additionally, SkyPilot now supports MOUNT_CACHED mode for Nebius cloud. (#6456)
In addition to running jobs on clusters, you can also ask autostop/autodown to wait for active SSH sessions or none of them in SkyPilot YAML (#6361, #6485).
# task.sky.yaml
resources:
autostop:
wait_for: jobs_and_ssh
You can now dump your SkyPilot cluster/job logs to external logging services like AWS CloudWatch and GCP Cloud Logging (#6331, #6405, #6411, #6369). Configure it in your ~/.sky/config.yaml:
logs:
store: aws # Or 'gcp', etc.
aws:
... # Service-specific options; see below.
sky status: Intelligent caching reduces status command latency by 50% (#6166)sky api logout to logout from API server (#6284, #6327, #6412)We thank all contributors who made this release possible!
New Contributors**: @LokmaSpeedy, @jacobergzhou, @amd-pratmish, @tedspare, @jimbz, @makhalin
All Contributors: @alex000kim, @amd-pratmish, @andylizf, @aylei, @bikramnehra, @cblmemo, @cg505, @clayrosenthal, @concretevitamin, @DanielZhangQD, @jacobergzhou, @jimbz, @kevinmingtarja, @kyuds, @lloyd-brown, @LokmaSpeedy, @lucamanolache, @makhalin, @Maknee, @Michaelvll, @rohansonecha, @romilbhardwaj, @SalikovAlex, @SeungjinYang, @tedspare, @zhenjiasun, @zpoint
Special thanks to the community for bug reports, feature requests, and pull requests that helped improve SkyPilot!
For a complete list of changes, see the commit history.
--cloud/--region/--zone flags have been deprecated in favor of --infra
We are excited to announce SkyPilot 0.10! This release is the largest release by far, bringing enterprise-ready features including API server deployment in production with SSO, feature-rich dashboard, external PostgreSQL, workspace isolation and graceful upgrade, automatic network setup, and SSH Node Pools.
Get it now:
pip install -U skypilot
SkyPilot now integrates with enterprise SSO providers like Okta, Google Workspace, enabling secure authentication with automatic account creation and access control.
<img width="1675" height="1067" alt="image" src="https://github.com/user-attachments/assets/6bb3b043-c5b9-4054-af39-7468f1ccf0b6" />
Log in to the API server with SSO enabled:
$ sky api login -e https://skypilot.example.com
A web browser has been opened to http://skypilot.example.com/token. Please continue the login in the web browser.
To manually copy the token, press ctrl+c.
Logged into SkyPilot API server at: http://skypilot.example.com
└── Dashboard: http://skypilot.example.com/dashboard
Users authenticate via their organization's SSO provider, and their identities are automatically tracked across all SkyPilot resources.
SkyPilot dashboard now includes significant amount of new features:
<img width="1675" height="1067" alt="image" src="https://github.com/user-attachments/assets/c98dcbc6-a6c9-4c6e-806d-9694087e89b5" />
<img width="2948" height="1779" alt="infra-gpu" src="https://github.com/user-attachments/assets/0d2cfcd4-78ac-4304-831f-30ba01967310" />
SkyPilot 0.10 adds support for persisting API server state to an external PostgreSQL database, enabling high availability and disaster recovery for production deployments.
Configure your deployment to use a managed database service (e.g., AWS RDS, Cloud SQL) to ensure your cluster and job state survive API server restarts or migrations.
db: postgresql://myusername:mypassword@hostname:5432/database
Workspaces provides a declarative way to define isolated environments with custom cloud configurations for different teams or projects.
Configure workspaces to control which teams can access which infrastructure:
# API server config
workspaces:
research-private:
private: true
allowed_users:
- alice@skypilot.co
- mike@skypilot.co
gcp:
project_id: skypilot-research-private
aws:
disabled: true
ml-team:
gcp:
project_id: skypilot-ml-team-prod
Teams simply set their active workspace to use their workspace configuration:
# In team's .sky.yaml
active_workspace: ml-team
<img width="1675" height="1067" alt="image" src="https://github.com/user-attachments/assets/76fb347f-bfe5-41f4-a890-8b9b1a1ee021" />
SkyPilot 0.10 introduces robust graceful upgrade of API server:
<p align="center"> <img src="https://i.imgur.com/jUjXu0J.gif" alt="Graceful upgrade demo" /> </p>
network_tier: best)SkyPilot v0.10.0 can now automatically configure high-performance network with a single network_tier: best config. Supported infra:
Turn your existing machines — on-premises servers, cloud reserved instances or even your personal workstation — into SSH Node Pools to run SkyPilot clusters and jobs on them.
Configure your machines in ~/.sky/ssh_node_pools.yaml:
# ~/.sky/ssh_node_pools.yaml
my-node-pool:
hosts:
- 1.2.3.4
- 1.2.3.5
Deploy SkyPilot on them with a single command:
$ sky ssh up
$ sky launch --infra ssh/my-node-pool -- python train.py
Your machines now appear as infra choices alongside cloud providers, complete with GPU availability tracking and resource management.
--infra option to specify infrastructure instead of separate --cloud/--region/--zone flags (#5602, #5656, #5695, #5703)
--infra aws/us-west-2/us-west-2a, --infra aws/*/us-west-2b--infra k8s/my-k8s-context--infra ssh/my-ssh-pool--gpus 80GB+ (#5948)SKYPILOT_DEBUG=1 (#6121)sky cancel now supports glob patterns for cluster names (#5933)SkyPilot 0.10 brings major enhancements to Kubernetes support:
SkyPilot 0.10 includes comprehensive documentation improvements across all major features:
--cloud/--region/--zone flags have been deprecated in favor of --infra
--infra aws/us-west-2 instead of --cloud aws --region us-west-2SkyPilot 0.10.0 maintains backward compatibility with existing clusters and jobs. However, there are a few considerations:
--cloud/--region/--zone to the new --infra flagWhen upgrading from 0.9.x to 0.10.0, both API server and client need to be upgraded, i.e. API server 0.10.0 does not support 0.9.x clients.
CONTROLLER_NAME=$(sky status | grep "sky-jobs-controller" | awk '{print $1}')
sky start -f $CONTROLLER_NAME
pip install -U skypilot
Or, upgrade your existing API server to 0.10.0 (see upgrade guide):
NAMESPACE=skypilot
RELEASE_NAME=skypilot
VERSION=0.10.0
helm repo update skypilot
helm upgrade -n $NAMESPACE $RELEASE_NAME skypilot/skypilot \
--set apiService.image=berkeleyskypilot/skypilot:$VERSION \
--version $VERSION --devel --reuse-values
Note: --set apiService.image=berkeleyskypilot/skypilot:$VERSION is needed in case your previous API server was using the nightly build.
This release includes contributions from many new and returning contributors. Special thanks to everyone who helped make SkyPilot 0.10 possible!
New contributors: @lucamanolache, @ykocaogullar, @omahs, @kilavvy, @vtjl10, @mundaym, @leopardracer, @dhiaEddineRhaiem, @crStiv, @zhenjiasun, @AngadSethi, @davidknittel728, @bikramnehra, @hyoxt121, @turtlebasket, @clayrosenthal, @kevinmingtarja, and many others from the community who contributed through issues, discussions, and feedback.
Thanks to all contributors: @SeungjinYang, @zpoint, @Michaelvll, @aylei, @cg505, @romilbhardwaj, @DanielZhangQD, @kyuds, @concretevitamin, @cblmemo, @rohansonecha, @lucamanolache, @Maknee, @SalikovAlex, @kevinmingtarja, @turtlebasket, @JiangJiaWei1103, @AngadSethi, @zhenjiasun, @ykocaogullar, @vtjl10, @vnavkal, @omahs, @mundaym, @leopardracer, @kilavvy, @hyoxt121, @greendev0127, @ggilley, @funkypenguin, @dhiaEddineRhaiem, @crStiv, @colinjc, @clayrosenthal, @bikramnehra, @andylizf, @Kovbo, @KeplerC
To stay updated, star and watch our GitHub repo, follow @skypilot_org, or join our community Slack.
Full Changelog: https://github.com/skypilot-org/skypilot/commits/v0.10.0
[Core] Do not initialize conda for users if using docker image by @SeungjinYang in https://github.com/skypilot-org/skypilot/pull/5303
gcloud when installed using wget by @SeungjinYang in https://github.com/skypilot-org/skypilot/pull/5335all_users in wait for job status in back compact tests by @zpoint in https://github.com/skypilot-org/skypilot/pull/5346asyncio_default_fixture_loop_scope by @zpoint in https://github.com/skypilot-org/skypilot/pull/5348test_gcp_disk_tier by @zpoint in https://github.com/skypilot-org/skypilot/pull/5393pytest.ini to remove test warning by @zpoint in https://github.com/skypilot-org/skypilot/pull/5379api info: display dashboard on last line by @concretevitamin in https://github.com/skypilot-org/skypilot/pull/5417test_multi_echo -- change sshd config to support large number of jobs by @zpoint in https://github.com/skypilot-org/skypilot/pull/5323test_kubernetes_context_failover by @zpoint in https://github.com/skypilot-org/skypilot/pull/5455test_cancel_launch_and_exec_async by @zpoint in https://github.com/skypilot-org/skypilot/pull/5456sky check parallel by @kyuds in https://github.com/skypilot-org/skypilot/pull/5483ordered to a section by @zpoint in https://github.com/skypilot-org/skypilot/pull/5540show-gpus, and add back the 0 GPU nodes by @Michaelvll in https://github.com/skypilot-org/skypilot/pull/5490test_gcp_disk_tier by @zpoint in https://github.com/skypilot-org/skypilot/pull/5592Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.9.2...v0.9.3
This patch release is a minor bump from v0.9.1 to resolve an issue that could affect users with old clusters from v0.7 and earlier (#5439).
This patch release is a minor bump from v0.9.1 to resolve an issue that could affect users with old clusters from v0.7 and earlier (#5439).
See the full v0.9 release notes for everything new in SkyPilot v0.9!
experimental.config_overrides has been deprecated. Use the config field instead.
We're excited to announce the release of SkyPilot v0.9.1! This update brings major improvements to SkyPilot, making it faster, more powerful and flexible for production-ready deployment.
<p align="left"> <img src="https://blog.skypilot.co/client-server/images/previous-now.png" alt="Client-Server Architecture" width="400"/> </p>
The new client-server model transforms SkyPilot from a single-user system into a scalable, multi-user platform, making it easier for individuals and teams to run and manage their workloads.
SkyPilot has a new dashboard! Easily view and manage your clusters, jobs and logs.
<p align="left"> <img src="https://i.imgur.com/lEQa50B.png" alt="SkyPilot Web Dashboard" width="600"/> </p>
Access it with sky dashboard.
<p align="left"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://docs.skypilot.co/en/latest/_images/config-cheatsheet-dark.svg"> <source media="(prefers-color-scheme: light)" srcset="https://docs.skypilot.co/en/latest/_images/config-cheatsheet-light.svg"> <img src="https://docs.skypilot.co/en/latest/_images/config-cheatsheet-dark.svg" alt="New configuration system" width="400"/> </picture> </p>
SkyPilot now supports specifying configuration at various levels: CLI, SkyPilot YAML, project-level config, client-level global config and server-side config.
You can now have a project configuration storing default values for all jobs in a project, a user configuration to apply globally to all projects and Task YAML overrides for specific jobs.
mount_cached storage - 9.6x faster checkpointingNew storage mode mount_cached uses the local disk as a cache for cloud storage buckets. Boosts GPU utilizationby making cloud I/O asynchronous.
file_mounts:
/checkpoints:
source: gs://my-checkpoints-bucket
mode: MOUNT_CACHED # Will asynchronously upload all writes to the bucket
SkyPilot now supports Nebius cloud! Getting started is easy:
$ sky check nebius
$ sky launch --gpus H200:8 --cloud nebius
New native images for ARM instances allows you to run SkyPilot on your GH200s, GB200s on Lambda cloud, GCP or your own Kubernetes clusters! (#4835)
sky CLI now returns non-zero exit code on launch/exec/logs/jobs launch/jobs logs failures (#4846)
sky CLI in automated workflows.sky check now separately checks storage and compute capabilities (#4996, #4977)--all option for sky jobs queue to show all jobs (#4923)resources.gpus can now be used to alias resources.accelerators in the SkyPilot YAML (#5207)# ~/.sky/config.yaml
jobs:
controller:
# autostop: false # to disable completely
autostop:
idle_minutes: 5
down: true
sky jobs queue -u when using a shared controller (#4787)mount_cached (#4369)
.gitignore handling is now more robust (#4988)/dev/fuse access mechanism on k8s (#5028)
smarter-devices-fuse resource, making SkyPilot fuse mounting compatible on autoscaling clusters.sky check now detects and hints for unlabeled GPU nodes on GKE (#5065)initContainers can now be overriden through pod_config (#5247)sky exec now waits for the cluster to be started (#4867)sky local up --ips now supports specifying sudo password (#5030)dynamic_fallback (#4628)
spot_placer field can be set to dynamic_fallback to let SkyPilot automatically switch from spot to on-demand instances if spot instances are not available.any_of field order issue causing version bump to not work (#4978)a3-highgpu-8g) can now be directly selected from the CLI with -t flag (#5120)mount_cachedSKY_ are no longer supported. Use SKYPILOT_ env vars instead.kubernetes is no longer a valid region name. use the k8s context name to specify a kubernetes cluster if required.experimental.config_overrides has been deprecated. Use the config field instead.SkyPilot 0.9.1 introduces the asynchronous execution model, which may cause compatibility issues with user programs using SkyPilot SDKs <=0.8.1.
Refer to the migration guide to upgrade your code.
TL;DR: Wrap all SkyPilot SDK function calls (except tail_logs) with sky.stream_and_get() to make your program behave mostly the same as before:
# <= 0.8.1
job_id, handle = sky.launch(task)
# 0.9.1
job_id, handle = sky.stream_and_get(sky.launch(task))
New contributors: @kyuds, @BorenTsai, @funkypenguin, @JiangJiaWei1103, @SalikovAlex, @flaviomartins, @Ajay, @bradhilton, @SeungjinYang, @eltociear, @vvidovic, @KennBro, @DanielZhangQD
Many thanks to all contributors who contributed to this release!
Contributors: @aylei, @zpoint, @SeungjinYang, @cg505, @Michaelvll, @romilbhardwaj, @KeplerC, @concretevitamin, @SalikovAlex, @DanielZhangQD, @kyuds, @cblmemo, @andylizf, @clayrosenthal, @JiangJiaWei1103, @bradhilton, @funkypenguin, @vvidovic, @cbrownstein, @flaviomartins, @KennBro, @mjibril, @kristopolous, @Ajay, @landscapepainter, @eltociear, @BorenTsai
Nothing published for this version
This patch release is a minor bump over v0.8.0 to get you the latest fixes as soon as possible.
This patch release is a minor bump over v0.8.0 to get you the latest fixes as soon as possible.
wheel<0.46.0 to mitigate build errors when launching clusters in environments with wheel>=0.46.0 (#5153)Stay tuned for a major upgrade coming up in v0.9.0!
LocalDockerBackend is deprecated. To run locally, use sky local up to setup a local k8s cluster.
We’re thrilled to release SkyPilot v0.8.0! This update makes SkyPilot faster and more robust, with major improvements to Managed Jobs, Kubernetes support, and new cloud integrations.
sky launch on existing clusters is 5x faster when using --fast flag.# ~/.sky/config.yaml
jobs:
bucket: s3://my-bucket
load_balancing_policy field to choose from multiple policies (round_robin, least_load)file_mount/workdir.sky jobs logs has a new flag --sync-down to download logs to local machine (#4527)sky launch on existing clusters is 5x faster when using --fast flag. We have reworked the provisioning logic to be more efficient when reusing clusters (#4328, #4289)uv under the hood for 3x faster setup phase (#4414)remote_identity: NO_UPLOAD option to skip uploading credentials to the remote VM (#4307)sky check now shows enabled contexts (#4587)
<img width="400" alt="image" src="https://github.com/user-attachments/assets/1ce7155f-cedd-4e74-8be3-8d0a46ce029b" />lsof in k8s environments (#4304)sky show-gpus --cloud kubernetes now handles limited permissions gracefully (#4208)CUSTOM_GPU_RESOURCE_NAME environment variable (#4337)nvidia.com/product labels (#4511)pod_config specified in config.yaml is now validated before launching clusters (#4466)sky logs has a new --tail parameter to stream job logs (#4241)sky.jobs.launch from the Python API now returns the job id (#4620)SkyServe now supports choosing a load balancing policy to be used by the service (#4439)
service:
load_balancing_policy: round_robin # round_robin, least_load
| Policy | Description |
|---|---|
least_load |
(New default) Routes requests to replicas with the lowest current load, optimizing for latency and throughput |
round_robin |
Distributes requests evenly across all replicas in a circular order |
Improved security with TLS support on the load balancer (#3380)
You can now expose multiple ports on replicas: useful for running monitoring, UI or other services on the replicas (#4356)
latest and robustness improvements (#4581, #4411, #4457)sky local up to setup a local k8s cluster.sky spot CLI is now removed. Use sky jobs launch --use-spot to launch spot instances.New contributors: @weih1121, @clayrosenthal, @manbeardave, @bend, @nkwangleiGIT, @kristopolous, @sachiniyer, @KeplerC, @aylei, @Yisaer, @cbrownstein, @chesterli29, @sfrolich, @AlexCuadron
Many thanks to all contributors who contributed to this release!
Contributors: @romilbhardwaj, @cg505, @Michaelvll, @zpoint, @HysunHe, @cblmemo, @andylizf, @concretevitamin, @KeplerC, @yika-luo, @cbrownstein, @weih1121, @nkwangleiGIT, @aylei, @clayrosenthal, @sethkimmel3, @landscapepainter, @Conless, @sfrolich, @AlexCuadron, @shashank2000, @mjibril, @asaiacai, @chesterli29, @Yisaer, @sachiniyer, @manbeardave, @bend, @kristopolous
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.7.0...v0.8.0
All SKY_* environment variables are deprecated in favor of SKYPILOT_* variables.
We are excited to announce the release of SkyPilot v0.7.0! This release brings significant performance improvements and many new features:
sky CLIand many bug fixes and enhancements!
We have made 2-3x performance improvements across cloud providers through optimizations in our provisioning stack and the images we use.
| Cloud | Provisioning Time | Speedup |
|---|---|---|
| AWS | 1 min 10s | 3x |
| GCP | 1 min 15s | 3x |
| Azure | 2 min 16s | 2x |
| Kubernetes | 52s | 2.5x |
SkyPilot now supports short-term and long-term reservations across clouds:
SkyPilot's failover includes these reservations, so they can be combined with spot instances or any other resources/clouds to create a resilient and cost-effective infrastructure.
SkyPilot now has two new observability features on Kubernetes:
sky status --kubernetes shows all SkyPilot resources on the cluster. (#4040, #4079)$ sky status --cloud kubernetes
Kubernetes cluster state (context: mycluster)
SkyPilot clusters
USER NAME LAUNCHED RESOURCES STATUS
alice infer-svc-1 23 hrs ago 1x Kubernetes(cpus=1, mem=1, {'L4': 1}) UP
alice sky-jobs-controller-80b50983 2 days ago 1x Kubernetes(cpus=4, mem=4) UP
alice sky-serve-controller-80b50983 23 hrs ago 1x Kubernetes(cpus=4, mem=4) UP
bob dev 1 day ago 1x Kubernetes(cpus=2, mem=8, {'H100': 1}) UP
bob multinode-dev 1 day ago 2x Kubernetes(cpus=2, mem=2) UP
bob sky-jobs-controller-2ea485ea 2 days ago 1x Kubernetes(cpus=4, mem=4) UP
Managed jobs
In progress tasks: 1 STARTING
USER ID TASK NAME RESOURCES SUBMITTED TOT. DURATION JOB DURATION #RECOVERIES STATUS
alice 1 - eval 1x[CPU:1+] 2 days ago 49s 8s 0 SUCCEEDED
bob 4 - pretrain 1x[H100:4] 1 day ago 1h 1m 11s 1h 14s 0 SUCCEEDED
bob 3 - bigjob 1x[CPU:16] 1 day ago 1d 21h 11m 4s - 0 STARTING
bob 2 - failjob 1x[CPU:1+] 1 day ago 54s 9s 0 FAILED
bob 1 - shortjob 1x[CPU:1+] 2 days ago 1h 1m 19s 1h 16s 0 SUCCEEDED
sky show-gpus --cloud kubernetes shows detailed GPU availability information on the cluster. (#3816, #4085)$ sky show-gpus --cloud kubernetes
Kubernetes GPUs
GPU REQUESTABLE_QTY_PER_NODE TOTAL_GPUS TOTAL_FREE_GPUS
L4 1, 2, 4 8 8
H100 1, 2, 4, 8 16 16
Kubernetes per node GPU availability
NODE_NAME GPU_NAME TOTAL_GPUS FREE_GPUS
my-cluster-0 L4 4 4
my-cluster-1 L4 4 4
my-cluster-2 H100 8 8
my-cluster-3 H100 8 8
SkyPilot has a new admin policy mechanism (#3966) that admins can use to enforce policies on users’ SkyPilot usage. These policies apply custom validation and mutation logic to a user’s tasks and SkyPilot config.
Example policies:
In addition to S3, GCS and R2, you can now use Azure Blob Storage as a storage backend for storing and accessing data. (#3032)
ultra (#3860) for GCP and AWS.SkyPilot CLI is cleaner, simpler and even easier to parse now (#4023)
<img src="https://i.imgur.com/fg8tOYq.gif" width="600"/>
SKY_* environment variables are deprecated in favor of SKYPILOT_* variables.
SKY_* variables will be removed in v0.9.0.New Features
max_restarts_on_errors to specify the number of times SkyPilot should try to restart the job.resources:
job_recovery:
max_restarts_on_errors: 3 # Retry 3 times before marking the job as failed
SKYPILOT_NUM_NODES to fetch the number of nodes in the cluster. (#3656)experimental.config_override (#3689)experimental:
config_override:
docker:
run_options: ...
kubernetes:
pod_config: ...
provision_timeout: ...
gcp:
managed_instance_group: ...
nvidia_gpus:
disable_ecc: ...
Enhancements
docker.run_options now allows users to pass additional options when running docker containers. (#3682)Fixes
sky cancel not terminating all child processes (#3919)New Features
sky status --cloud kubernetes shows all SkyPilot resources on the Kubernetes cluster. (#4040, #4079)sky show-gpus --cloud kubernetes shows detailed GPU availability information on the cluster. (#3816, #4085)sky local up can now automatically set it up as a cluster to be used for running jobs. (#3926)kubectl logs, filebeat, etc.) to view SkyPilot job outputs.nvidia.com/gpu.product) for detecting GPU types. (#3493)
Performance improvements:
sky local up for GPUs is now ~5x faster, provisioning in 2min 30s instead of 12min (#3664)port-forward mode (#3657)Enhancements and fixes
--k8s is now a valid alias for --cloud kubernetes (#4151)apparmor, SkyPilot will now retry without requesting it. (#4176)New Features
pd-extreme disks with disk_tier: ultra (#3860)gcp.force_enable_external_ips to force enable external IPs (#3699)
Enhancements
New Features
io2 disks with disk_tier: ultra (#3860)Enhancements
: and other special characters. (#3734)New Features
--image-id (#4145)Premium_LRS disks with disk_tier: high (#3921)Enhancements
sky serve down --replica-id (#4032).skyignore support (#4038)
.skyignore file to skip uploading them to cloud storage.New contributors: @winglian, @Ultramann, @jucor, @BitPhinix, @sethkimmel3, @hyoxt121, @BabyChouSr, @wizenheimer, @gurcangercek, @shashank2000, @ckgresla, @bernardwin, @kmushegi, @Conless, @JayThomason, @colinjc, @mtaran, @Haijian06, @KrishivPiduri, @zpoint
Many thanks to all contributors who contributed to this release!
Contributors: @Michaelvll, @romilbhardwaj, @cblmemo, @landscapepainter, @asaiacai, @andylizf, @yika, @concretevitamin, @colinjc, @fozziethebeat, @MaoZiming, @JGSweets, @Ultramann, @Conless, @jucor, @wizenheimer, @Haijian06, @HysunHe, @gurcangercek, @bernardwin, @JungleCatSW, @BabyChouSr, @hyoxt121, @winglian, @sethkimmel3, @mjibril, @shashank2000, @ckgresla, @zpoint, @mtaran, @KrishivPiduri, @JayThomason, @BitPhinix, @kmushegi
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.6.0...v0.7.0
This patch release brings many improvements and fixes to SkyPilot, including major performance improvements for Kubernetes and Azure and new features
This patch release brings many improvements and fixes to SkyPilot, including major performance improvements for Kubernetes and Azure and new features for AWS and GCP.
Stay tuned for a detailed changelog coming up in v0.7.0!
The following features have been deprecated and will be removed in the next minor release:
We are excited to release SkyPilot v0.6.0! This release includes a number of new features:
sky jobs launch instead of sky spot launch.sky jobs API is identical to the sky spot API, but also supports on-demand instances.sky jobs launch or sky serve up, and SkyPilot will automatically deploy the controller on your Kubernetes cluster if available and run jobs on the cheapest available location.~/.paperspace/config.json and run sky check paperspace to get started.The following features have been deprecated and will be removed in the next minor release:
sky spot CLI: use sky jobs CLI instead.core.spot_xxx APIs: refactored to jobs.xxx.qps_lower_threshold and auto_restart in service: use target_qps_per_replica instead.SKYPILOT_TASK_ID environment variable (#3424)core.spot_xxx to jobs.xxx (#3417)New Features
Enhancements
sky show-gpus now shows realtime availability of GPUs in the cluster (#3499)remote_identity in ~/.sky/config.yaml (#3377, #3527)sky local up now also automatically installs the Nginx Ingress Controller (#3223)pod_config (#3244)
HTTP_PROXY and more! See example pod_config here.Enhancements
skypilot-user to identify the owner of the pod (#3576)New Features
resources now supports labels field to set labels (instance tags on aws, labels on gcp and k8s) on cloud resources (#3464, #3505)sky check now supports checking credentials for specific clouds, e.g. sky check aws gcp (#3229)
allowed_clouds in ~/.sky/config.yaml. (#3556)any_of or ordered fields in resources can now have clouds that are not enabled (#3567)SKYPILOT_CLUSTER_INFO, containing cluster name, cloud, region and zone is now available in all tasks (#3424)Enhancements
DEVICE_MEM in sky show-gpus (#3375)sky show-gpus (#3492)Optimizations
image_id in the resources field. (#3362)container-role IAM roles (#3503)New contributors: @MysteryManav, @JGSweets, @Harthgar, @mjkanji
Many thanks to all contributors who contributed to this release!
Contributors: @Michaelvll, @romilbhardwaj, @concretevitamin, @cblmemo, @MaoZiming, @shethhriday29, @asaiacai, @JGSweets, @mjkanji, @MysteryManav, @landscapepainter, @Harthgar, @mjibril, @dtran24, @fozziethebeat, @JungleCatSW
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.5.0...v0.6.0
Deprecate cpunode/gpunode/tpunode, hide admin
We are excited to release SkyPilot v0.5.0, where we introduce a significant amount of new features and enhancements, including:
and more!
any_of or ordered in resources), allowing users to significantly enlarge the resource pool and get higher availability.best disk tier for the best performance and cost, so you can choose the best disk for any cloud. (#2434)SkyServe is a serving system on top of SkyPilot that deploys and scales any HTTP services across one or more regions or clouds, with autoscaling, load balancing, and more.
Other Enhancements
Kubernetes support received a number of New Features and Enhancements.
sky local up (#2890)Other Enhancements
KUBECONFIG env var for config file specification (#3169)SkyPilot now supports 13 cloud providers, including 4 new provider-contributed clouds: VMWare vSphere, RunPod, Fluidstack and Cudo Compute.
New Features
Enhancements
Fixes
New Features
Enhancements
Fixes
--disk-size for Custom Machine Images (#2718)Enhancements
Fixes
sky check (#3038)New Features
sky status --endpoints CLI (#3199)sky show-gpus (#2583, #2892, #2933, #2946, #3083, #3149, #3113)--commit and --version for sky CLI (#2720, #2731, #2733)Enhancements
--disk-tier none override (#2906)sky check improvement (#3174, #3212, #3160)Fixes
sky_logs and mounting directory (#2667, #2845)sky logs with --sync-down (#2660)Deprecations
cpunode/gpunode/tpunode, hide admin (#2800)Local cloud which is now replaced by Kubernetes support (#3037, #3186)New Features
Enhancements
~/.ssh/generated/ssh instead of directly editing ~/.ssh/config (#2706, #3069)Fixes
New Features
Enhancements
Fixes
~/.sky/config.yaml for spot jobs (#2876)New Features
Enhancements
Fixes
Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.4.0...v0.5.0
New contributors: @rtalaricw, @jackyk02, @Vaibhav2001, @rohanvaidya45, @Shrinandan, @manishiitg, @amitkumarj441, @tgaddair, @aseriesof-tubes, @changxiaohui, @thams, @kishb87, @PratikKumar125, @mmcclean, @dtran24, @davidwagnerkc, @mjibril, @kbrgl, @msehsah1, @JungleCatSW, @Ying1123
Many thanks to all contributors who contributed to this release!
Contributors: @Michaelvll, @concretevitamin, @cblmemo, @romilbhardwaj, @MaoZiming, @landscapepainter, @sunny0826, @suquark, @Vaibhav2001, @infwinston, @hemildesai, @asaiacai, @Shrinandan, @kishb87, @rtalaricw, @iojw, @aseriesof-tubes, @manishiitg, @jackyk02, @mmcclean, @thams, @amitkumarj441, @rohanvaidya45, @saihtaungkham, @tgaddair, @davidwagnerkc, @PratikKumar125, @dtran24, @changxiaohui, @mjibril, @kbrgl, @msehsah1, @JungleCatSW, @Ying1123
This is a patch release to ship bug fixes faster to our users! This release includes many feature updates and bug fixes, including the new provisioner
This is a patch release to ship bug fixes faster to our users! This release includes many feature updates and bug fixes, including the new provisioner for AWS, fixing OOM and credential issues for long-running spot jobs, and some additional improvements.
Detailed changelog coming up in v0.5!
SkyPilot On-prem is now deprecated and Kubernetes will be the recommended mode of running SkyPilot on on-prem clusters.
We are excited to release SkyPilot v0.4.0, which brings a host of new features and improvements, including Kubernetes support, native container support, ability to open ports, and more.
sky check and sky launch --cloud kubernetes to run your task on Kubernetes.ports field. These ports are publicly accessible and can be used for hosting LLM inference endpoints, Jupyter notebooks, web servers, Tensorboard, and other services.setup and run commands can now directly be executed in that container. This allows you to wrap your environment in a container and run it on any cloud with SkyPilot.SkyPilot now supports 8 clouds, including community contributed support for two new clouds:
SkyPilot now also supports IBM COS buckets (#1966).
--ip flag for sky status returns the public IP address of the cluster (e.g., sky status --ip mycluster). Use this to access services such as LLM inference endpoints, jupyter notebooks and more.file_mounts can be dynamically defined with environment variables (docs, example), environment variables can be set through a dotenv file with the new --env-file flag (#2296).sky status updates for stopped clusters are 10x faster (#2288), and the job queue is more memory efficient (#1636).pip install skypilot-nightly (#1446)Below is a detailed list of changes.
sky spot dashboard: you can now see all your spot jobs in GUI (#2103, #2136)sky status can now show the head IP of the cluster with -a or --ip flags (#2305, #2563)sky down/stop/start defaults to a unique cluster if it exists and sky cancel without cluster cancels the latest task (#2325)sky check output is now friendlier with more hints for disabled clouds (#2002, #2017, #2196, #2114, #2221, #2377)sky down progress bar now reflects clusters failed to terminate (#1595, #2005)--cpus is provided (#2037)sky launch is interrupted (#2206, #2252)ports field (docs, #2210, #2477)image_id - tasks can now be run inside docker containers (docs, #1910)--clone-disk-from flag (#2098)sky launch by caching cluster IP address (#2400)sky status --refresh for STOPPED cluster is 10x faster (#2079)sky spot launch will now exclude files from .gitignore (#2018)sky storage CLI (#2063, #2177)<2.0 (#2157)>3.13, != 5.4.* to avoid issues with Cython 3 (#2256, #2514)<= 2.6.3 is supported on local machines (#2401)pycryptodome, oauth2client are no longer required (#2515)New contributors: @JGoo1, @tobi, @HysunHe, @blucz, @shethhriday29, @MaoZiming, @ksasi, @pushmatrix, @hzeng-0, @saihtaungkham, @fozziethebeat, @n10dollar, @asaiacai, @mtaku3, @gbmarc1, @alex000kim, @steve-marmalade, @xzrderek, @sunny0826.
Many thanks to all contributors who contributed to this release!
@Michaelvll, @concretevitamin, @romilbhardwaj, @cblmemo, @HysunHe, @landscapepainter, @shethhriday29, @infwinston, @alex000kim, @suquark, @sunny0826, @gbmarc1, @MaoZiming, @xzrderek, @tobi, @steve-marmalade, @saihtaungkham, @pushmatrix, @n10dollar, @mtaku3, @ksasi, @hzeng-0, @fozziethebeat, @blucz, @asiaacai, @WoosukKwon, @JGoo1, @mraheja, @iojw, @hemildesai, @ewzeng, @aviweit, @Saikrishna-Achalla, @Cohen-J-Omer
This patch release brings many bug fixes and features, including new mechanics for stop/down, callbacks for spot jobs and a critical dependency fix fo
This patch release brings many bug fixes and features, including new mechanics for stop/down, callbacks for spot jobs and a critical dependency fix for PyYAML after the release of cython 3.
Detailed changelog coming up in v0.4!
This is a patch release to ship bug fixes faster to our users! This release includes many feature updates and bug fixes, including the pedantic depend
This is a patch release to ship bug fixes faster to our users! This release includes many feature updates and bug fixes, including the pedantic dependency issue, disk cloning, file mounts, and cloud-specific improvements.
Detailed changelog coming up in v0.4!
This is a patch release to ship several important enhancements and bug fixes:
This is a patch release to ship several important enhancements and bug fixes:
Enhancements
sky launch --gpus h100
rm -rf ~/.sky/catalogs/v5/lambdaFAILED_SETUP error (#1998)Fixes
$PWD/~/sky_logs in some cases (#2009)sky spot launch --retry-until-up to make it actually retry until up (#2004)sky check has never been called (#2017)Full Changelog: https://github.com/skypilot-org/skypilot/compare/v0.3.0...v0.3.1
We are excited to release SkyPilot v0.3, the most significant release thus far in the project's history.
We are excited to release SkyPilot v0.3, the most significant release thus far in the project's history.
v0.3 focuses on:
See the release blog post for a deep-dive into highlights.
Release notes below are as compared to v0.2 (full changelog).
sky check to set it up. Docs here.sky cost-report; fine-grained optimizer; user identity; AWS SSO; private IP-only VPCs; Ray runtime is decoupled from user's Ray clusters; ...New Features
sky cost-report: show the estimated cost of launched clusters (#1301, #1621, #1780, #1680, #1788)
sky launch / YAML resources: field
--cpus support https://github.com/skypilot-org/skypilot/pull/1622--memory support https://github.com/skypilot-org/skypilot/pull/1746--disk-tier support https://github.com/skypilot-org/skypilot/pull/1812--detach-setup and --detach-run to sky launch https://github.com/skypilot-org/skypilot/pull/1379--retry-until-up, --region, --zone, and --idle-minutes-to-autostop for interactive nodes https://github.com/skypilot-org/skypilot/pull/1297sky status/sky.status() on specific clusters https://github.com/skypilot-org/skypilot/pull/1568--region in sky show-gpus https://github.com/skypilot-org/skypilot/pull/1187image_id field under resources https://github.com/skypilot-org/skypilot/pull/1384Enhancements
sky show-gpus
sky show-gpus <gpu>:<num> (same syntax as sky launch --gpus) https://github.com/skypilot-org/skypilot/pull/1924sky down -p bypass identity mismatch errors. https://github.com/skypilot-org/skypilot/pull/1892Fixes
sky {cpu,gpu,tpu}node commands correctly reuse existing cluster if possible https://github.com/skypilot-org/skypilot/pull/1787New Features
sky status (#1270, #1467, #1691)sky spot queue -a (#1655)Enhancements
sky spot launch default -r/--retry-until-up to True. https://github.com/skypilot-org/skypilot/pull/1781sky start on the spot controller resets the default autostop https://github.com/skypilot-org/skypilot/pull/1453sky spot queue displays job states with colors (#1473)sky spot queue no longer shows a cached (and possibly stale) version of the jobs (#1742)sky down on spot controller when in-progress spot jobs exist https://github.com/skypilot-org/skypilot/pull/1667FAILED_SETUP for spot jobs that fail during setup (#1479)CANCELLING for spot jobs that are being cancelled (#1785)SKYPILOT_JOB_ID the same for all recoveries of the same job https://github.com/skypilot-org/skypilot/pull/1400Fixes
-n) possibly overwriting each other https://github.com/skypilot-org/skypilot/pull/1782ssh_proxy_command if specified https://github.com/skypilot-org/skypilot/pull/1792Robustness is enhanced for TPUs in various modes: VMs, pods, spot (#1500, #1279, #1359, #1483, #1562, ...).
Enhancements
apt install ... in setup may non-deterministically fail due to APT lock being held by background unattended upgradescloud-init ensures unattended-upgrade is disabled at boot (#1949, #1954); for other clouds we kill the processes (#1347)Fixes
New Features
source of a storage mount, e.g., source: [~/mydir/myfile.txt, ~/datasets] https://github.com/skypilot-org/skypilot/pull/1311 #1677Enhancements
.git folder for cloud storage mounts https://github.com/skypilot-org/skypilot/pull/1494file_mounts destination path is a relative path, it is treated as being under workdir #1315Fixes
sky storage delete for externally deleted buckets https://github.com/skypilot-org/skypilot/pull/1875New Features
Enhancements
sky launch/startray[default]>=2.2.0,<=2.4.0 to fix some dependency conflicts with click/grpcio/protobufSKY_NUM_GPUS_PER_NODE https://github.com/skypilot-org/skypilot/pull/1337PYTHONUNBUFFERED=1 in task execution to disable python output buffer by default https://github.com/skypilot-org/skypilot/pull/1290Fixes
.ssh/config more robust (#1763, #1683)New Features
Enhancements
Fixes
Enhancements
sky check https://github.com/skypilot-org/skypilot/pull/1772skypilot-user tags to VMs on these clouds. https://github.com/skypilot-org/skypilot/pull/1593Fixes
Enhancements
skypilot-user tags to VMs on these clouds. https://github.com/skypilot-org/skypilot/pull/1593Fixes
New Features
Enhancements
mv ~/.sky/catalogs/v5/gcp ~/.sky/catalogs/v5/gcp.backup so that new catalogs will be auto-fetchedFixes
New contributors: @dongreenberg, @turian, @scruel, @vivekkhimani, @stephenbalaban, @landscapepainter, @cblmemo, @Saikrishna-Achalla, @datlife, @Cohen-J-Omer (IBM Cloud support!), @zetavg.
Many thanks to all contributors who contributed to this release!
@Michaelvll, @concretevitamin, @romilbhardwaj, @infwinston, @ewzeng, @michaelzhiluo, @WoosukKwon, @iojw, @sumanthgenz, @landscapepainter, @suquark, @dongreenberg, @cblmemo, @mraheja, @vivekkhimani, @turian, @stephenbalaban, @scruel, @lhqing, @datlife, @Saikrishna-Achalla, @Cohen-J-Omer, @zetavg
Another patch release to ship bug fixes faster to our users! This release includes many fixes, including those for managed spot and cloud specific imp
Another patch release to ship bug fixes faster to our users! This release includes many fixes, including those for managed spot and cloud specific improvements.
Detailed changelog coming up in v0.3!
This patch release brings more bug fixes, including fixes for cloud-specific networking and VPC configuration and managed spot.
This patch release brings more bug fixes, including fixes for cloud-specific networking and VPC configuration and managed spot.
Detailed changelog coming up in v0.3!
This is a patch release with lots of bug fixes across the board, including many cloud-specific networking and VPC fixes.
This is a patch release with lots of bug fixes across the board, including many cloud-specific networking and VPC fixes.
Stay tuned for a detailed changelog coming up in v0.3!
This is a patch release with several bug fixes for TPU, Spot, Onprem and Storage.
This is a patch release with several bug fixes for TPU, Spot, Onprem and Storage.
Detailed announcements will be made in 0.3.0.
Nothing published for this version
Nothing published for this version
Nothing published for this version
We are excited to release SkyPilot 0.2.0, which receives a host of new features, with many enhancements and fixes.
We are excited to release SkyPilot 0.2.0, which receives a host of new features, with many enhancements and fixes.
sky spot launch on your existing yamls!accelerators: tpu-v2-8 to accelerators: tpu-v2-32.sky bench to easily measure the performance and cost of different cloud resources for your task.A100-80GB is now available on 3 clouds. Check out sky show-gpus -a for GPU prices.New Features
--no-setup option to sky launch to allow for remounting of files without running setup commands again #1184sky start --all to start all clusters #1065sky storage delete #1117--no-follow option to sky logs and sky spot logs (print logs so far and exit)Enhancements
sky launch <flags> '', simply do sky launch <flags> #1191sky check automatically enable necessary GCP APIs (#1197, #1209); make it more robust for AWS checks (#1194)New Features
sky spot launch now automatically translates file_mounts in a YAML to use cloud storage. #1081 #1215
sky launch can now be launched by sky spot launch.--retry-until-up for sky spot launch; improve the responsiveness for sky spot cancel https://github.com/skypilot-org/skypilot/pull/1098$SKYPILOT_RUN_ID environment variable shared by all recoveries of the same spot job (useful for identifying it in Weights & Biases) #1196
Enhancements
spot launch with <= 0.1.2.Fixes
sky spot status -a for resources and region information https://github.com/skypilot-org/skypilot/pull/1135Enhancements
Fixes
Enhancements
sky admin deploy now automatically installs skypilot, ray (and python3 and pip3) on the local cluster under admin user #1116Fixes
Enhancements
Fixes
pip install skypilot now installs skypilot[aws] by default https://github.com/skypilot-org/skypilot/pull/1055~/.ssh/config permissions https://github.com/skypilot-org/skypilot/pull/1174New contributors
Many thanks to all contributors who contributed to this release!
@Michaelvll, @concretevitamin, @infwinston, @michaelzhiluo, @WoosukKwon, @romilbhardwaj, @sumanthgenz, @ewzeng, @iojw, @franklsf95
Nothing published for this version
This is our first release for SkyPilot -- a framework for easily running machine learning workloads on any cloud through a unified interface. No knowl
This is our first release for SkyPilot -- a framework for easily running machine learning workloads on any cloud through a unified interface. No knowledge of cloud offerings is required or expected – you simply define the workload and its resource requirements, and SkyPilot will automatically execute it on AWS, Google Cloud Platform or Microsoft Azure.
Many thanks to all those who contributed to this release! @concretevitamin @romilbhardwaj @Michaelvll @infwinston @michaelzhiluo @WoosukKwon @suquark @mraheja @gmittal @iojw @lhqing @franklsf95
Full Changelog: https://github.com/skypilot-org/skypilot/commits/v0.1.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →