NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4561 most downloaded on PyPI
Pytorch version of Stable Baselines, implementations of reinforcement learning algorithms.
Last release 3 months ago
15 Jun 2026
Release timing varies
gaps range from 2 weeks to 4 months
Nearly every release is documented
notes for 30 of 30 stable releases
3 versions withdrawn
withdrawn after publishing
6 years old
115 releases · first in 2020
One column per quarter.
Fixed deprecated error Taxi-v3 from gymnasium v1.3.0 in tests
"gymnasium>=0.29.1,<1.3.0" to "gymnasium>=0.29.1,<2.0")pandas and matplotlib are no longer core dependencies; they are now optional and only required for loading results and plotting (moved to stable-baselines3[extra]).read_json and read_csv helper functions to test filestorch minimum version from 2.3 to 2.8 to mitigate GHSA-887c-mr87-cxwpRecurrentPPO.rollout_buffer_class and rollout_buffer_kwargs arguments in PPO and OnPolicyAlgorithmJax constructors, as in Stable Baselines3. (@Trenza1ore)Full Changelog: v2.8.0...v2.9.0
Updated dependencies (pandas is now optional, gymnasium 1.3.0 support, torch>=2.8)
Relaxed Gymnasium version range (from "gymnasium>=0.29.1,<1.3.0" to "gymnasium>=0.29.1,<2.0" )
pandas and matplotlib are no longer core dependencies; they are now optional and only required for loading results and plotting (moved to stable-baselines3[extra] ).
Moved read_json and read_csv helper functions to test files
Raised torch minimum version from 2.3 to 2.8 to mitigate https://github.com/advisories/GHSA-887c-mr87-cxwp
Fixed deprecated error Taxi-v3 from gymnasium v1.3.0 in tests
Optimized tests (faster to run)
Fixed dead link for RecurrentPPO .
Added support for rollout_buffer_class and rollout_buffer_kwargs arguments in PPO and OnPolicyAlgorithmJax constructors, as in Stable Baselines3. (@Trenza1ore)
Updated Jax dependency
Optimized tests (faster to run)
Add workflow to automatically publish to PyPi when a new tag is created
Added example for using torch.compile
Fixed many broken links and updated links to https whenever possible
Nothing published for this version
Nothing published for this version
Removed support for Python 3.9, please upgrade to Python >= 3.10
strict=True for every call to zip(...)pygame-ce when installing extrasth.compile()) by updating get_parameters()pandas.concat futurewarnings occuring when dataframes are empty by removing empty frames from the list before concatenatingstrict=True for every call to zip(...)MaskablePPO and RecurrentPPO inaccurate n_updates counting when target_kl early exits the training loopRecurrentPPO and MaskablePPO forward and predict not reshaping the action before clipping it (@immortal-boy)forward() method directly in RecurrentPPO (@immortal-boy)MaskableCategorical.apply_masking() crashing with ValueError: Simplex when cached probs deviate from sum=1 in float32 with large action spaces (torch 2.9+) (@kirann-05)strict=True for every call to zip(...)env_kwargs in the hyperparam configzip_strict() is not needed anymore since Python 3.10, please use zip(..., strict=True) insteadweights_only=True (PyTorch 2.x)Full Changelog: v2.7.1...v2.8.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Warning Stable-Baselines3 (SB3) v2.7.1 will be the last one supporting Python 3.9 (end of life in October 2025) We highly recommended you to upgrade t
Warning
Stable-Baselines3 (SB3) v2.7.1 will be the last one supporting Python 3.9 (end of life in October 2025)
We highly recommended you to upgrade to Python >= 3.10.
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
RolloutBuffer and DictRolloutBuffer now uses the actual observation / action space dtype (instead of float32), this should save memory (@Trenza1ore)Sequence observation spaces when nested inside composite spaces (Dict, Tuple, OneOf) (@copilot)VecVideoRecorder where recorded_frames stayed in memory due to reference in the moviepy clip (@copilot)StopTrainingOnRewardThreshold callback message (@sea-bass)MaskablePPOenv.reset() may perform a no-op step instead of truly resetting when terminal_on_life_loss=True (default), and how to avoid this behavior by setting terminal_on_life_loss=False_sample_action() method to better explain action scaling behavior for off-policy algorithms (@copilot)pip install) of CONTRIBUTING.md more robustFull Changelog: v2.7.0...v2.7.1
get_schedule_fn() , get_linear_fn() , constant_fn() are deprecated, please use FloatSchedule() , LinearSchedule() , ConstantSchedule() instead
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
n_steps parameterfrom stable_baselines3 import SAC
# SAC with n-step returns
model = SAC("MlpPolicy", "Pendulum-v1", n_steps=3, verbose=1)
model.learn(10_000)NStepReplayBuffer that allows to compute n-step returns without additional memory requirement (and without for loops)n_steps parameterFloatSchedule and LinearSchedule classes instead of lambdas in the ARS, PPO, and QRDQN implementations to improve model portability across different operating systemslinear_schedule now returns a SimpleLinearSchedule object for better portabilityLunarLander-v2 to LunarLander-v3 in hyperparametersCarRacing-v2 to CarRacing-v3 in hyperparametersConstantSchedule, and SimpleLinearSchedule instead of constant_fn and linear_scheduleCarRacing-v3 hyperparameters for newer Gymnasium versionn_steps parameterget_schedule_fn(), get_linear_fn(), constant_fn() are deprecated, please use FloatSchedule(), LinearSchedule(), ConstantSchedule() insteadevaluate_policy documentationtotal_timesteps parameterLunarLander and LunarLanderContinuous environment versions to v3 (@j0m0k0)Full Changelog: v2.6.0...v2.7.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
algo._dump_logs() is deprecated in favor of algo.dump_logs() and will be removed in SB3 v2.7.0
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
has_attr method for VecEnv to check if an attribute existsLogEveryNTimesteps callback to dump logs every N timesteps (note: you need to pass log_interval=None to avoid any interference)SubProcVecEnv will now exit gracefully (without big traceback) when using KeyboardInterrupt_dump_logs() to dump_logs()SubprocVecEnv and MaskablePPO by using vec_env.has_attr() (pickling issues, mask function not present)--trial-id argument of train.py.VecEnv class use to instantiate the env in the ExperimentManager--log-interval -2 (useful when logging things manually)get_hf_trained_models()net_arch, and additional fixespredict() for env that were not normalized (action spaces with limits != [-1, 1])algo._dump_logs() is deprecated in favor of algo.dump_logs() and will be removed in SB3 v2.7.0VecNormalize)set_wrapper_attr should be used nowmake_vec_env in the section on Vectorized Environments (@pstahlhofen)EveryNTimestepsVecEnv callsMultiInputPolicy (@darkopetrovic)Full Changelog: v2.5.0...v2.6.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
VecNormalize now cast normalized rewards to float32, updated bit flipping env to avoid overflow issues toopolicy_kwargs parameter (@kplers)Discrete action spaces with start!=0Full Changelog: v2.4.0...v2.5.0
Nothing published for this version
Nothing published for this version
Fixed a bug introduced in v2.4.0 where the VecVideoRecorder would override videos
VecVideoRecorder would override videosFull Changelog: v2.4.0...v2.4.1
Warning Stable-Baselines3 (SB3) v2.4.0 will be the last one supporting Python 3.8 (end of life in October 2024) and PyTorch < 2.3. We highly recommend
Warning
Stable-Baselines3 (SB3) v2.4.0 will be the last one supporting Python 3.8 (end of life in October 2024)
and PyTorch < 2.3.
We highly recommended you to upgrade to Python >= 3.9 and PyTorch >= 2.3 (compatible with NumPy v2).
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
Note
DQN (and QR-DQN) models saved with SB3 < 2.4.0 will show a warning about truncation of optimizer state when loaded with SB3 >= 2.4.0.
To suppress the warning, simply save the model again.
You can find more info in PR #1963
pre_linear_modules and post_linear_modules in create_mlp (useful for adding normalization layers, like in DroQ or CrossQ)MultiDiscrete spacesset_parameters() does not try to load the object data anymoreCallbackList now sets the .parent attribute of child callbacks to its own .parent. (will-maclean)net_arch manually set to None (@jak3122)test_buffers.py::test_device which was not actually checking the device of tensors (@rhaps0dy)CrossQ algorithm, from "Batch Normalization in Deep Reinforcement Learning" paper (@danielpalen)BatchRenorm PyTorch layer used in CrossQ (@danielpalen)target_update_interval (@jak3122)MlpPolicycopy_obs_dict method for SubprocVecEnv, remove the use of ordered dict and rename flatten_obs to stack_obsMlpPolicyFull Changelog: v2.3.2...v2.4.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Reverted torch.load() to be called weights_only=False as it caused loading issue with old version of PyTorch. #1913
torch.load() to be called weights_only=False as it caused loading issue with old version of PyTorch. #1913Full Changelog: v2.3.0...v2.3.2
torch.load() to be called weights_only=False as it caused loading issue with old version of PyTorch. https://github.com/DLR-RM/stable-baselines3/pull/1913Full Changelog: https://github.com/DLR-RM/stable-baselines3/compare/v2.3.0...v2.3.2
Reverted torch.load() to be called weights_only=False as it caused loading issue with old version of PyTorch.
Added ER-MRL to the project page (@corentinlger)
Updated Tensorboard Logging Videos documentation (@NickLucche)
- Updated SBX documentation (CrossQ and deprecated DroQ)
Cast return value of learning rate schedule to float, to avoid issue when loading model because of weights_only=True (@markscsmith)
Updated SBX documentation (CrossQ and deprecated DroQ)
Updated RL Tips and Tricks section
Warning Because of weights_only=True , this release breaks loading of policies when using PyTorch 1.13. Please upgrade to PyTorch >= 2.0 or upgrade SB
Warning
Because of weights_only=True, this release breaks loading of policies when using PyTorch 1.13.
Please upgrade to PyTorch >= 2.0 or upgrade SB3 version (we reverted the change in SB3 2.3.2)
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
TD3 and DDPG have been changed to be more consistent with SAC # SB3 < 2.3.0 default hyperparameters
# model = TD3("MlpPolicy", env, train_freq=(1, "episode"), gradient_steps=-1, batch_size=100)
# SB3 >= 2.3.0:
model = TD3("MlpPolicy", env, train_freq=1, gradient_steps=1, batch_size=256)Note
Two inconsistencies remain: the default network architecture for TD3/DDPG is [400, 300] instead of [256, 256] for SAC (for backward compatibility reasons, see report on the influence of the network size ) and the default learning rate is 1e-3 instead of 3e-4 for SAC (for performance reasons, see W&B report on the influence of the lr )
learning_starts parameter of DQN have been changed to be consistent with the other offpolicy algorithms # SB3 < 2.3.0 default hyperparameters, 50_000 corresponded to Atari defaults hyperparameters
# model = DQN("MlpPolicy", env, learning_starts=50_000)
# SB3 >= 2.3.0:
model = DQN("MlpPolicy", env, learning_starts=100)torch.load() is now called with weights_only=True when loading torch tensors,load() still uses weights_only=False as gymnasium imports are required for it to workhuggingface_sb3, you will now need to set TRUST_REMOTE_CODE=True when downloading models from the hub, as pickle.load is not safe.rollout/success_rate when available for on policy algorithms (@corentinlger)monitor_wrapper argument that was not passed to the parent class, and dones argument that wasn't passed to _update_into_buffer (@corentinlger)rollout_buffer_class and rollout_buffer_kwargs arguments to MaskablePPOtrain_freq type annotation for tqc and qrdqn (@Armandpl)sb3_contrib/common/maskable/*.py type annotationssb3_contrib/ppo_mask/ppo_mask.py type annotationssb3_contrib/common/vec_env/async_eval.py type annotationsMaskablePPO (evaluation and multi-process) (@icheered)setup.py (@power-edge)requirements.txt (remove duplicates from setup.py)MultiDiscrete and MultiBinary action spaces to PPOtrain() signature and update type hintsCrossQrender_mode="human" in the README example (@marekm4)log_interval in the base class (@rushitnshah).Full Changelog: v2.2.1...v2.3.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
> Stable-Baselines3 (SB3) v2.2.0 was yanked after a breaking change was found in GH#1751.
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
[!NOTE] Stable-Baselines3 (SB3) v2.2.0 was yanked after a breaking change was found in GH#1751. Please use SB3 v2.2.1 and not v2.2.0.
ruff for sorting imports (isort is no longer needed), black and ruff version now require a minimum versionx is False in favor of not x, which means that callbacks that wrongly returned None (instead of a boolean) will cause the training to stop (@iwishiwasaneagle)env_checker for env wrongly detected as GoalEnv (compute_reward() is defined)options at reset with VecEnv via the set_options() method. Same as seeds logic, options are reset at the end of an episode (@ReHoss)rollout_buffer_class and rollout_buffer_kwargs arguments to on-policy algorithms (A2C and PPO)_setup_learn() in OffPolicyAlgorithm (@PatrickHelm)callback.update_locals() before callback.on_rollout_end() in OnPolicyAlgorithm (@PatrickHelm)render_mode which was not properly loaded when using VecNormalize.load()SimpleMultiObsEnv (@NixGD)set_options for AsyncEvalrollout_buffer_class and rollout_buffer_kwargs arguments to TRPOgym dependency, the package is still required for some pretrained agents.--eval-env-kwargs to train.py (@Quentin18)ppo_lstm to hyperparams_opt.py (@technocrat13)pybullet_envs_gymnasium>=0.4.0optuna.suggest_uniform(...) by optuna.suggest_float(..., low=..., high=...)DDPG and TD3 algorithmsstable_baselines3/common/callbacks.py type hintsstable_baselines3/common/utils.py type hintsstable_baselines3/common/vec_envs/vec_transpose.py type hintsstable_baselines3/common/vec_env/vec_video_recorder.py type hintsstable_baselines3/common/save_util.py type hintsstable_baselines3/common/buffers.py type hintsstable_baselines3/her/her_replay_buffer.py type hints.copy() when storing new transitionsActorCriticPolicy.extract_features() signature by adding an optional features_extractor argumentsphinx_autodoc_typehints)stable_baselines3/common/off_policy_algorithm.py type hintsstable_baselines3/common/distributions.py type hintsstable_baselines3/common/vec_env/vec_normalize.py type hintsstable_baselines3/common/vec_env/__init__.py type hintsstable_baselines3/common/policies.py type hintsmypy only for checking typesFull changelog: https://github.com/DLR-RM/stable-baselines3/compare/v2.1.0...v2.2.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
stats_window_size argumentenv_checker.py warning messages for out of bounds in complex observation spaces (@Gabo-Tor)test_spaces.py testsFull Changelog: https://github.com/DLR-RM/stable-baselines3/compare/v2.0.0...v2.1.0
The deprecated online_sampling argument of HerReplayBuffer was removed
[!WARNING] Stable-Baselines3 (SB3) v2.0 will be the last one supporting python 3.7 (end of life in June 2023). We highly recommended you to upgrade to Python >= 3.8.
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo Stable-Baselines Jax (SBX): https://github.com/araffin/sbx
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
shimmy package (@carlosluis, @arjun-kg, @tlpss)online_sampling argument of HerReplayBuffer was removedstack_observation_space method of StackedObservationsevaluate_policy to prevent shadowing the input observations during callbacks (@npit)HumanOutputFormat file check: now it verifies if the object is an instance of io.TextIOBase instead of only checking for the presence of a write method.vec_env.seed(seed=seed) will only be effective after then env.reset() call.shimmy package)CarRacing-v1 to CarRacing-v2 in hyperparameters--n-timesteps argument to adjust the length of the videorecord_video steps (before it was stepping in a closed env)VecExtractDictObs does not handle terminal observation (@WeberSamuel)>=1.20 due to use of numpy.typing (@troiganto)target_update_interval (@tobirohrer)step() when checking
for Inf and NaN (@lutogniew)truncate_last_trajectory() (@lbergmann1)stable_baselines3/a2c/*.py type hintsstable_baselines3/ppo/*.py type hintsstable_baselines3/sac/*.py type hintsstable_baselines3/td3/*.py type hintsstable_baselines3/common/base_class.py type hintsstable_baselines3/common/logger.py type hintsstable_baselines3/common/envs/*.py type hintsstable_baselines3/common/vec_env/vec_monitor|vec_extract_dict_obs|util.py type hintsstable_baselines3/common/vec_env/base_vec_env.py type hintsstable_baselines3/common/vec_env/vec_frame_stack.py type hintsstable_baselines3/common/vec_env/dummy_vec_env.py type hintsstable_baselines3/common/vec_env/subproc_vec_env.py type hintsVecEnv and VecEnvWrapperseed() method return type from List to SequenceVecEnv API vs Gym APIVecEnv vs Gym envEvalCallback example (@sidney-tio)pink-noise-rl to projects pageortho_init was ignoredFull Changelog: https://github.com/DLR-RM/stable-baselines3/compare/v1.8.0...v2.0.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Replaced deprecated optuna.suggest_loguniform(...) by optuna.suggest_float(..., log=True)
[!WARNING] Stable-Baselines3 (SB3) v1.8.0 will be the last one to use Gym as a backend. Starting with v2.0.0, Gymnasium will be the default backend (though SB3 will have compatibility layers for Gym envs). You can find a migration guide here. If you want to try the SB3 v2.0 alpha version, you can take a look at PR #1327.
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
mlp_extractor (@AlexPasqua)StackedObservations (it now handles dict obs, StackedDictObservations was removed)features_extractor parameter when calling extract_features()HerReplayBufferHerReplayBuffer was refactored to support multiprocessing, previous replay buffer are incompatible with this new versionHerReplayBuffer doesn't require a max_episode_length anymorerepeat_action_probability argument in AtariWrapper.NoopResetEnv and MaxAndSkipEnv when needed in AtariWrapperVecCheckNan, the check is now active in the env_checker() (@DavyMorgan)HerReplayBufferHerReplayBuffer now supports all datatypes supported by ReplayBufferobservation_space of custom gym environments using check_env (@FieteO)stats_window_size argument to control smoothing in rollout logging (@jonasreiher)check_env in the MaskablePPO docs (@AlexPasqua)sb3_contrib/qrdqn/*.py type hintsmlp_extractor (@AlexPasqua)dtype (default to float32) to the noise for consistency with gym action (@sidney-tio)DictRolloutBuffer.add with multidimensional action space (@younik)tests/test_tensorboard.py type hinttests/test_vec_normalize.py type hintstable_baselines3/common/monitor.py type hintsetup.cg to pyproject.toml configuration fileflake8 to ruffstable_baselines3/dqn/*.py type hintsextra_no_roms option for package installation without Atari Romsload_parameters to set_parameters (@DavyMorgan)A2C docstring (@AlexPasqua)log_interval description (@theSquaredError)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
> Shared layers in MLP policy (mlp_extractor) are now deprecated for PPO, A2C and TRPO. This feature will be removed in SB3 v1.8.0 and the behavior of…
SB3 Contrib (more algorithms): https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3 (training framework): https://github.com/DLR-RM/rl-baselines3-zoo
To upgrade:
pip install stable_baselines3 sb3_contrib rl_zoo3 --upgrade
or simply (rl zoo depends on SB3 and SB3 contrib):
pip install rl_zoo3 --upgrade
Warning Shared layers in MLP policy (
mlp_extractor) are now deprecated for PPO, A2C and TRPO. This feature will be removed in SB3 v1.8.0 and the behavior ofnet_arch=[64, 64]will create separate networks with the same architecture, to be consistent with the off-policy algorithms.
Note A2C and PPO models saved with SB3 < 1.7.0 will show a warning about missing keys in the state dict when loaded with SB3 >= 1.7.0. To suppress the warning, simply save the model again. You can find more info in issue #1233
create_eval_env, eval_env, eval_log_path, n_eval_episodes and eval_freq parameters,
please use an EvalCallback insteadsde_net_arch parameterret attributes in VecNormalize, please use returns insteadVecNormalize now updates the observation space when normalizing imageswith_bias argument to create_mlpspaces.MultiBinary observationsnormalize_images=Falsenormalized_image parameter to NatureCNN and CombinedExtractorRecurrentPPO where the lstm states where incorrectly reshaped for n_lstm_layers > 1 (thanks @kolbytn)RuntimeError: rnn: hx is not contiguous while predicting terminal values for RecurrentPPO when n_lstm_layers > 1monitor_kwargs parameterProgressBarCallback under-reporting (@dominicgkerr)evaluate_actions in ActorCritcPolicy to reflect that entropy is an optional tensor (@Rocamonde)policy in BaseAlgorithm and OffPolicyAlgorithmcustom_objects workaroundmodel in evaluate_policySelf return type using TypeVarnormalize_images which was not passed to parent class in some casesload_from_vector that was broken with newer PyTorch version when passing PyTorch tensorfeatures_extractor parameter when calling extract_features()MlpExtractor (@AlexPasqua)compute_reward method, rather than by their inheritance to gym.GoalEnvCartPole-v0 by CartPole-v1 is teststests/test_distributions.py type hintsstable_baselines3/common/type_aliases.py type hintsstable_baselines3/common/torch_layers.py type hintsstable_baselines3/common/env_util.py type hintsstable_baselines3/common/preprocessing.py type hintsstable_baselines3/common/atari_wrappers.py type hintsstable_baselines3/common/vec_env/vec_check_nan.py type hints__init__.py with the __all__ attribute (@ZikangXiong)np.bool = bool so gym 0.21 is compatible with NumPy 1.24+from gym import spacesget_system_info to avoid issue linked to copy-pasting on GitHub issueenv to vec_env when environment is vectorizedmlp_extractor's dimensions (@AlexPasqua)Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →