NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4561 most downloaded on PyPI
Pytorch version of Stable Baselines, implementations of reinforcement learning algorithms.
Last release 3 months ago
15 Jun 2026
Release timing varies
gaps range from 2 weeks to 4 months
Nearly every release is documented
notes for 30 of 30 stable releases
3 versions withdrawn
withdrawn after publishing
6 years old
115 releases · first in 2020
One column per quarter.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Added deprecation warning if parameters eval_env, eval_freq or create_eval_env are used (see #925) (@tobirohrer)
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib RL Zoo3: https://github.com/DLR-RM/rl-baselines3-zoo
progress_bar argument in the learn() method, displayed using TQDM and rich packagespip install rl_zoo3)self.num_timesteps was initialized properly only after the first call to on_step() for callbacks~=4.13 to be compatible with gym=0.21eval_env, eval_freq or create_eval_env are used (see #925) (@tobirohrer)env_id parameter in make_vec_env and make_atari_env (@AlexPasqua)wrapper_class parameter in make_vec_env (@AlexPasqua)Progress bar in the learn() method, RL Zoo3 is now a package
Added progress_bar argument in the learn() method, displayed using TQDM and rich packages
Added progress bar callback
The RL Zoo can now be installed as a package ( pip install rl_zoo3 )
RL Zoo is now a python package and can be installed using pip install rl_zoo3
self.num_timesteps was initialized properly only after the first call to on_step() for callbacks
Set importlib-metadata version to ~=4.13 to be compatible with gym=0.21
Added deprecation warning if parameters eval_env , eval_freq or create_eval_env are used (see #925) (@tobirohrer)
Fixed type hint of the env_id parameter in make_vec_env and make_atari_env (@AlexPasqua)
Extended docstring of the wrapper_class parameter in make_vec_env (@AlexPasqua)
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
VecNormalize statistics (@anand-bala)Monitor to append to existing file instead of overriding (@sidney-tio)CnnLstmPolicy or MultiInputLstmPolicy with RecurrentPPO (@mlodel)PPO gives NaN if rollout buffer provides a batch of size 1 (@hughperkins)predict does not always return action as np.ndarray (@qgallouedec)EvalCallback constructor (@burakdmb)running_mean and running_var properties of batch norm layers are not updated (@honglu2875)common.OffPolicyAlgorithm initializer, where an instance instead of a class was required (@Rocamonde)forward() abstract method declaration from common.policies.BaseModel (already defined in torch.nn.Module) to fix type errors in subclasses (@Rocamonde).load() and .learn() methods in BaseAlgorithm so that they now use TypeVar (@Rocamonde)common.logger.HumanOutputFormat (@Rocamonde and @AdamGleave)DictReplayBuffer.next_observations typing (@qgallouedec)device="auto" in buffers and made it default (@qgallouedec)ResultsWriter` (used internally by Monitorwrapper) to automatically create missing directories whenfilename`` is a path (@dominicgkerr)Bug fix release
Switched minimum tensorboard version to 2.9.1
Support logging hyperparameters to tensorboard (@timothe-chaumont)
Added checkpoints for replay buffer and VecNormalize statistics (@anand-bala)
Added option for Monitor to append to existing file instead of overriding (@sidney-tio)
The env checker now raises an error when using dict observation spaces and observation keys don’t match observation space keys
Fixed the issue of wrongly passing policy arguments when using CnnLstmPolicy or MultiInputLstmPolicy with RecurrentPPO (@mlodel)
Fixed issue where PPO gives NaN if rollout buffer provides a batch of size 1 (@hughperkins)
Fixed the issue that predict does not always return action as np.ndarray (@qgallouedec)
Fixed division by zero error when computing FPS when a small number of time has elapsed in operating systems with low-precision timers.
Added multidimensional action space support (@qgallouedec)
Fixed missing verbose parameter passing in the EvalCallback constructor (@burakdmb)
Fixed the issue that when updating the target network in DQN, SAC, TD3, the running_mean and running_var properties of batch norm layers are not updated (@honglu2875)
Fixed incorrect type annotation of the replay_buffer_class argument in common.OffPolicyAlgorithm initializer, where an instance instead of a class was required (@Rocamonde)
Fixed loading saved model with different number of environments
Removed forward() abstract method declaration from common.policies.BaseModel (already defined in torch.nn.Module ) to fix type errors in subclasses (@Rocamonde)
Fixed the return type of .load() and .learn() methods in BaseAlgorithm so that they now use TypeVar (@Rocamonde)
Fixed an issue where keys with different tags but the same key raised an error in common.logger.HumanOutputFormat (@Rocamonde and @AdamGleave)
Set importlib-metadata version to ~=4.13
Fixed DictReplayBuffer.next_observations typing (@qgallouedec)
Added support for device="auto" in buffers and made it default (@qgallouedec)
Updated ResultsWriter (used internally by Monitor wrapper) to automatically create missing directories when filename is a path (@dominicgkerr)
Added an example of callback that logs hyperparameters to tensorboard. (@timothe-chaumont)
Fixed typo in docstring “nature” -> “Nature” (@Melanol)
Added info on split tensorboard logs into (@Melanol)
Fixed typo in ppo doc (@francescoluciano)
Fixed typo in install doc(@jlp-ue)
Clarified and standardized verbosity documentation
Added link to a GitHub issue in the custom policy documentation (@AlexPasqua)
Update doc on exporting models (fixes and added torch jit)
Fixed typos (@Akhilez)
Standardized the use of " for string representation in documentation
Nothing published for this version
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
register_policy helper, policy_base parameter and using policy_aliases static attributes instead (@Gregwar)CnnPolicy or MultiInputPolicy with SAC or DDPG/TD3,
share_features_extractor is now set to False by default and the net_arch=[256, 256] (instead of net_arch=[] that was before)DummyVecEnv's and SubprocVecEnv's seeding function. None value was unchecked (@ScheiklP)EvalCallback would crash when trying to synchronize VecNormalize stats when observation normalization was disabledkl_divergence check that would fail when using numpy arrays with MultiCategorical distributionpyupgradeBaseAlgorithm._wrap_env (@TibiGG)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
StopTrainingOnNoModelImprovement to callback collection (@caburu)HumanOutputFormat configurable,
depending on desired maximum width of output.VecMonitor. The monitor did not consider the info_keywords during stepping (@ScheiklP)HumanOutputFormat. Distinct keys truncated to the same prefix would overwrite each others value,
resulting in only one being output. This now raises an error (this should only affect a small fraction of use cases
with very long keys.)nn.Module calls through implicit rather than explict forward as per pytorch guidelines (@manuel-delverme)VecNormalize where error occurs when norm_obs is set to False for environment with dictionary observation (@buoyancy99)env argument to None in HerReplayBuffer.sample (@qgallouedec)batch_size typing in DQN (@qgallouedec)DictReplayBuffer (@qgallouedec)remove_time_limit_termination in off policy algorithms since it was dead code (@Gregwar)Directly Accessing The Summary Writer in tensorboard integration (@xy9485)Full Changelog: https://github.com/DLR-RM/stable-baselines3/compare/v1.4.0...v1.5.0
Nothing published for this version
…length constrain. (see PR #704) This will be a backward incompatible change (model trained with previous version of HER won't work with the new versio…
SB3 Contrib: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib
mask argument of the predict() method to episode_start (used with RNN policies only)action, done and reward were renamed to their plural form for offpolicy algorithms (actions, dones, rewards),
this may affect custom callbacks.episode_reward field from RolloutReturn() typeWarning:
An update to the HER algorithm is planned to support multi-env training and remove the max episode length constrain.
(see PR #704)
This will be a backward incompatible change (model trained with previous version of HER won't work with the new version).
norm_obs_keys param for VecNormalize wrapper to configure which observation keys to normalize (@kachayev)HerReplayBuffer currently not supported)TimeLimit)skip option to VecTransposeImage to skip transforming the channel order when the heuristic is wrongcopy() and combine() methods to RunningMeanStdset_env() with VecNormalize would result in an error with off-policy algorithms (thanks @cleversonahum)learn call, even when reset_num_timesteps is set to False (@kachayev)VecFrameStack with channel first image envs, where the terminal observation would be wrongly created.np.float32 for continuous actionsnewline="\n" when opening CSV monitor files so that each line ends with \r\n instead of \r\r\n on Windows while Linux environments are not affected (@hsuehch)device argument inconsistency (@qgallouedec)BaseAlgorithm.load docstring (@Demetrio92)load behavior in the examples (@Demetrio92)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Cap gym max version to 0.19 to avoid issues with atari-py and other breaking changes
WARNING: This version will be the last one supporting Python 3.6 (end of life in Dec 2021). We highly recommend you to upgrade to Python >= 3.7.
SB3-Contrib changelog: https://github.com/Stable-Baselines-Team/stable-baselines3-contrib/releases/tag/v1.3.0
sde_net_arch argument in policies is deprecated and will be removed in a future version.
_get_latent (ActorCriticPolicy) was removed
All logging keys now use underscores instead of spaces (@timokau). Concretely this changes:
time/total timesteps to time/total_timesteps for off-policy algorithms (PPO and A2C) and the eval callback (on-policy algorithms already used the underscored version),rollout/exploration rate to rollout/exploration_rate androllout/success rate to rollout/success_rate.get_distribution and predict_values for ActorCriticPolicy for A2C/PPO/TRPO (@cyprienc)forward_actor and forward_critic for MlpExtractorsb3.get_system_info() helper function to gather version information relevant to SB3 (e.g., Python and PyTorch version)print_system_info parameter to help debugging load issues.dtype of observations for SimpleMultiObsEnvVecNormalize to wrap discrete-observation environments to normalize reward
when observation normalization is disabled.DQN would throw an error when using Discrete observation and stochastic actionsforce_reset argument to load() and set_env() in order to be able to call learn(reset_num_timesteps=False) with a new environmentEvalCallback with two envs not wrapped the same way.setup.pydocutils issue)Nothing published for this version
Nothing published for this version
Nothing published for this version
SB3 now requires PyTorch >= 1.8.1
VecNormalize ret attribute was renamed to returnsVecNormalize where the observation filter was not updated at reset (thanks @vwxyzjn)train() and eval() (@davidblom603)gradient_steps=0 to an off-policy algorithm will result in no gradient steps being taken (vs as many gradient steps as steps done in the environment
during the rollout in previous versions)predict() by moving the preprocessing to obs_to_tensor() methodVecEnvWrapperAll customs environments (e.g. the BitFlippingEnv or IdentityEnv) were moved to stable_baselines3.common.envs folder
BitFlippingEnv or IdentityEnv) were moved to stable_baselines3.common.envs folderHER which is now the HerReplayBuffer class that can be passed to any off-policy algorithmTimeLimit)_last_dones and dones to _last_episode_starts and episode_starts in RolloutBuffer.ObsDictWrapper as Dict observation spaces are now supported her_kwargs = dict(n_sampled_goal=2, goal_selection_strategy="future", online_sampling=True)
# SB3 < 1.1.0
# model = HER("MlpPolicy", env, model_class=SAC, **her_kwargs)
# SB3 >= 1.1.0:
model = SAC("MultiInputPolicy", env, replay_buffer_class=HerReplayBuffer, replay_buffer_kwargs=her_kwargs)
channels_last from is_image_space as it can be inferred.model.logger that be set by the user using model.set_logger()logger.configure and utils.configure_logger, they now return a Logger objectLogger.CURRENT and Logger.DEFAULTwarn(), debug(), log(), info(), dump() methods to the Logger class.learn() now throws an import error when the user tries to log to tensorboard but the package is not installedDict observation space (@JadenTravnik)DictRolloutBuffer DictReplayBuffer to support dictionary observations (@JadenTravnik)StackedObservations and StackedDictObservations that are used within VecFrameStackHerReplayBuffer now supports VecNormalize when online_sampling=FalseHERreplay_buffer_class and replay_buffer_kwargs arguments to off-policy algorithmskl_divergence helper for Distribution classes (@09tangriro)num_envs > 1 (@benblack769)wrapper_kwargs argument to make_vec_env (@amy12xx)ent_coef for SAC and TQC, it was not optimized anymore (thanks @Atlis)A2C and PPO policy when using gSDE (thanks @liusida)verbose>=1 after passing verbose=0 onceflake8-bugbear to tests dependencies to find likely bugsenv_checker to reflect support of dict observation spacesbatch_size > 1 in PPO to avoid NaN in advantage normalizationProcgenEnvdocutils==0.16 to avoid issue with rtd themesave_freq definitionA2C docs (@bstee615)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
…(see PR #243 and #351) This will be a backward incompatible change (model trained with previous version of HER won't work with the new version).
First Major Version
Blog post: https://araffin.github.io/post/sb3/
100+ pre-trained models in the zoo: https://github.com/DLR-RM/rl-baselines3-zoo
stable_baselines3.common.cmd_util (already deprecated), please use env_util instead<div class="warning">
<div class="admonition-title">
Warning
</div>
A refactoring of the HER algorithm is planned together with support for dictionary observations (see PR #243 and
#351)
This will be a backward incompatible change (model trained with previous version of HER won't work with the new version).
</div>
custom_objects when loading modelsDQN predict method when using deterministic=False with image spaceNothing published for this version
Nothing published for this version
Nothing published for this version
evaluate_policy now returns rewards/episode lengths from a Monitor wrapper if one is present, this allows to return the unnormalized reward in the cas
evaluate_policy now returns rewards/episode lengths from a Monitor wrapper if one is present,
this allows to return the unnormalized reward in the case of Atari games for instance.common.vec_env.is_wrapped to common.vec_env.is_vecenv_wrapped to avoid confusion
with the new is_wrapped() helper_get_data() to _get_constructor_parameters() for policies (this affects independent saving/loading of policies)n_episodes_rollout and merged it with train_freq, which now accepts a tuple (frequency, unit):replay_buffer in collect_rollout is no more optional
# SB3 < 0.11.0
# model = SAC("MlpPolicy", env, n_episodes_rollout=1, train_freq=-1)
# SB3 >= 0.11.0:
model = SAC("MlpPolicy", env, train_freq=(1, "episode"))
VecFrameStack to stack on first or last observation dimension, along with
automatic check for image spaces.VecFrameStack now has a channels_order argument to tell if observations should be stacked
on the first or last observation dimension (originally always stacked on last).common.env_util.is_wrapped and common.env_util.unwrap_wrapper functions for checking/unwrapping
an environment for specific wrapper.env_is_wrapped() method for VecEnv to check if its environments are wrapped
with given Gym wrappers.monitor_kwargs parameter to make_vec_env and make_atari_envMonitor wrapper when possible.EvalCallback now logs the success rate when available (is_success must be present in the info dict)Logger. (@lorenz-h)DQN predict method when using single gym.Env with deterministic=Falseexplained_variance() in ppo.py and a2c.py is not correct (@thisray)HerReplayBuffer leads to an index error. (@megan-klaiber)PPO construction error in edge-case scenario where n_steps * n_envs = 1 (size of rollout buffer),
which otherwise causes downstream breaking errors in training (@decodyng)train_freq=1)np.bool with bool)VecNormalize was not normalizing the terminal observationVecTranspose was not transposing the terminal observationaction_noise was not used when using HER (thanks @ShangqunYu)train_freq was not properly converted when loading a saved modelNatureCNNtrain() method of SAC, TD3 and DQN to match SB3-Contrib.PPO when n_steps * n_envs is not a multiple of batch_size (last mini-batch truncated) (@decodyng)A2C (epsilon parameter)clip_range docstringEvalCallback docstring (thanks @tfederico)- evaluate_policy now returns rewards/episode lengths from a Monitor wrapper if one is present, this allows to return the unnormalized reward in the c
evaluate_policy now returns rewards/episode lengths from a Monitor wrapper if one is present, this allows to return the unnormalized reward in the case of Atari games for instance.
Renamed common.vec_env.is_wrapped to common.vec_env.is_vecenv_wrapped to avoid confusion with the new is_wrapped() helper
Renamed _get_data() to _get_constructor_parameters() for policies (this affects independent saving/loading of policies)
Removed n_episodes_rollout and merged it with train_freq , which now accepts a tuple (frequency, unit) :
replay_buffer in collect_rollout is no more optional
Add support for VecFrameStack to stack on first or last observation dimension, along with automatic check for image spaces.
VecFrameStack now has a channels_order argument to tell if observations should be stacked on the first or last observation dimension (originally always stacked on last).
Added common.env_util.is_wrapped and common.env_util.unwrap_wrapper functions for checking/unwrapping an environment for specific wrapper.
Added env_is_wrapped() method for VecEnv to check if its environments are wrapped with given Gym wrappers.
Added monitor_kwargs parameter to make_vec_env and make_atari_env
Wrap the environments automatically with a Monitor wrapper when possible.
EvalCallback now logs the success rate when available ( is_success must be present in the info dict)
Added new wrappers to log images and matplotlib figures to tensorboard. (@zampanteymedio)
Add support for text records to Logger . (@lorenz-h)
Fixed bug where code added VecTranspose on channel-first image environments (thanks @qxcv)
Fixed DQN predict method when using single gym.Env with deterministic=False
Fixed bug that the arguments order of explained_variance() in ppo.py and a2c.py is not correct (@thisray)
Fixed bug where full HerReplayBuffer leads to an index error. (@megan-klaiber)
Fixed bug where replay buffer could not be saved if it was too big (> 4 Gb) for python<3.8 (thanks @hn2)
Added informative PPO construction error in edge-case scenario where n_steps * n_envs = 1 (size of rollout buffer), which otherwise causes downstream breaking errors in training (@decodyng)
Fixed discrete observation space support when using multiple envs with A2C/PPO (thanks @ardabbour)
Fixed a bug for TD3 delayed update (the update was off-by-one and not delayed when train_freq=1 )
Fixed numpy warning (replaced np.bool with bool )
Fixed a bug where VecNormalize was not normalizing the terminal observation
Fixed a bug where VecTranspose was not transposing the terminal observation
Fixed a bug where the terminal observation stored in the replay buffer was not the right one for off-policy algorithms
Fixed a bug where action_noise was not used when using HER (thanks @ShangqunYu)
Add more issue templates
Add signatures to callable type annotations (@ernestum)
Improve error message in NatureCNN
Added checks for supported action spaces to improve clarity of error messages for the user
Renamed variables in the train() method of SAC , TD3 and DQN to match SB3-Contrib.
Updated docker base image to Ubuntu 18.04
Set tensorboard min version to 2.2.0 (earlier version are apparently not working with PyTorch)
Added warning for PPO when n_steps * n_envs is not a multiple of batch_size (last mini-batch truncated) (@decodyng)
Removed some warnings in the tests
Updated algorithm table
Minor docstring improvements regarding rollout (@stheid)
Fix migration doc for A2C (epsilon parameter)
Fix clip_range docstring
Fix duplicated parameter in EvalCallback docstring (thanks @tfederico)
Added example of learning rate schedule
Added SUMO-RL as example project (@LucasAlegre)
Fix docstring of classes in atari_wrappers.py which were inside the constructor (@LucasAlegre)
Added SB3-Contrib page
Fix bug in the example code of DQN (@AptX395)
Add example on how to access the tensorboard summary writer directly. (@lorenz-h)
Updated migration guide
Updated custom policy doc (separate policy architecture recommended)
Added a note about OpenCV headless version
Corrected typo on documentation (@mschweizer)
Provide the environment when loading the model in the examples (@lorepieri8)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Warning: Renamed common.cmd_util to common.env_util for clarity (affects make_vec_env and make_atari_env functions)
common.cmd_util to common.env_util for clarity (affects make_vec_env and make_atari_env functions)net_arch=dict(qf=[400, 300], pi=[64, 64]) for off-policy algorithms (SAC, TD3, DDPG)HER. (@megan-klaiber)VecNormalize now supports gym.spaces.Dict observation spacesshare_features_extractor argument to SAC and TD3 policiesmake_vec_env support the env_kwargs argument when using an env ID str (@ManifoldFR)device="cpu" is providedcheck_env not checking if the env has a Dict actionspace before calling _check_nan (@wmmc88)SAC, DDPG and TD3 when using CnnPolicy (or custom feature extractor)CnnPolicy, the passed env was not wrapped properly
(the bug was introduced when implementing HER so it should not be present in previous versions).vscode to the gitignoreCnnPoliciesNothing published for this version
Removed device keyword argument of policies; use policy.to(device) instead. (@qxcv)
device keyword argument of policies; use policy.to(device) instead. (@qxcv)BaseClass.get_torch_variables -> BaseClass._get_torch_save_params and
BaseClass.excluded_save_params -> BaseClass._excluded_save_paramstensors to pytorch_variables for claritymake_atari_env, make_vec_env and set_random_seed must be imported with (and not directly from stable_baselines3.common):from stable_baselines3.common.cmd_util import make_atari_env, make_vec_env
from stable_baselines3.common.utils import set_random_seed
unwrap_vec_wrapper() to common.vec_env to extract VecEnvWrapper if neededStopTrainingOnMaxEpisodes to callback collection (@xicocaio)device keyword argument to BaseAlgorithm.load() (@liorcohen5)get_parameters and set_parameters for accessing/setting parameters of the agentevaluate_policyclip_fraction in PPO (@diditforlulz273)device="cuda:0" (@liorcohen5)VecEnvmake_vec_env (@ManifoldFR)AlreadySteppingError and NotSteppingError that were not usedBaseClass (save/load functions close to each other, private
functions at top)save_to_zip_file function by removing duplicate codeStopTrainingOnMaxEpisodes details and example (@xicocaio)sphinx_autodoc_typehintsNothing published for this version
AtariWrapper and other Atari wrappers were updated to match SB2 ones
AtariWrapper and other Atari wrappers were updated to match SB2 onessave_replay_buffer now receives as argument the file path instead of the folder path (@tirafesi)Critic class for TD3 and SAC, it is now called ContinuousCritic
and has an additional parameter n_criticsSAC and TD3 now accept an arbitrary number of critics (e.g. policy_kwargs=dict(n_critics=3))
instead of only 2 previouslyDQN Algorithm (@Artemis-Skade)ReplayBufferpsutil is availableDDPG algorithm as a special case of TD3.BaseModel abstract parent for BasePolicy, which critics inherit from.close() method of SubprocVecEnv, causing wrappers further down in the wrapper stack to not be closed. (@NeoExtended)cloudpickle.load instead of pickle.load in CloudpickleWrapper. (@shwang)bias=False in custom policy (@rk37)dones in on-policy algorithm rollout collection. (@andyshih12).learn() methodcollect_rollout() method for off-policy algorithms_on_step() for off-policy base classnext_observations numpy arrayblack codestyle and added make format, make check-codestyle and commit-checksgSDEcommon.sb2_compat.RMSpropTFLike optimizer, which corresponds closer to the implementation of RMSprop from Tensorflow.Nothing published for this version
Nothing published for this version
Nothing published for this version
render() method of VecEnvs now only accept one argument: mode
render() method of VecEnvs now only accept one argument: mode
Created new file common/torch_layers.py, similar to SB refactoring
MlpExtractor, create_mlp, NatureCNNRenamed BaseRLModel to BaseAlgorithm (along with offpolicy and onpolicy variants)
Moved on-policy and off-policy base algorithms to common/on_policy_algorithm.py and common/off_policy_algorithm.py, respectively.
Moved PPOPolicy to ActorCriticPolicy in common/policies.py
Moved PPO (algorithm class) into OnPolicyAlgorithm (common/on_policy_algorithm.py), to be shared with A2C
Moved following functions from BaseAlgorithm:
_load_from_file to load_from_zip_file (save_util.py)_save_to_file_zip to save_to_zip_file (save_util.py)safe_mean to safe_mean (utils.py)check_env to check_for_correct_spaces (utils.py. Renamed to avoid confusion with environment checker tools)Moved static function _is_vectorized_observation from common/policies.py to common/utils.py under name is_vectorized_observation.
Removed {save,load}_running_average functions of VecNormalize in favor of load/save.
Removed use_gae parameter from RolloutBuffer.compute_returns_and_advantage.
render() method for VecEnvsseed() method for SubprocVecEnvdeterministic=Falseregister_policy to allow re-registering same policy for same sub-class (i.e. assign same value to same key).gSDE with PPO/A2C, this does not affect SACfork start method in the tests (was causing a deadlock with tensorflow)SubprocVecEnv and renderingprogress (value from 1 in start of training to 0 in end) to progress_remaining.policies.py files for A2C/PPO, which define MlpPolicy/CnnPolicy (renamed ActorCriticPolicies).VecNormalize, VecCheckNan and PPO.Nothing published for this version
Nothing published for this version
Remove State-Dependent Exploration (SDE) support for TD3
TD3logkv -> record, writekvs -> write, writeseq -> write_sequence,logkvs -> record_dict, dumpkvs -> dump,getkvs -> get_log_dict, logkv_mean -> record_mean,VecCheckNan and VecVideoRecorder (Sync with Stable Baselines)cmd_util and atari_wrappersMultiDiscrete and MultiBinary observation spaces (@rolandgvc)MultiCategorical and Bernoulli distributions for PPO/A2C (@rolandgvc)VectorizedActionNoise for continuous vectorized environments (@PartiallyTyped)EvalCallback using the loggersde_sample_freq that was not taken into account for SACBaseCallback otherwise they cannot write in the one used by the algorithmsVecEnvs with Stable-Baselinesgym>=0.17.readthedoc.yml fileflake8 and make lint commandtrain_freq and n_episodes_rollout to Off-Policy AlgorithmsTD3 example code blockNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →