DLR-RM/stable-baselines3

[Feature Request] independently configurable learning rates for actor and critic

Ouverte

#338 ouverte le 3 mars 2021

 (11 commentaires) (1 réaction) (0 personne assignée)Python (1 407 forks)batch import
enhancementhelp wanted

Métriques du dépôt

Stars
 (6 550 étoiles)
Métriques de merge PR
 (Merge moyen 11j 13h) (3 PRs mergées en 30 j)

Description

🚀 Feature

independently configurable learning rates for actor and critic in AC-style algorithms

Motivation

In literature the actor is often configured to learn slower, such that the critics responses are more reliable. At least it would be nice if i could allow my hyperparameter optimizer to decide which learning rates he wants to use for actor or critic.

Pitch

https://github.com/DLR-RM/stable-baselines3/blob/65100a4b040201035487363a396b84ea721eb027/stable_baselines3/ddpg/ddpg.py#L12-L26

Additional context

https://spinningup.openai.com/en/latest/algorithms/ddpg.html#documentation-pytorch-version

Guide contributeur