pytorch/examples

Ambiguous code in reinforce

Ouverte

#297 ouverte le 2 févr. 2018

 (2 commentaires) (0 réaction) (0 personne assignée)Python (9 429 forks)batch import
good first issue

Métriques du dépôt

Stars
 (21 634 étoiles)
Métriques de merge PR
 (Métriques PR en attente)

Description

In /reinforcement_learning/reinforce.py, line 91:

running_reward = running_reward * 0.99 + t * 0.01

The variable running_reward seems to used for record average episodic rewards(not actually average, but I think the concept is similar), 0.01, is the scalar to update the average episodic rewards and t is done step. Add some comment or refactor naming may help beginners to understand this example.

Guide contributeur