pytorch/examples

Ambiguous code in reinforce

Aberta

#297 aberto em 2 de fev. de 2018

 (2 comentários) (0 reação) (0 responsável)Python (9.429 forks)batch import
good first issue

Métricas do repositório

Stars
 (21.634 estrelas)
Métricas de merge de PR
 (Métricas PR pendentes)

Description

In /reinforcement_learning/reinforce.py, line 91:

running_reward = running_reward * 0.99 + t * 0.01

The variable running_reward seems to used for record average episodic rewards(not actually average, but I think the concept is similar), 0.01, is the scalar to update the average episodic rewards and t is done step. Add some comment or refactor naming may help beginners to understand this example.

Guia do colaborador