pytorch/examples

Ambiguous code in reinforce

オープン

#297 opened on 2018/02/02

 (2 件のコメント) (0 件のリアクション) (0 人の担当者)Python (9,429 件のフォーク)batch import
good first issue

Repository metrics

Stars
 (21,634 個のスター)
PR merge metrics
 (PR metrics pending)

説明

In /reinforcement_learning/reinforce.py, line 91:

running_reward = running_reward * 0.99 + t * 0.01

The variable running_reward seems to used for record average episodic rewards(not actually average, but I think the concept is similar), 0.01, is the scalar to update the average episodic rewards and t is done step. Add some comment or refactor naming may help beginners to understand this example.

コントリビューターガイド