pytorch/examples
A question about weight initialization of embedding layer.
オープン
#595 opened on 2019/07/18
good first issue
Repository metrics
- Stars
- (21,634 個のスター)
- PR merge metrics
- (PR metrics pending)
説明
I have a question about the weight intialization of embedding layer.
In source code of PyTorch, the weight of embedding layer is initialized by N(0, 1). In this code, the weight of embedding layer is intialized by uniform.
I trained the model two times with different initialization method, and I found that the default initialization makes the model converges too slow.
So, why the embedding is initialized by N(0, 1), which seems not a good start point?