good first issue
仓库指标
- 星标
- (21,634 个星标)
- PR 合并指标
- (PR 指标待抓取)
描述
I have a question about the weight intialization of embedding layer.
In source code of PyTorch, the weight of embedding layer is initialized by N(0, 1). In this code, the weight of embedding layer is intialized by uniform.
I trained the model two times with different initialization method, and I found that the default initialization makes the model converges too slow.
So, why the embedding is initialized by N(0, 1), which seems not a good start point?