good first issue
倉庫指標
- 星標
- (21,634 顆星)
- PR 合併指標
- (PR 指標待抓取)
描述
I have a question about the weight intialization of embedding layer.
In source code of PyTorch, the weight of embedding layer is initialized by N(0, 1). In this code, the weight of embedding layer is intialized by uniform.
I trained the model two times with different initialization method, and I found that the default initialization makes the model converges too slow.
So, why the embedding is initialized by N(0, 1), which seems not a good start point?